Available projects

Click title to see full project details

DC1 Exascale-enabled scattering codes and pion-nucleon scattering
Supervisors:
G. Koutsou (The Cyprus Institute) and D. Pleiter (University of Groningen)

Objectives: i) Optimize multi-hadron contractions for exascale technologies overcoming bottlenecks in state-of-the-art scattering applications; ii) Carry out a study of pion-nucleon ($\pi N$) scattering.

Description: Studying meson-baryon scattering and baryon resonances in Lattice QCD is an exascale challenge. It requires i) increased statistics to overcome exponentially increasing noise in baryon and multi-hadron states; ii) multi-hadron correlation functions with more Wick-contractions than single-hadron states; and iii) multiple spatial volumes of (5 to 7 fm)$^3$ and/or momenta, and physical pion mass ($m_\pi$). A bottleneck lies in the fermion contractions for multi-hadron observables, memory bandwidth-bound operations, which need to be optimized since future supercomputing architectures will not significantly increase main memory bandwidth. Rather, one expects die-stacking for large L3 caches (several GBytes) while maintaining cost-effective but not faster DDR5/LPDDR5-based memories. The project will implement data layout transformations to optimize contractions for CPUs and GPUs so that the tightly coupled CPU--GPU resources on exascale systems are fully utilized.

Expected results: i) Meson-baryon contraction codes optimized for the memory characteristics of exascale systems expected by mid-term (2 years), and demonstrated in a production setup on current systems, such as Jupiter; ii) Study of $\pi N$ scattering for isospin 1/2 and 3/2 using publicly available MILC ensembles with physical $m_\pi$.

Secondments: NVIDIA labs at FZJ.

DC2 Next-generation CPU and accelerator technologies for exploring the sea quark content of the nucleon
Supervisors:
D. Pleiter (University of Groningen) and S. Bacchio (The Cyprus Institute)

Objectives: i) Enable outer-product vector and matrix instructions; ii) compute nucleon matrix elements, including the nucleon $\sigma$-terms and the strange quark contributions to the electromagnetic form factors to high precision.

Description: To go beyond recent LQCD successes and increase further the precision in baryon studies we need to address new challenges that emerge from i) using multi-GPUs and ensembles with small lattice spacing and large volumes that require compromises due to the relatively limited GPU memory, e.g. in the number of low modes that can be stored, and ii) the need to implement next-generation CPU and accelerator technologies, based e.g. on ARM or RISC-V, that most LQCD codes are lacking. DC2 will address these challenges by developing optimizations to use ARM and RISC-V processors, enabling SME instructions in the computationally demanding kernels of the DD$\alpha$AMG multigrid code and quark loop calculation, which complements other advancements in multi-grid solvers of (DC3).

The resulting codes will be used to study the nucleon sea structure that requires computation of fermion loops, and in particular nucleon $\sigma$-terms. These are of high relevance to ongoing and future experiments at major facilities such as Jefferson Lab, CERN and EIC.

Expected results: i) Production-ready LQCD codes taking advantage of SME outer-product instructions; ii) High-statistics results on the nucleon tensor and scalar charges and $\sigma$-terms as well as on the strange nucleon EM form factors in the continuum limit will be obtained.

Secondments: M11-M13 at HPE to be trained on emerging supercomputing technologies.

DC3 Efficient multigrid solvers for multiple Krylov inversions to compute electromagnetic corrections in neutron $\beta$-decay
Supervisors:
A. Frommer (University of Wuppertal) and C. Alexandrou (University of Cyprus)

Objectives: i) Assess and integrate different approaches for block Krylov solvers within multigrid methods for LQCD for efficient inversion of multiple right-hand-sides (RHS) and multiple mass shifts; ii) Apply the accelerated solver to compute four-point correlators and electromagnetic (EM) corrections to neutron $\beta$-decay achieving sub-percent precision.

Description: Multigrid solvers provide indispensable acceleration over conventional Krylov methods in state-of-the-art LQCD calculations, even though it underutilizes GPU resources due to its relatively low arithmetic intensity. We will extend simultaneous multiple RHS solves by going beyond the simple batched schemes and by developing block Krylov solvers for the coarse-grid operator. Specifically, we will i) implement schemes where all iterates for all RHS are taken from the sum of all Krylov spaces, ii) integrate block-Krylov methods into the coarse-grid solve, iii) investigate their effectiveness and scalability on modern GPU systems, and iv) optimize the coarse-grid correction using deflation-restart techniques to compute low eigenspaces dynamically during the Krylov process. This optimized solver will be used to calculate, for the first time within LQCD, EM corrections to neutron $\beta$-decay, which have become relevant since the axial charge can now be calculated to the percent level. The axial charge provides one of the most precise tests of the SM.

Expected results: i) Implementation of robust block Krylov solvers with deflation-restart strategies within QUDA, and benchmarks demonstrating their scalability and efficiency on GPUs; ii) Controlled evaluation, for the first time, of EM corrections to the neutron axial charge using four-point functions.

Secondments: ControlExpert

DC4 Calculation of parton distribution functions via Mellin moments and Machine Learning-improved Wilson flow
Supervisors:
H. Panagopoulos (University of Cyprus) and J. Finkenrath (University of Wuppertal)

Objectives: i) Improve the Wilson flow of the gauge and fermion fields using Machine Learning (ML)-improved resulting in higher accuracy for the moments at lower cost; ii) Compute higher Mellin moments of the nucleon and reconstruct the three Parton Distribution Functions (PDFs)

Description: In Lattice QCD, nucleon PDFs are computed either via spatial matrix elements of non-local operators at large nucleon boosts or reconstructed via Mellin moments computed from matrix elements of local operators. The former suffers from large systematics, while the latter is presently limited to the lowest four moments. Going beyond the fourth moment is essential and timely, since these quantities are central to current and future experiments e.g. at the future EIC at BNL. However, computing higher than the fourth moment is not possible due to severe operator mixing caused by the breaking of O(4) symmetry on the lattice. To address this, we will extend a newly proposed flow-based approach, that has been accompanied with a proof-of-concept application for the pion Mellin moments up to sixth. We will go beyond this initial proof-of-concept in two key directions: i) employ higher-order continuous flows and machine-learned flow transformations more effectively restoring O(4) symmetry and thus leading to a more reliable continuum extrapolation using coarser lattice spacings at reduced computational cost, and ii) apply the improved flow-based approach to the nucleon, for the first time. Improving the fidelity of results on higher moments will allow for reliable reconstruction of the proton PDFs.

Expected results: i) ML-optimized flows that effectively restore continuum symmetries in both gauge and fermion sectors; ii) First calculation of nucleon Mellin moments beyond the fourth with ETMC ensembles and $m_\pi=260$ MeV; iii) Benchmarking of novel inverse-problem methods, such as Bayesian and Gaussian processes for PDF reconstruction; iv) Direct reconstruction of nucleon unpolarized, helicity, and transversity PDFs with improved precision. PDFs serve as crucial input for processes in searches beyond the Standard Model e.g. at CERN.

Secondments: NVIDIA Labs at FZJ

DC5 Quantum inspired methods for out-of-equilibrium dynamics in high dimensional Lattice Gauge Theories
Supervisors:
E. Rico Ortega (University of the Basque Country) and S. Montangero (University of Padova)

Objectives: i) Develop novel tensor network methods to simulate real-time dynamics in high dimensional Lattice Gauge Theories (LGTs); ii) Apply these algorithms to simulate scattering processes and string breaking.

Description: LGTs have enabled the computation of static properties of quantum field theories, such as mass spectra and thermodynamic behavior in strongly interacting systems using Monte Carlo simulations. However, real-time dynamics and finite chemical potential cannot be studied due to the sign problem. Quantum and quantum-inspired approaches provide a promising new approach as demonstrated by groundbreaking results in low-dimensional abelian LGTs. Building on the recently developed augmented aTTN, we will design novel methods to improve efficiency and accuracy in capturing complex entanglement. The resulting software will be integrated into HPC platforms with multi-node, multi-GPU capabilities, enabling large-scale tensor network simulations. These approaches will enable the study of real-time 2D dynamics in LGTs, such as scattering, string breaking and hadron formation from quarks and gluons, thus going beyond the state-of-the-art.

Expected results: i) Novel Tensor Network algorithms for simulating real-time dynamics of high-dimensional LGTs; ii) New physical insights on key high-energy problems, namely determination of scattering cross-sections in non-perturbative regimes at strong coupling, exploring the emergence of bound states and resonances, and quantifying the quantum correlations generated in scattering events.

Secondments: CINECA and CERN

DC6 Turbulence super-resolution, based on combining deep learning and turbulent statistics conditioning
Supervisors:
K. Um (Télécom Paris-Institut Mines-Télécom) and A. Lanotte (Consiglio Nazionale delle Ricerche)

Objectives: Generate high-resolution snapshots of turbulent flows from low-resolution inputs by exploiting the statistical self-similarity of the turbulent cascade, effectively filling in details at scales that are not dynamically resolved. To this end, i) generative models will be scoped with appropriate conditioning on the statistical properties of the flow and wavelet-based learning models; and ii) transformer-based models will be investigated with a mixture of local and global attention mechanisms for learning under-resolved details.

Description: Enhancing resolution in turbulent flow simulations, while retaining its key statistical properties, is desirable since these tend to be lost at insufficient resolutions. The project will help isolate high-resolution details within distinct spatial regions, surpassing state-of-the-art methods that have thus far been limited to low-resolution datasets. To this end, generative models, which have been successfully deployed for image and natural language processing, will be explored and adapted to architectures that are better suited for turbulent flows. A set of representative ML models and their variants (e.g., the Latent Diffusion Model and Transolver) will be examined for super-resolution tasks on existing turbulence simulation datasets, such as the Smart-Turb software infrastructure, maintained by collaborators at the University of Rome Tor Vergata, the PDEBench repository, and the Johns Hopkins Turbulence Database. The ML models will be conditioned on the statistical properties of the flow and wavelet-based learning workflows will be explored to enrich wavenumber/scale characterization of the turbulent fields. New tokenization and attention mechanisms informed by flow statistics will be developed within the transformer architecture.

Expected results: i) Advanced ML tools for turbulent fluid dynamics, leading to, e.g., new conditioning, tokenizing, and attention mechanisms for deep learning models suitable for turbulence super-resolution; ii) qualitative and quantitative assessment of new ML-based super-resolution, multi-scale models for strongly correlated fields.

Secondments: E4

DC7 Lagrangian particle spatial distribution on coarse-grained flow fields: a Machine Learning approach
Supervisors:
A. Lanotte (Consiglio Nazionale delle Ricerche) and K. Um (Télécom Paris-Institut Mines-Télécom)

Objectives: Develop Machine Learning (ML) strategies to generate both 2D (and possibly 3D) turbulent Eulerian velocity fields and the correct distribution of inertial particles that evolve in these fields, enabling: i) the production of unstructured point clouds using ML approaches (e.g. Graph Neural Networks (GNN) and PointNet); and ii) their conditional deployment, by querying the spatial distribution of particles based on coarse-grained Eulerian fields.

Description: In various applications in fluid dynamics, flow information is often available on coarse grids, while there is a need to know the evolution of the spatial distribution of discrete particles on much finer grids. The proposed research will enable efficient numerical modelling in LES in the presence of Lagrangian particles (e.g., particles in the atmosphere, microplastics or sediments in the ocean), allowing, e.g., to evolve only the LES Eulerian fields, without the need to explicitly evolve the particle distributions. This will be accomplished by utilizing synergistically state-of-the-art ML models for CFD and point-cloud-based ML approaches for evolving the distribution of Lagrangian particles. To this end, GNN and PointNet approaches will be trained to learn turbulent statistical properties, such as scale invariance for both the velocity fields and the particle spatial distributions.

Expected results: i) Advanced ML tools for multi-phase turbulent flow description, leading to new approaches to the modelling of particle transport in poorly resolved velocity fields (as in LES); ii) qualitative and quantitative assessment of new ML-based approaches for inertial particle transport, relevant to sub-grid scale modelling.

Secondments: E4

DC8 Mechanistic Interpretability for data-driven models in fluid dynamics
Supervisors:
A. Gabbana (university of Ferrara) and M. Nicolaou (University of Cyprus)

Objectives: Develop and apply interpretability methodologies for Machine Learning (ML) models in Computational Fluid Dynamics (CFD), using ideas from the natural language and vision domains to: i) enhance interpretability of trained ML models, through linking internal representations to physically relevant structures; and ii) improve generalizability across flow regimes; and iii) assess the applicability of the learned representations for other engineering systems.

Description: Current ML models in CFD focus on predictive accuracy, but remain largely black-box, offering limited insight into the physical mechanisms they capture. Physics informed Machine Learning has improved physical consistency by embedding governing equations and constraints; yet systematic frameworks for interpretability are still lacking. Early approaches, such as combining U-net architectures with SHAP (SHapley Additive exPlanations) in turbulence, provide only shallow post-hoc explanations without uncovering deeper computational mechanisms. In contrast, mechanistic interpretability methods have advanced significantly in language and vision domains, providing a new paradigm for granular and precise understanding of ML models. The project aspires to bridge this gap, introducing mechanistic interpretability for CFD through Sparse AutoEncoders (SAEs) and their variants, including multilinear mixture-of-experts ($\mu$MoE) and mixture-of-decoders (MoD) architectures. Candidate fluid systems will be selected from high-fidelity datasets and pre-trained ML models provided by the University of Ferrara, reflecting ongoing computational campaigns or existing collections.

Expected results: i) Deliver interpretable, feature-level insights into fluid AI models, demonstrating mappings between latent dimensions and coherent flow phenomena; ii) apply interpretability techniques in data-driven fluid dynamics models and highlight their role in building trustworthy, generalizable models.

Secondments: RetailZoom

DC9 Integrating Machine Learning and biophysical simulations for accurate free energy predictions
Supervisors:
P. Carloni (Forschungszentrum Jülich) and A. V. Vargiu (University of Cagliari)

Objectives: i) Develop a massively parallel, Machine Learning-based Free Energy Prediction (FEP) approach to obtain protein-ligand binding free energies at the Quantum Mechanics/Molecular Mechanics (QM/MM) level; ii) Validate the ML-based FEP method on selected examples.

Description: In contrast to traditional alchemical perturbation, which relies on force fields and many MD simulations in which one ligand is slowly transformed into another, our ML-based approach uses a configurational map that can transform the Boltzmann distribution of the first ligand into that of the second, bypassing the need to run multiple MD simulations. This approach can compute affinities with quantum level accuracy at moderate cost. We will integrate our massively parallel MiMiC QM/MM framework into our ML-based method to predict ligand binding affinity differences with QM/MM accuracy. We will validate the approach on pharmacologically relevant neuronal proteins for which ligand/target binding free energies are known and for which the FZJ group's lab has already performed enhanced sampling calculations using traditional methods. These include G protein-coupled receptor, such as the adenosine A2A receptor and macrodomains. The newly developed workflows will complement the activities of DC10.

Expected results: i) An HPC workflow for ML-aided QM/MM ligand affinity prediction with demonstrated scalability on current supercomputers; ii) Affinity predictions of ligands towards their receptors.

Secondments: CINECA

DC10 Integration of simulations and Machine Learning algorithms: accurate descriptors from enhanced conformational sampling
Supervisors:
A. V. Vargiu (University of Cagliari) and P. Carloni (Forschungszentrum Jülich)

Objectives: i) Set up an automated pipeline to predict druggable protein conformations; ii) Map protein mutations on the correct isoform and predict the impact of mutations on conformational equilibrium of proteins; iii) Determine molecular descriptors of ligand-protein interactions to be used in ML approaches for structure-based drug design.

Description: Recent investigations demonstrated that current ML-based methods struggle to predict allosteric interactions, are often limited to the structure of protein canonical isoforms and are relatively insensitive to mutations. Physics-based methods can be used to provide accurate structural data for challenging targets and to rationalize the effect of mutations, thus enriching the diversity of training sets and ultimately boosting the predictive power of ML-based algorithms. We will integrate methods and pipelines to set up a fully automated computational workflow for the accurate description of ligand-protein binding events, as well as to predict the effect of mutations in different isoforms associated with neurodevelopmental diseases. The results, including a set of molecular descriptors associated with collected protein-ligand interactions (also coming from DC09), will be collected in a publicly accessible database.

Expected results: i) A modular and automated workflow to accurately predict ligand-protein interactions accounting for different protein isoforms and flexibility of targets. The workflow will also provide a wide and customizable range of molecular interaction descriptors; ii) A database of protein and protein-ligand complexes conformations with associated molecular descriptors; iii) An assessment of the molecular mechanism by which specific mutations are involved in neural development.

Secondments: ControlExpert

DC11 Tensor networks for biomolecular configuration space analysis
Supervisors:
S. Montangero (University of Padova) and S. Maniscalco (University of Helsinki)

Objectives: i) Develop quantum and quantum-inspired algorithms for optimizing over equational theories based on rewriting systems; ii) Apply these algorithms to quantum circuit compilation by exploring equivalence classes of circuits; iii) Model and analyze configuration spaces of biomolecules and/or polymers with interacting species.

Description: A new class of quantum and quantum-inspired algorithms for symbolic reasoning and optimization will be implemented targeting use cases from polymers, proteins, or classes of drug-like molecules. It builds on recent advances in quantum algorithms for circuit compilation, where equivalence classes of quantum circuits are explored using rewriting rules encoded in quantum dynamics. We will extend these algorithms to prepare and manipulate quantum states encoding all expressions equivalently under a given rewriting system. We will begin by formalizing the connection between string rewriting systems and quantum Hamiltonians whose ground states encode equivalence classes of symbolic expressions. The use of symbolic rewriting enables connections with formal languages and rule-based modeling in biomolecular and complex systems. Quantum algorithms based on variational and adiabatic techniques will be developed to prepare these states, referred to as orbit states. Tensor network simulations will be used to study the performance and scalability of these methods on representative rewriting systems for different applications both in compilation problems and systems.

Expected results: i) A new class of quantum and quantum-inspired algorithms for symbolic reasoning tasks; ii) Application prototypes showing how rewriting-based quantum dynamics can be used for quantum circuit compilation and for analyzing complex configuration spaces in polymer physics and biomolecular systems.

Secondments: ALGO

DC12 Tensor-network-based Ansätze and hardware-aware compilation methods for ground-state preparation for quantum chemistry
Supervisors:
Z. Zimboras (University of Helsinki) and S. Montangero (University of Padova)

Objectives: Develop: i) TTN/aTTN Ansätze tailored to molecular electronic-structure Hamiltonians for accurate ground-state preparation; and ii) quantum and quantum-inspired compilation techniques for optimized quantum circuits for ground-state preparation of quantum chemistry Hamiltonians on real devices.

Description: Quantum algorithms for electronic-structure calculations, such as VQE and QPE, that build fermion-space Ansätze, map them to qubits and enforce particle-number and spin symmetries lead to large shot counts, noise sensitivity, and barren plateaus. Additional constraints come from the specific hardware. We will extend TTN/aTTN capabilities in the Quantum Tea suite developed by UNIPD to obtain accurate ground-state approximations for molecular Hamiltonians, extract warm-start parameters and reduced data for circuit design. We will pair this with hardware-specific fermion-to-qubit mappings and compilation techniques with a classical algorithm for fermionic circuits inspired by Majorana Propagation to preselect structures and parameters in the quantum circuit to produce compact, hardware-aware state-preparation circuits. The development of techniques for simulating and compiling quantum circuits will address: i) state-preparation circuit and pre-optimization of its parameters via classical simulations; ii) pre-compilation of the circuit to the device (mapping, routing/SWAP scheduling, native gates); and iii) pre-screening observables and grouping strategies to cut on-device measurements. These techniques will help develop chemistry-aware TTN/aTTN ground-state solvers and hardware-specific compilation that minimize depth and measurement cost while maintaining trainability on current hardware.

Expected results: i) TTN/aTTN ground-state solvers achieving (near) chemical-accuracy on benchmark molecules; and ii) a hardware-specific compilation and measurement pipeline targeting quantum chemistry circuits.

Secondments: ALGO