All invited talks run in two parallel tracks — Session A in 105 Jordan and Session B in 101 Jordan — across 18 parallel slots of 30 minutes each. Parallel sessions are deliberately paired so that the two concurrent topics are clearly distinct. The opening remarks, lightning talks and poster session are held jointly, and a 5-minute transition follows the opening remarks so attendees can reach their session room. Click any session to jump to the detailed program with abstracts.
Jump to: Day 1 – Saturday, Sept 26 Day 2 – Sunday, Sept 27 Poster Presentations
| Time | Session A — 105 Jordan | Session B — 101 Jordan |
|---|---|---|
| 8:50 – 9:00 am | Opening Remarks 📍 105 Jordan | |
| 9:00 – 9:05 am | Move to session rooms | |
| 9:05 – 10:35 am | 1AData-Driven Modeling & Digital TwinsChair: Guang Lin9:05 Houman Owhadi — Graphical conditional generative modeling for digital twin modeling9:35 Dongbin Xiu — Modeling Observable Dynamics: Flow Map Learning, Delay Representation and Digital Twins10:05 Tan Bui — Real-time SciML-based forecast, calibration, and optimal control algorithms for Digital Twins | 1BSparse Data, Streaming & Probabilistic LearningChair: Zhiliang Xu9:05 Daniel Tartakovsky — Learning from Sparse Data9:35 Alireza Doostan — Randomized Sketching for Streaming Data Analytics10:05 Siting Liu — From ICON to GenICON: Probabilistic Operator Learning with Uncertainty Quantification |
| 10:35 – 11:00 am | Coffee Break | |
| 11:00 am – 12:00 pm | 2AAI Agents for Scientific DiscoveryChair: Daniele Schiavazzi11:00 Francisco Villaescusa — The Denario project11:30 Qile Jiang — Building and analyzing collaborative multi-agents for emergent discovery in scientific machine learning | 2BGeometry, Sampling & High-Dimensional InferenceChair: Di Qi11:00 Yifan Chen — Accelerating Sampling and Generative Diffusions for High-Dimensional Scientific Inference11:30 Anna Little — A Geometric Framework for Dimension Reduction and Clustering via Path Metrics |
| 12:00 – 1:00 pm | Lunch Break | |
| 1:00 – 2:30 pm | 3AAgentic AI & Foundation Models for Science & EngineeringChair: Daniele Schiavazzi1:00 Shivam Barwey — A verifiable automation approach for the analysis of high-fidelity fluid dynamics simulations1:30 Marta D'Elia — Minimizing the Thought-to-Thing time in engineering workflows2:00 Paul Brenner & Charles Vardeman — Human and AI Teaming for Discovery - Shared Memory | 3BMathematical Foundations of Machine LearningChair: Zecheng Zhang1:00 Wenjing Liao — Understanding Neural Scaling Laws of Transformers for In-Context Learning1:30 Guowei Wei — Mathematical AI for Biosciences2:00 Rahul Parhi — What Kinds of Functions Do Neural Networks Learn? Low-Norm vs. Flat Solutions |
| 2:30 – 3:00 pm | Coffee Break | |
| 3:00 – 3:30 pm | 3AAgentic AI & Foundation Models for Science & Engineering (continued)Chair: Daniele Schiavazzi3:00 Meng Jiang — Advances in Molecular Graph Foundation Models | 3BMathematical Foundations of Machine Learning (continued)Chair: Zecheng Zhang3:00 Elizabeth Newman — Boost Like a (Var)Pro: Trust-Region Gradient Boosting via Variable Projection |
| 3:30 – 4:30 pm | 4APDE–ML Coupling & Operator Learning for EngineeringChair: Guang Lin3:30 Panos Stinis — Stabilizing PDE-ML systems with applications to fluid dynamics4:00 Nick Winovich — Operator Learning for Non-Destructive Testing of Mechanical Components | 4BReliable Prediction & Uncertainty QuantificationChair: Di Qi3:30 Wenrui Hao — Toward Reliable Scientific Machine Learning: Identifiability and Neural Operator Learning4:00 Alexander Scheinker — Round-Trip Consistency: Error Bars for Digital Twins via Bidirectional Diffusion Models that Can Predict Their Own Rollout Errors |
| 4:30 – 5:00 pm | 5ALearning-Enhanced Numerical Methods for PDEsChair: Zhiliang Xu4:30 Li Wang — Learning-enhanced structure preserving methods for kinetic plasma models | 5BStochastic Modeling, Transport & Uncertainty QuantificationChair: Di Qi4:30 Gianluca Geraci — Multi-fidelity triangular transport formulations |
| 5:00 – 6:00 pm | Lightning Talks I — 6 talks × 10 min 📍 105 JordanChair: Zecheng Zhang | |
| 6:00 – 7:30 pm | Poster Session and Speakers Banquet 📍 Jordan Hall | |
| Time | Session A — 105 Jordan | Session B — 101 Jordan |
|---|---|---|
| 9:30 – 10:30 am | 5ALearning-Enhanced Numerical Methods for PDEs (continued)Chair: Zhiliang Xu9:30 Matthew Zahr — Optimization-Based Simulation of High-Speed Flows10:00 Yingjie Liu — Neural Networks with Local Converging Inputs (NNLCI) for Predicting Smooth and Non-smooth PDE Solutions with Minimal Training Data, Strong Generalization, and Drastically Reduced Complexity | 5BStochastic Modeling, Transport & Uncertainty Quantification (continued)Chair: Di Qi9:30 Erhan Bayraktar — Analytical Approach to Continuous-Time Causal Optimal Transport10:00 Xiaofan Li — Structure-Aware Variational Learning of a Class of Generalized Diffusions |
| 10:30 – 11:00 am | Coffee Break | |
| 11:00 am – 12:00 pm | 6AOperator Learning: Architectures & In-Context MethodsChair: Zecheng Zhang11:00 Yue Yu — Smoother or rougher: how to generate data to learn nonlocal operators?11:30 Guang Lin — LegONet: Plug-and-Play Structure-Preserving Neural Operator Blocks for Compositional PDE Learning | 6BSurrogate Models, Reduced-Order Methods & Efficient AlgorithmsChair: Daniele Schiavazzi11:00 Ibrahim Ekren — Consensus-based algorithms for min–max problems: uniform-in-time propagation of chaos11:30 Patrick Brewick — Toward Data-Efficient Scientific Machine Learning for Nonlinear Dynamics: Active Solution Operators and Bi-Fidelity Generative Models |
| 12:00 – 1:00 pm | Lunch Break | |
| 1:00 – 2:00 pm | 6AOperator Learning: Architectures & In-Context Methods (continued)Chair: Zecheng Zhang1:00 Jonathan Siegel — Operator-to-operator learning using DeepOSets and FNOSets1:30 Rongjie Lai — Self-supervised In-context Operator Learning on Probability Measure Space: Theory and Applications | 6BSurrogate Models, Reduced-Order Methods & Efficient Algorithms (continued)Chair: Daniele Schiavazzi1:00 Alexandros Taflanidis — Leveraging scientific machine learning techniques to support planning and emergency response management for storm surge risk1:30 Qi Di — Reduced-order models and data assimilation for prediction and uncertainty quantification of multiscale turbulent systems |
| 2:00 – 3:00 pm | Lightning Talks II — 6 talks × 10 min 📍 105 JordanChair: Daniele Schiavazzi | |
| 3:00 pm | Departure | |
Digital twin modeling, including control and data assimilation under model uncertainty, often faces an open-ended fidelity problem: adding variables, data streams, and time scales can indefinitely increase model complexity, ultimately producing systems that are difficult to maintain, validate, interpret, and use for stress or safety testing. As an alternative, one can seek parsimonious stochastic surrogate models built only on the variables needed to describe the relevant quantities of interest. We introduce a framework for discovering such variables from observational data by identifying which candidate inputs influence the full conditional law of a target quantity, rather than only its conditional mean. This distinction is essential in stochastic, coarse-grained, or partially observed systems, where dependencies may appear through changes in variability, tail behavior, multimodality, or uncertainty rather than through deterministic functional relationships. The framework couples conditional generative modeling, which learns the conditional distribution of the target given candidate inputs, with Gaussian-process-based analysis of variance (through kernel mode decomposition), which enables iterative pruning of non-influential inputs and interpretable structure discovery. In control settings, the resulting surrogate can be interpreted as a learned Markov decision process: the method identifies not only a transition model, but also the state, action, and memory variables needed to make the learned dynamics effectively Markovian. Across examples involving stochastic dynamical systems, missing variables, PDE control, reinforcement learning, and financial data, the discovered structures yield interpretable stochastic surrogates whose downstream performance is comparable to models trained on the full variable set. This talk is based on joint work with Zongren Zou, Theo Bourdais and Ricardo Baptista.
We present a mathematical and numerical framework for modeling dynamics of observables in complex systems, termed Flow Map Learning (FML). The method constructs discrete-time evolution operators directly from data, enabling accurate prediction of the observables dynamics without the need to solve the underlying governing equations. We establish that, for a broad class of systems, observable dynamics admit finite-dimensional delay representations, leading to closed evolution equations with minimal memory. This perspective provides a foundation for incorporating temporal dependence in data-driven models and fast predictions of observable dynamics, particularly for Digital Twin applications where real-time prediction and control are critical. We then present various numerical examples to demonstrate the efficacy of FML for long-time prediction of observable dynamics.
Digital twins (DTs) are high-fidelity virtual representations of physical systems and processes. At their foundation lie mathematical and physical models that describe system behavior across multiple spatial and temporal scales. A central purpose of DTs is to enable “what-if” analyses through hypothetical simulations, supporting lifecycle monitoring, parameter calibration against observational data, and systematic uncertainty quantification (UQ). For DTs to serve as a reliable basis for real-time forecasting, optimization, and decision-making, they must reconcile two traditionally competing requirements: mathematical rigor and physical fidelity, and computational efficiency at scale. This has motivated a new generation of approaches that combine classical tools from numerical analysis, partial differential equations, inverse problems, and optimization with the expressive power of Scientific Machine Learning (SciML). In this talk, I will outline a principled pathway from traditional computational mathematics to rigorously grounded SciML. I will then present recent Scientific Deep Learning (SciDL) methods for forward modeling, inverse and calibration problems, and uncertainty quantification, emphasizing mathematical structure, stability, and generalization. Both theoretical results and numerical demonstrations will be shown for representative problems governed by transport, heat, Burgers, Euler (including transonic and hypersonic regimes), and Navier–Stokes equations.
The development of efficient surrogates of partial differential equations (PDEs) is a critical step toward scalable modeling of complex, multiscale systems-of-systems. Convolutional neural networks (CNNs) have gained popularity as the basis for such surrogate models due to their success in capturing high-dimensional input–output mappings and the negligible cost of a forward pass. However, the high cost of generating training data—typically via classical numerical solvers—raises the question of whether these models are worth pursuing over more straightforward alternatives with well-established theoretical foundations such as Monte Carlo (MC) methods. To reduce the cost of data generation, we propose training a CNN surrogate model on a mixture of high and low fidelity data. These data are generated as numerical solutions obtained on fine and coarse meshes or as a (d−1)-dimensional approximation of the d-dimensional problem. We demonstrate our approach on a multiphase flow test problem, using transfer learning to train a dense, fully convolutional encoder-decoder CNN on the two classes of data. Numerical results from a sample uncertainty quantification (UQ) task demonstrate that our surrogate model outperforms MC with several times the data generation budget. This presentation is based on two publications: Song and Tartakovsky, J. Mach. Learn. Model. Comput., 3(1), 31-47, 2021; and Propp and Tartakovsky, J. Mach. Learn. Model. Comput., 6(2), 13-27, 2025.
Large-scale scientific data streams – from high-fidelity simulations to sensor measurements – now arrive faster than they can be stored, whether the goal is compressing them for later reconstruction or discovering the governing equations that produced them. Streaming and in situ methods aim to address this limitation by never retaining the entire raw high-dimensional record. This talk presents randomized sketching as a common tool for that constraint, applied as early as possible in the pipeline so a fixed, reusable random projection compresses each incoming snapshot on arrival, a lightweight buffer accumulates primarily the compressed record, and any operation coupling different time steps is deferred to a single finalization pass over the retained sketches. This principle recurs across low-rank data reduction, neural representation learning, and governing-equation discovery. This is a joint work with Angran Li, Stephen Becker, Cooper Simpson, and Kyuwon Lee.
In-context operator networks (ICON) learn solution operators for ODEs/PDEs by conditioning on example initial/boundary data and their solutions. I will present a probabilistic interpretation: ICON approximates the posterior predictive mean given the context. Using random differential equations, this connects ICON to Bayesian inference and motivates GenICON, a generative extension that samples from the posterior predictive distribution for principled uncertainty quantification. The framework unifies operator learning under a Bayesian lens and provides uncertainty-aware predictions.
Science advances by formulating and testing hypotheses, collecting and analyzing data, and drawing conclusions—yet much of a scientist’s time is spent in tasks such as coding analyses, writing and revising text, reviewing the literature, and learning new concepts. Can recent advances in AI help reclaim some of that time? In this talk, I will show how large language models and AI agents may help scientists with these tasks. I will first describe what AI agents are and their applications in science. Next, I will present Denario, a complex, publicly available, multi-AI-agent system designed to function as a research assistant. Developed and evaluated by a diverse team of scientists, mathematicians, and philosophers, Denario is an interdisciplinary tool capable of generating ideas, searching the literature, developing research plans, writing and executing code, crafting plots, drafting and reviewing scientific papers. To showcase its capabilities, I will conduct a live demonstration tasking the system with turning a dataset into ideas, codes, plots, and paper drafts in real time. I will then present and discuss some of the documents generated by Denario in disciplines such as astrophysics, biology, biophysics, biomedical informatics, chemistry, machine learning, material science, mathematical physics, medicine, neuroscience, planetary physics, and quantum physics. I’ll close by discussing how tools like Denario may help researchers accelerate scientific research and invite the audience to a broad discussion on the benefits and risks of this technology.
Collaborative multi-agent systems like AgenticSciML have shown that teams of LLM agents can propose, critique, and refine scientific machine learning methods, discovering solutions that outperform human-designed baselines by orders of magnitude. As we delegate more of the scientific process to such systems, in this talk, we will also present ways that we can analyze the internal dynamics of multi-agent collaboration and understand they reach their conclusions. Together, these directions point toward multi-agent systems that are not only powerful engines of scientific discovery but also transparent and reliable enough to be genuine research partners.
Probability measures are central computational objects in scientific computing and generative AI, where scientific conclusions often depend not only on a single sample but on calibrated distributions. The typical computational route to such a distribution is to evolve a probability law along a dynamics of measures, specified by a model, learned from data, or both. This talk presents some recent work on accelerating these dynamics through parallel, multiscale, and geometric ideas, with theoretical guarantees and high-dimensional applications such as probabilistic imaging and forecasting.
In this talk, I will introduce a geometric framework for unsupervised learning that leverages power-weighted path metrics to effectively embed high-dimensional data while preserving its intrinsic geometric structure. I will begin by demonstrating how discrete, sample-based path distances converge to a density re-weighted geodesic distance on manifolds. Spectral convergence results for the associated Graph Laplacian operators then lead to theoretical guarantees for spectral convergence with these metrics. Building on these insights, I will showcase an application to single-cell RNA-sequencing data, where power-weighted path metrics outperform standard clustering and embedding methods by maintaining both local cluster fidelity and global data geometry. Finally, I will discuss a multi-manifold clustering approach based on path distances in a simplex graph, which ensures robust separation of intersecting manifolds even under noise and curvature. Collectively, these results illustrate the versatility and scalability of path-metric-based techniques in tackling modern challenges in manifold learning, clustering, and data visualization.
The focus of this talk is on acceleration of the analysis stage of the high-fidelity computational fluid dynamics workflow: a severely human-bandwidth-limited step at extreme scales that sits outside of the physics simulation itself. Central to this work is the formalization of ``analysis'' as the construction of natural-language descriptions of relative changes in either large-scale or fine-scale quantities of interest observed throughout many simulations. Based on this formalization, a graph representation of simulations is used to power agent-based automation of this process using a two-step procedure consisting of local neighborhood aggregations and global summarizations. The result is a graph-anchored harness (admitting different types of tools and language model backends) capable of generating verifiable scientific insights from complex datasets directly in natural language through a query refinement procedure. Demonstrations of the framework on high-fidelity computational fluid dynamics data (focusing on reacting flows) will be shown.
Atomic Machines is developing a completely new AI-native digital micro-manufacturing technology stack that will enable new classes of microdevices. Critical to the mission is minimizing “thought-to-thing” time, i.e., the latency between an idea and a manufacturable device. Rather than maximize AI usage, we pursue a generate-to-solve strategy that uses as little AI as necessary, using AI to blend models grounded in physics, constraints, and verification to ensure our designs are trustworthy by construction. In this talk we present an end-to-end workflow that converts an informal description into a precise set of engineering requirements. From this input, an agentic system generates a multi-physics simulation of a first candidate component drawn from a catalog, and closes the loop modifying it with a manufacturability-first generative-design optimization approach. We automatically move from informal descriptions to producing parts that meet performance targets while satisfying process constraints, tolerances, and assembly requirements.
As AI agents move from running instruments to arguing about results, the bottleneck in scientific discovery is shifting from raw capability to trust, provenance, and shared understanding. This talk frames human–AI teaming for discovery around a single organizing idea: a durable, shared memory that scientists and agents both write to and can trust over time. Drawing on a curated, continuously maintained knowledge base of the field's load-bearing work from 2022 to 2026, we trace an arc from "retrieval-augmented generation" toward memory as a governed, persistent system — the substrate on which cumulative science actually compounds.
We organize the problem around three mechanisms that keep collaboration trustworthy: change provenance (who or what contributed a claim, and on what basis), contextual-expertise measurement (who is qualified to weigh in, and how much), and contribution segmentation (trusted versus untrusted, converged versus contested). Shared memory is where these become concrete: temporal knowledge graphs that preserve when a fact was true, multi-user memory with dynamic access control and provenance tiers, lifecycle governance including deliberate forgetting, and benchmarks that test whether teams actually retain and reuse what they learn.
We close with a working example(s) — an LLM-maintained scientific wiki that instantiates these ideas as a lab's own compounding memory — and with open questions for the faculty: how do we credit contributions, retire stale claims, and calibrate reliance when a discovery is co-produced by people and machines? The goal is a practical agenda for AI teammates that remember, and that we can hold accountable.
When training deep neural networks, a model’s generalization error is often observed to follow a power scaling law dependent on the model size and the data size. A theoretical interest in machine learning is to understand why transformer scaling laws emerge. In this talk, we present a mathematical framework for analyzing neural scaling laws of transformers for in-context learning. Our framework exploits low-dimensional structure in both the data and the underlying family of learning tasks, and quantify the benefits of pretraining on a family of tasks with structures and shared information. From this perspective, transformer scaling laws arise naturally from the geometry and complexity of the task space. Our analysis also provides a theoretically grounded interpretation of pretraining: its power lies in efficiently exploring a task space.
Artificial intelligence (AI) has fundamentally transformed the landscape of science, engineering, and technology over the past decade and holds tremendous promise for data-driven discovery. However, AI-driven discovery faces significant challenges arising from the intricate complexity, high dimensionality, nonlinearity, and multiscale nature of biological data. We address these challenges through mathematical AI paradigms. By developing and integrating tools from algebraic topology, differential geometry, geometric topology, and commutative algebra, we have significantly enhanced the ability of AI to tackle complex biological data.
Using our mathematical AI approaches, my team was consistently a top winner in the D3R Grand Challenges, a worldwide annual competition series in computer-aided drug design and discovery. By further integrating mathematical AI with millions of patient-derived viral genomes, we uncovered the mechanisms underlying SARS-CoV-2 evolution and accurately predicted emerging dominant SARS-CoV-2 variants months in advance. I will also discuss applications of mathematical AI to other areas, including single-cell biology, genomics, and biomedical imaging.
This talk investigates the fundamental differences between low-norm and flat solutions of shallow ReLU networks training problems, particularly in high-dimensional settings. We sharply characterize the regularity of the functions learned by neural networks in these two regimes. This enables us to show that global minima with small weight norms exhibit strong generalization guarantees that are dimension-independent. In contrast, local minima that are “flat” can generalize poorly as the input dimension increases. We attribute this gap to a phenomenon we call neural shattering, where neurons specialize to extremely sparse input regions, resulting in activations that are nearly disjoint across data points. This forces the network to rely on large weight magnitudes, leading to poor generalization. Our analysis establishes an exponential separation between flat and low-norm minima. In particular, while flatness does imply some degree of generalization, we show that the corresponding convergence rates necessarily deteriorate exponentially with input dimension. These findings suggest that flatness alone does not fully explain the generalization performance of neural networks.
Foundation models are transforming molecular science from prediction to scientific reasoning and design. This talk reviews recent advances in molecular graph foundation models that enable AI to understand molecular structures, generate new compounds, and accelerate materials discovery. The presentation covers three core applications: virtual screening, inverse molecular design, and retrosynthetic planning, with a particular focus on the challenge of designing molecular structures that satisfy multiple functional objectives. I will introduce recent work on graph diffusion transformers for conditional molecular generation and in-context molecular design, multimodal large language models for joint molecular design and synthesis planning, and graph foundation models for polymer informatics. These techniques have enabled advances in polymer membrane discovery, explainable materials design, and reasoning over polymer structures. Finally, I will discuss the next generation of molecular graph foundation models that integrate graph learning, multimodal reasoning, and scientific knowledge to serve as intelligent assistants for chemistry and materials science, opening new opportunities for autonomous scientific discovery.
Training machine learning models is computationally challenging; one has to solve a high-dimensional, nonconvex optimization problem repeatedly to appropriately calibrate hyperparameters. Our goal is to reduce the training burden by strategic model design paired with structure-exploiting optimization. To this end, we introduce VPBoost (Variable Projection Boosting), a gradient boosting algorithm for separable smooth approximators, i.e., models with a smooth nonlinear featurizer followed by a final linear mapping. VPBoost fuses variable projection, a training paradigm for separable models that enforces optimality of the linear weights, with a second-order weak learning strategy. The combination of second-order boosting, separable models, and variable projection give rise to a natural interpretation of VPBoost as a functional trust-region method. We leverage trust-region theory to prove VPBoost converges under mild regularity conditions. Through numerical experiments on synthetic data, image classification, and scientific machine learning, we demonstrate that VPBoost outperforms gradient-descent-based boosting and attains competitive performance relative to an industry-standard decision tree boosting algorithm.
A long-standing obstacle in the use of machine-learnt (ML) surrogates with partial differential equations (PDEs) is the onset of instabilities when the coupled system is simulated. We present a collection of different approaches to stabilize PDE-ML coupled systems. The first approach treats closure as a multifidelity problem, where a low-fidelity traditional PDE solver makes a prediction which is corrected by an operator network. The coupled system uses in-the-loop training which results in stable rollouts. The second approach couples a low-fidelity numerical solver with a multifidelity correction based on Kolmogorov-Arnold networks and is able to extrapolate accurately well beyond the training interval. The third approach identifies the instability of PDE-ML coupled systems as a result of the spectral bias of ML surrogates. It employs a projection formalism to allow the PDE and ML parts to operate at the same restricted range of frequencies. This leads to stable simulations whose accuracy is further augmented by the introduction of memory terms which account for the interaction with the neglected frequencies. We also offer an efficient way of approximating the memory using ML. Applications to benchmark problems in fluid mechanics are used for illustrative purposes.
For complex engineered systems to operate reliably, it is essential that each constituent component performs as intended and remains within its allowable tolerances. A single component failure may cause complete system failure or, more insidiously, result in inaccurate outputs and unreliable performance. However, validating components within an assembled system can be time-consuming, difficult, and potentially destructive. Non-destructive testing offers an alternative that can assess component behavior while preserving the system for continued operation. This poses a natural inverse problem: can the motion and condition of internal components be inferred from indirect signals measured outside the system? In this talk, we investigate operator learning as a framework for performing non-destructive testing on ratcheting mechanisms enclosed in sealed assemblies. Neural operators are trained to learn mappings from externally observable signals to the internal motion of the mechanism. We compare several operator-learning architectures, including Deep Operator Networks and Fourier Neural Operators, along with two transformer-based approaches. In addition, we examine methods for quantifying predictive uncertainty, including ensemble techniques and conformal prediction, with the goal of identifying when model-based assessments can be considered reliable. Together, these results demonstrate the potential for operator learning to support efficient, uncertainty-aware, non-destructive evaluation of mechanical components.
Scientific machine learning (SciML) provides powerful tools for modeling complex dynamical systems, yet its reliability remains challenged by parameter non-identifiability and the computational difficulty of learning nonlinear operators. This talk presents a unified computational framework toward reliable SciML by addressing these two fundamental challenges. First, we develop practical approaches for parameter identifiability analysis in data-driven models based on the Fisher Information Matrix and its connections to coordinate identifiability. Regularization strategies are further introduced to improve parameter inference, uncertainty quantification, and model robustness in the presence of non-identifiable parameters. Second, we present the Laplacian Eigenfunction-Based Neural Operator (LE-NO), an efficient operator-learning framework for nonlinear reaction–diffusion systems. By exploiting spectral structures of differential operators, LE-NO achieves improved computational efficiency and generalization across varying conditions while reducing dependence on large-scale training data. Finally, we demonstrate the framework in Alzheimer's disease modeling, showing how reliable inference and efficient nonlinear dynamics learning can enable high-fidelity digital twins for complex biological systems.
Learned surrogates and generative models increasingly stand in for numerical solvers across computational science, and as digital twins of physical systems such as particle accelerators, plasmas in tokamaks, weather, and even human videos. In these cases, they are often deployed autoregressively: each prediction feeds the next, small errors compound, and at deployment there is no ground truth to check against. The model cannot tell you how far into the future it can still be trusted. This talk presents a simple structural fix [1]. We train a single conditional latent diffusion model that steps a dynamical system both forward and backward in physical time via a direction flag, so one network is both a surrogate solver and an inverse solver. The method is relatively model-agnostic and can be applied in the same way to a wide class of generative autoregressive models, including flow models. The main idea is that an accurate model composed with its own inverse is the identity, therefore rolling forward i steps and backward i steps must return to the start, and the size of the miss becomes a measurement-free, test-time error signal. Across turbulent magnetohydrodynamics, an astrophysical mixing layer, a public Navier–Stokes benchmark, and natural video, this signal predicts the true rollout error of held-out trajectories to within about 15%, flags out-of-distribution dynamics immediately in exactly the regime where standard sampling-spread uncertainty fails, and lets a single model approach a ten-model ensemble's accuracy at a tenth of the training cost. We show both theoretically and experimentally why bidirectional training beats direction-specialist models in both directions. We also present theoretical bounds on the predictive ability of this approach which shows that the check certifies error only while the learned inverse remains well-conditioned, it dissolves for strongly information-destroying dynamics, and both regimes can be identified offline with representative training data. [1] Preprint: arXiv:2608.00675
Plasma fusion holds great promise as a future source of clean energy. Despite decades of progress in developing reliable computational methods for plasma models, kinetic models—which provide a first-principles description of interacting charged particles—remain prohibitively expensive to solve. Their computational cost is far from meeting the millisecond-scale requirements of real-time burning-plasma control. In this talk, I will describe our recent efforts to integrate deep learning with conventional structure-preserving numerical methods to address some of the key challenges in simulating these kinetic models.
Many computational science and engineering applications can benefit from dimension reduction techniques, as complex quantities of interest often admit accurate representations on low-dimensional manifolds. Motivated by the need to design and analyze workflows that integrate heterogeneous sources of information represented on their respective manifolds, this talk focuses on strategies for linking ensembles of data generated by different models.
Specifically, we investigate measure transport approaches for constructing deterministic maps between model-specific probability distributions through a shared reference distribution. We consider triangular transport maps for this purpose, with particular emphasis on their multi-fidelity construction in settings where high-fidelity model data are limited. Although recent advances in triangular transport algorithms have been substantial, their data requirements remain challenging for complex scientific applications. We show that combining information from multiple sources can provide an effective strategy for mitigating this limitation.
We will discuss two classes of parameterizations for triangular transport maps: hierarchical and non-hierarchical approaches. Their respective advantages and limitations will be examined through a range of numerical test cases, from analytical examples to applications involving amortized inference.
SNL is managed and operated by NTESS under DOE NNSA contract DE-NA0003525.
In this talk, we present a class of optimization-based numerical methods for the simulation of high-speed compressible flows. The proposed framework leverages nonlinear finite element manifolds that adapt the approximation space to the most salient flow features, including boundary layers, shock waves, and rarefaction regions. By embedding problem-specific structure directly into the trial space (e.g., analytical functions or data-driven models), these methods enhance accuracy and efficiency, particularly in regimes characterized by strong gradients and discontinuities. Both the linear coefficients and the nonlinear parameters defining the approximation manifold are determined through the solution of an optimization problem. Numerical results demonstrate the potential of this approach to deliver highly accurate solutions on extremely coarse grids.
This talk presents a series of joint works with my collaborators. Artificial neural networks map inputs to outputs, but when those inputs, outputs, or loss functions depend on global information, network complexity can grow rapidly. Such complexity typically requires large, dense training datasets to achieve meaningful accuracy. In contrast, our recently developed NNLCI method is a local neural network approach: it operates like a scanning microscope that, at each location, examines two coarse-grid numerical solutions—one more accurate than the other—and then predicts the corresponding highfidelity solution at that location. The method has several distinguishing advantages: 1. Small, simple neural network architecture thanks to its strictly local design. 2. High data efficiency: a single finegrid simulation produces hundreds to thousands of local training samples. 3. Sparse training across parameter space: finegrid simulations can be widely spaced, yet the method generalizes strongly due to its locality. 4. High accuracy for both smooth and nonsmooth features: NNLCI resolves discontinuities and complex shock interactions sharply. 5. Geometric flexibility: it works naturally on complex domains, and training and prediction can occur on different geometries. 6. Substantial computational savings: in 2D experiments, NNLCI achieves a twoorderofmagnitude reduction in complexity for shockinteraction problems, and roughly a 500× reduction for smooth solutions such as electromagnetic waves. I will present applications of NNLCI to a broad range of problems, including 1D and 2D Euler equations with shock interactions, unstructured grids, electromagnetic wave scattering from curved conductors with corners, the Poisson–Nernst–Planck ion channel model, 3D Stokes flows, and 3D Black–Scholes equations.
In this talk, we study causal optimal transport in continuous time, with Markovian cost, between a finite-state Markov source and a diffusion target. By replacing the source with its conditional law given the observation of the target, we characterize the value of this transport problem through a fully nonlinear parabolic master equation on an enlarged state space. We further show that this value coincides with those of two equivalent stochastic control problems on the simplex: a control of the Kushner--Stratonovich filtering equation with a zero-mean condition, and a state-constrained stochastic optimal control problem. The proof relies on establishing a comparison principle for suitable classes of subsolutions and supersolutions to the master equation. Both formulations give rise to implementable numerical schemes that approximate the value from above and below.
This is joint work with Julio Backhoff-Veraguas, Ibrahim Ekren, and Antonios Zitridis
Learning the underlying potential energy of stochastic gradient systems from partial and noisy observations is a fundamental problem arising in physics, chemistry, and data-driven modeling. Classical approaches often rely on direct regression of governing equations or velocity fields, which can be sensitive to noise and external perturbations and may fail when observations are incomplete. In this work, we propose a structure-aware, energy-based learning framework for inferring unknown potential functions in generalized diffusion processes, grounded in the energetic variational approach. Starting from the energy–dissipation law associated with the Fokker–Planck equation, we construct loss functions based on the De Giorgi dissipation functional, which consistently couple the free energy and the dissipation mechanism of the system. This formulation avoids explicit enforcement of the governing partial differential equation and preserves the underlying variational structure of the dynamics.
The problem of learning operators from input-output data is fundamental to many areas of scientific computing and system identification. In this work, we consider the specific inverse problem of recovering an unknown kernel function that governs a nonlocal operator, given a dataset of noisy observations. This formulation encompasses a wide range of applications, from constitutive modeling in continuum mechanics to the identification of nonlocal interaction laws in particle systems. Based on this problem, we aim to address two fundamental questions in scientific machine learning: 1) Is the smoother or rougher data more favorable for a robust operator learning? 2) How is data quality, such as its regularity and noise levels, going to impact the learning results?
The primary objective of this work is to rigorously quantify how the roughness of input data influences the learnability of kernels in nonlocal operators. We posit that rough data expands the range of learnable kernels by slowing the spectral decay of the data operator, thereby increasing the effective dimension of the inverse problem. Firstly, a theoretical bridge between the continuous regularity of the input data and the spectral decay rate of the discretized operator is established. We define the effective dimension of the inverse problem via the spectrum of the normal operator, and analyze how it scales with the regularization parameter. It was found that rougher data leads to a slower polynomial decay of eigenvalues, denoted by a smaller decay parameter. Second, we derive explicit convergence rates for the estimation error in the small noise limit. By analyzing the bias-variance trade-off under a source condition characterized by the kernel's smoothness, we demonstrate that the optimal convergence rate depends critically on the interplay between the kernel smoothness and the data roughness. The analysis is numerical verified using both Tikhonov regularized estimators and neural network-based approaches. We confirm that as the data becomes rougher, the effective dimension increases, enabling the accurate recovery of increasingly complex kernel functions and the corresponding nonlocal operators.
Learned PDE solvers are often trained as monolithic surrogates for a specific equation, boundary condition and discretization. This makes them difficult to reuse when mechanisms change and it can limit stability under long-horizon rollout. We introduce Lego-like Operator Network (LegONet), a compositional framework that builds PDE solvers from plug-and-play, structure-preserving operator blocks defined on shared boundary-adapted spectral representations. LegONet separates boundary handling from mechanism learning, satisfying boundary conditions by construction. It also separates mechanism learning from time integration, enabling pretrained blocks to be assembled into new solvers without retraining. We also derive a finite-horizon error decomposition that separates block mismatch from splitting error and provides mechanism-level diagnostics for long-horizon predictions. Across ten time-dependent PDEs, LegONet delivers accurate closed-loop rollouts with improved stability under cross-PDE recombination and boundary reconfiguration. More broadly, this modular formulation suggests a path from task-specific neural solvers towards plug-and-play operator libraries for scientific computing.
Consensus-based methods are derivative-free particle algorithms for non-convex optimization and min–max problems in machine learning. The focus of this talk is a consensus-based algorithm for saddle-point problems, in which two interacting particle populations play the competing roles. Via a coupling argument exploiting the decay and concentration of particle variances, we establish uniform-in-time propagation of chaos with L2 error of order O(1/N1 + 1/N2): finitely many particles remain near a saddle point over arbitrarily long horizons, confirming the method's feasibility. I will also briefly mention a companion uniform-in-time weak propagation-of-chaos result, of order O(1/N), for consensus-based optimization. Joint work with Erhan Bayraktar, Zhiyan Ding, and Hongyi Zhou.
Scientific machine learning offers powerful tools for approximating nonlinear dynamical systems, but its practical impact is often limited by the cost, scarcity, and uneven informativeness of training data. This talk presents two complementary strategies for improving data efficiency in surrogate modeling. First, active learning is formulated as a sequential decision problem in which a reinforcement learning agent selects informative simulations for training neural solution operators. Using a Fourier neural operator surrogate, the framework learns transferable sampling policies that balance uncertainty, diversity, and exploration of extreme system responses across multiple partial differential equation families. Second, a bi-fidelity latent generative framework is introduced for nonlinear hysteretic structural systems, where abundant low-fidelity simulations are combined with limited high-fidelity numerical or experimental data. The model seeks to identify latent variables associated with meaningful nonlinear mechanisms while generating accurate high-fidelity response histories. Together, these studies illustrate a broader scientific machine learning paradigm in which computational resources are allocated strategically, low-cost data are exploited systematically, and learned representations retain connections to underlying physics. The combined perspective highlights how active data acquisition and multi-fidelity generative modeling can expand the applicability of surrogate models to complex, path-dependent, and high-dimensional engineering systems under strict simulation and experimental data budgets.
We will introduce and discuss operator-to-operator learning. In this methodology, which generalizes operator learning, a neural network is used to learn a map between two operators, or from an operator to a function, instead of a map between two function spaces. This can be applied to a variety of inverse problems, such as the inverse Darcy problem and Calderon problem, which naturally take as input an operator on a function space. In-context multi-operator learning can also be put into this framework. We will present two architectures, DeepOSets and FNOSets, for the operator-to-operator learning problem and prove their universality. Finally, we will present numerical experiments demonstrating the efficacy of these methods.
Many fundamental problems on probability measure spaces, including optimal transport, mean field control/games, and Wasserstein gradient flows, are computationally demanding. Existing learning methods often rely on solvers designed for individual problem instances and require costly retraining for each new instance. In context learning with transformers offers a new paradigm for approximating families of operators from only a few context examples, without task specific retraining. In this talk, I will present our recent work on an self-supervised in-context operator learning framework for approximating solution operators on probability measure spaces. The framework is independent of discretization, making it well suited to high dimensional measure transport problems, and requires no supervised solution labels, substantially reducing data generation costs. I will demonstrate its applications to optimal transport, mean field control, Wasserstein gradient flows, swarm control, score matching, and fluid mixing. I will also present a generalization error analysis of the proposed transformer model, connecting our results to the emerging theory of in context learning and highlighting their broader theoretical implications.
Prediction of storm-surge hazard and impacts within planning (pre-disaster), emergency management and post-disaster settings has emerged as a key priority in natural hazard risk mitigation efforts. Migration towards coasts as well as concerns related to the future effects of climate change, further stress the importance of research efforts that attempt to address this priority. Numerical advances in storm surge prediction are one of the most critical such efforts. These advances have produced high-fidelity simulation models that permit a detailed representation of hydrodynamic processes and therefore support high-accuracy surge forecasting. Unfortunately, the computational burden of such numerical models is large, requiring thousands of CPU hours for each simulation, something that limits their applicability for hurricane risk assessment and their broader use in regional planning or emergency response management efforts. This presentation will review how machine learning techniques have been recently promoted to address this challenge, and how the integration of such techniques can be established to better serve the needs of the relevant decision makers (e.g., planners and emergency managers). Emphasis is placed on technical aspects for integrating surrogate modeling techniques to provide surge predictions using a database of high-fidelity storm simulations. This ultimately supports great versatility in leveraging high-fidelity modeling to support regional flood studies (supported by FEMA or Army Corps of Engineers) and real-time emergency response management (supported by NOAA). The discussion then moves to discussing recent advancements for the implementation of graph neural networks as surrogate model in this context, examining its abilities to capture complex spatial dependencies in nearshore domains.
A new strategy is presented for the statistical forecasts of multiscale nonlinear systems involving non-Gaussian probability distributions. The capability of using reduced-order models to capture key statistical features is investigated. A closed stochastic-statistical modeling framework is proposed using a high-order statistical closure enabling accurate prediction of leading-order statistical moments and probability density functions in multiscale complex turbulent systems. A new efficient ensemble forecast algorithm is developed dealing with the nonlinear multiscale coupling mechanism as a characteristic feature in high-dimensional turbulent systems. To address challenges associated with closely coupled spatio- temporal scales in turbulent states and expensive large ensemble simulation for high-dimensional complex systems, we introduce efficient computational strategies using the random batch method. Effective nonlinear ensemble filters are developed based on the nonlinear coupling structures of the explicit stochastic and statistical equations, which satisfy an infinite-dimensional Kalman-Bucy filter with conditional Gaussian dynamics. It is demonstrated that crucial principal statistical quantities in the most important large scales can be captured efficiently with accuracy using the new reduced-order model in various dynamical regimes of the flow field with distinct statistical structures.
The poster session takes place on Saturday, September 26, 6:00 – 7:30 pm in Jordan Hall, together with the speakers' banquet. 35 posters will be presented, listed below in alphabetical order by surname.