Case study
EERIE
Leading the ocean–atmosphere evaluation of the EU's kilometre-scale coupled climate models.
- Coupled-model evaluation
- AI-emulator benchmarking
- Petabyte-scale engineering
- Air–sea coupling
- EERIE evaluation partner centresMPI-M · ECMWF · Met Office · BSC
- AI weather emulators benchmarked6
- EXCLAIM km-scale output processed~200 TB per simulated year
The role
This is the day job. At ETH Zürich and C2SM I lead the ocean–atmosphere evaluation for EERIE, the EU Horizon programme building the next generation of kilometre-scale coupled climate models, and as a Research Fellow of the ETH AI Centre I benchmark the AI weather emulators moving into the same territory. Three strands, one question: whether these models can be trusted where the physics is hardest.
Evaluating a new generation of climate models
EERIE runs coupled climate models at kilometre scale, fine enough to resolve the ocean mesoscale eddies and the air–sea feedbacks that coarser models can only parameterise. My role is to judge whether the emergent air–sea physics in those runs is right, across the models produced by MPI-M · ECMWF · Met Office · BSC. The result is the first multi-model quantification of mesoscale air–sea feedbacks in this class of model, written up as 2 first-author papers submitted.
Evaluation at this resolution is not a matter of eyeballing maps. To confirm that a model reproduces a physical mechanism rather than a lookalike statistic, I built a spatially resolved causal machine-learning framework that tests for the mechanism directly, presented at Climate Informatics 2026.
Benchmarking the AI emulators
The AI weather emulators are arriving faster than the tools to judge them. As a Research Fellow of the ETH AI Centre I engineered a multi-GPU inference and benchmarking pipeline spanning 6 operational emulators (AIFS, GraphCast, GenCast, NeuralGCM, Pangu-Weather and FourCastNet 3), and used it to quantify a systematic weakness: regression-trained emulators suppress the extremes, smoothing away exactly the tails that matter most for risk. The extreme-value and decision-value analysis behind that finding is its own case study, tailspec.
> ENGINE_ROOM Petabyte-scale evaluation infrastructure (EXCLAIM)
None of the evaluation happens without the engineering to move the data. Under EXCLAIM I built GPU-accelerated, Dask-distributed analysis packages that process ~200 TB per simulated year of kilometre-scale model output, so a diagnostic that would otherwise be intractable runs over the full archive rather than a subsample.
The same work became teaching. I designed and taught the C2SM workshop on distributed task-parallel computing with Dask and xarray, so the group can run the pipelines itself rather than depend on one person to feed them.
In collaboration with Oxford
With a group at the University of Oxford, I am an invited co-author on a paper in preparation that uses machine-learning weather emulators as diagnostic instruments for the ocean’s role in fast atmospheric evolution. My contribution is on the air–sea coupling side: the coupling theory and the multi-model feedback benchmarks that connect the emulator diagnostics back to the coupled-model physics.
What this demonstrates
The evaluation half of building trustworthy weather and climate AI: leading a cross-institutional assessment of kilometre-scale coupled models, benchmarking the operational emulators for the failure modes that matter, and engineering the petabyte-scale infrastructure that makes both tractable. The through-line is judging where a model can be trusted and where it cannot, which is the prerequisite for putting any of these models to work.