
Improving batteries, designing drugs and understanding the molecular basis of living systems — all require grappling with chemical complexity. Computational chemist Samuel Blau has built his career at Lawrence Berkeley National Laboratory around exactly that, developing the tools and datasets to study processes too intricate to measure experimentally.

“Molecular simulations can help us understand the microscopic processes governing these systems,” Blau says, “and eventually allow us to control and even design complex reaction cascades.”
That vision has driven his most ambitious project to date. In May 2025, Blau and collaborators from academia, national labs, and Meta and other commercial partners released Open Molecules 2025 (OMol25) — a trove of more than 100 million 3D snapshots of molecules and their quantum chemical properties. OMol25 is the largest and most chemically diverse molecular dataset for training machine learning models, Blau says. The achievement required more than three supercomputer years of compute time, using Meta’s resources.
OMol25 addresses a significant bottleneck in chemical simulations: density functional theory. DFT uses quantum mechanical equations to model how electrons arrange around atomic nuclei, revealing the precise details of atomic interactions that govern molecular behavior. But DFT’s computational cost scales as the cube of system size. As the number of electrons doubles, a calculation takes eight times longer. That scaling wall limits both the size and the number of chemical systems researchers can simulate.
Machine learning interatomic potentials, or MLIPs, offer a way around this wall. Trained on DFT data, he says, “MLIPs allow you to run molecular dynamics, optimize geometries or calculate free energies dramatically faster than DFT” — 10,000 times faster — “and their cost scales linearly, not cubically, with system size. That’s all you need to build a reaction network to pick apart a complex reactive system.”“Molecular simulations can help us understand the microscopic processes governing these systems,” Blau says, “and eventually allow us to control and even design complex reaction cascades.”
But like all AI models, an MLIP is only as useful as the data behind it. OMol25 fills that data gap. In assembling the project’s team, Blau drew on his alumni network from the Department of Energy Computational Science Graduate Fellowship (DOE CSGF), which supported him from 2012 to 2016. The OMol25 team included three other former fellows: Meta’s Zachary Ulissi and Berkeley Lab’s Santiago Vargas and Aditi Krishnapriyan, who is also an assistant professor at the University of California, Berkeley.
Patience is a thread running through Blau’s career.
“We didn’t just try and get a quick win,” Blau says. “By combining a world-leading team with Meta’s compute resources, we patiently created a transformative resource.”
That patience is a thread running through Blau’s career. “The fellowship gave me freedom in grad school to work on whatever I wanted because I wasn’t tied to a specific grant,” Blau says. “That led me to target a project which required data-driven approaches, which required patience and doing something carefully.”
The fellowship, he says, instilled a commitment to rigor and a quality-over-quantity approach to research. Whereas peers raced to publish on multiple projects, Blau spent years using quantum calculations to disprove a long-held assumption: that quantum coherence between light-absorbing molecules in photosynthetic systems was essential to their efficient energy transfer. The single first-author paper he published from that work shifted the field’s direction.
The same patient, data-driven mindset fueled his postdoctoral research in Kristin Persson’s group at Berkeley Lab. There, Blau developed software to automatically detect and fix failed DFT calculations — a critical capability when running thousands of simulations, roughly 20% of which fail on the first attempt. That infrastructure enabled him to simulate thousands of battery reactions simultaneously, shining light on a decades-old chemical mystery.
Blau’s fascination with complexity started early. At 15, he joined Berkeley Lab as a summer intern, working on thin-film devices for smart windows. Since then, he has gravitated toward problems at the edge of what computation can handle — among them, batteries, microchip patterning and catalysis.
“We’re seeing a ton of community engagement with both the models and the data,” Blau says. In its first six months, their OMol25 paper was cited nearly 100 times. Scientists at Pfizer used the data to train a biomolecular AI model. Microsoft trained MLIPs for simulating polymers.
And the MLIPs that Meta trained on OMol25 proved so accurate across such diverse chemistry that Markus Reiher, a prominent computational chemist at ETH Zurich, declared at an international chemistry conference in September 2025 that DFT would soon be obsolete.
Blau thinks both approaches are needed for now. “I think DFT will still be important for certain types of simulations and generating training data for some time. But undeniably, OMol25-trained MLIPs have already fundamentally changed the questions we can ask, and we’re just getting started.”
This article was adapted from the 2026 print edition of DEIXIS: The DOE CSGF Annual.

