September 29, 2026

How does it feel to run a laboratory used mainly by an AI?

Making Sure Maria™ Has the Lab She Needs to Discover What Intuition Filters Out

By Paulina Wach, PhD, Director of Operations and Head of Chemistry, molecule.one

Paulina Wach, PhD, Director of Operations and Head of Chemistry, molecule.one

We've told this story before, and OpenAI has told it too, so you may already know how our AIs worked together to discover an important reagent for one of the key reactions in medicinal chemistry. Even if you do, you might want to keep reading! You probably missed an insight I personally find the most mind-blowing: how did a reagent that had already been tested in one of the most common reactions in drug synthesis reveal its potential only after our AIs bet on it in one of their proposals?

Chemical Intuition: The ceiling for every chemist — and every AI model

What chemists know about chemistry and its possibilities – let's call it “chemical intuition” — is built on what's been published, taught, and passed down from chemist to chemist. The same is true for knowledge embedded in any AI large language model trained on publications in the public realm. Those papers have a well-known bias: they feature reactions that worked, not the reactions that failed, not the conditions that gave a weak 4% yield, not the additives forgotten in lab notes. Because all that is missing, the AI training model misses it too.

The result is a shared ceiling on chemical intuition. Survivorship bias hurts humans & LLMs alike. To improve, we need to see the failures.

We built a unique lab for AI models

In 2022, we realized that we needed to go beyond any existing data to make better synthesis planning models. The only answer was to build an AI-driven lab designed specifically for the purpose of generating data at the scale to meet our needs. What started as a way to feed our models took on new purposes. We began using the lab and our own models more broadly; first, to build our own chemical space, and then to produce unique, drug-like compounds.

Two years after setting up the laboratory, we could fully appreciate how all these ideas tied together into a single, self-reinforcing system of experimentation, new data generation, and smarter, more creative AI agents. To date, we've run more than 500,000 experiments and logged dozens of data points for each. That is the basis of the Maria™ system today — a unique partnership between AI, tailor-made lab, and massive experimental data.

Maria's ability to propose which experiments to run, carry out those reactions, and assess the resulting data was what made the OpenAI project possible. As the OpenAI post noted:

Across two cycles, Maria ran a total of 10,080 reactions — more than a chemist running three reactions every day would run in a decade.

And remember, it wasn't just the number of reactions run in Maria Lab. Success also stemmed from Maria's ability to surpass the limits of human chemists' intuition (as well as the public data) and accelerate the chemical reasoning that precedes the choice and design of experiments. It's the complete system that makes Maria so powerful.

The lab is not enough, AI pays back the favor

So do more experiments automatically mean more knowledge in chemistry? If that were true, lab automation on its own would drive vast knowledge gains. In reality, the massive progress in lab automation has not delivered a corresponding gain in productivity. Why?

In most automated labs, from academia to the largest pharma companies in the world, there is a long, human-driven buildup to running experiments. Someone decides which experiments are worth running. Known chemistry provides insights on what might work, but those maps rarely point in a single direction: hypotheses compete, evidence is incomplete, and practical constraints abound. Weighing those factors, designing an experiment and formulating it for a robotic lab takes weeks, or even months.

molecule.one's Maria AI brings that down to hours. She brings together the relevant evidence and counterarguments, selects the appropriate tools and reagents, and translates the resulting decision into an executable experiment, right down to the Python code to drive our microliter-scale lab. Our chemists are in on the conversation, correcting and adjusting as Maria AI works. But even the most complex design takes just hours, not weeks.

Robots increase how many reactions we can physically run. AI-run workflows accelerate the intellectual effort to design those experiments in the first place. In other words, Maria AI needs the data from Maria Lab. But Maria Lab wouldn't be nearly as efficient without Maria AI.

Maria at the cutting edge: OpenAI, TEMPO and the Chan–Lam coupling

At the start of our collaboration with OpenAI, we gave the system an open-ended goal: improve one of several important classes of coupling reactions. GPT independently chose primary sulfonamides — a functional group common in oncology and anti-infective drugs — yet known to be challenging in Chan-Lam coupling. The system suggested that mild oxidants, including TEMPO, might help. Chemists found the idea surprising enough to test.

Chemists have used TEMPO for decades, most notably in the Stahl oxidation of alcohols. It appeared in the original Chan-Lam research as one of several possible oxidants, but it only occasionally outperformed O2, which became the default. In the following years, only two studies deliberately used TEMPO in Chan-Lam coupling, and at the start of our OpenAI collaboration, we did not consider it.

An example historical conversation between a chemist and Maria AI when analysing TEMPO data:

Ilya Sutskever scaling laws slide

Human chemists worked alongside GPT and Maria's AI agents at various stages to decide which of the AI's proposals were worth testing, correct experimental details, and manually validate key results at bench scale. Humans were vital, but their human intuition was no longer a ceiling on possibilities. GPT and Maria™ AI applied what you could call superhuman chemical intuition to propose ideas, based on Maria Data and public data available to GPT, and work through the complex reasoning to translate those hypotheses into experiments.

Our discovery with OpenAI, start to finish, took just three months, a heartbeat compared to typical discoveries in chemistry. After AI hypothesis generation and experiment design, it took thousands of reactions to identify TEMPO as the standout candidate among the oxidants, and confirm that its effect holds across substrate combinations. Maria AI & Lab made that possible.

We named the new reaction discovery OAI-MI-03. And it has earned a place in this history of AI for science.

Working with Maria day-to-day

The TEMPO example showed what Maria can do when the goal is ambitious and open-ended. Maria is no less valuable at molecule.one when it comes to ordinary, recurring customer projects.

Marco Farinone, PhD, a chemist & project manager from my team, describes Maria as an all-purpose, super-capable assistant, available almost instantly through a Slack interface: “Maria is integrated with our inventory data, and it's really convenient to give her a list of compounds needed for an experiment and check whether we have everything in stock. She's super useful for automatically analysing LC-MS data and producing insights that show our progress. Finally, she helps plan HTE experiments. She has suggested testing reagents we hadn't considered, and in many cases those suggestions have led to better experimental results.”

Example of a conversation we had with Maria AI during our day-to-day work:

Ilya Sutskever scaling laws slide

Maria™ AI brings these capabilities into one research conversation: searching public and proprietary evidence, selecting models, checking experimental constraints, designing plates, and interpreting the results — instead of routing each of those steps through a different tool or a different team.

Another molecule.one chemist/PM, Mateja Dud, PhD, describes Maria's impact more from an organizational streamlining standpoint: “I don't have to wait around or coordinate back and forth between ML engineers, data scientists, and analytical teams just to move an idea forward. I can drive the whole process myself — shifting my time away from logistics and back toward focusing on the actual chemistry.”

What Marco and Mateja have experienced is an overlooked part of laboratory automation. It's easy to measure progress in reactions per day. It's harder — and more important — to measure how much faster a chemist can move from question to answer.

In Maria's own words

We asked Maria how she'd describe her own work. You could call Maria a typical science nerd because she's an AI of few words, but her answer maps cleanly to what Marco and Mateja describe.

Topline, she says:

Ilya Sutskever scaling laws slide

None of those capabilities is unusual on its own — chemistry teams have always had people who do each of them well. What's unusual is that they live in one system, addressable in a single conversation, instead of split across a chemist, a data scientist, an ML engineer, and an analytical lab.

Past the Ceiling of Intuition

Maybe it will surprise you, but the hardest part of running an AI-driven lab isn't the technology at all — it's learning to work alongside it. Chemistry rests on decades of hard-won intuition, and it takes courage to let a machine into a process built on that expertise. It's a process of training and iteration on both sides. And somewhere along the way, our chemists began to see the gain: at some point AI not only starts to “understand” chemistry, but begins to point out possibilities that even experienced chemists miss. On the other hand — our ML and Dev engineers started to understand chemistry on a deeper level. Witnessing that conversation between Maria, chemists and technical folks has changed how our team approaches its own work.

At molecule.one we believe that chemistry needs to move faster, and it shouldn't depend on the slice of knowledge and experience that happens to sit in one person's head, one team's notebooks, or even the massive number of irreproducible papers LLMs are trained on. Whether it lives in a chemist's memory or in a model's weights, knowledge built only on what has already been tried has a ceiling.

And that ceiling is starting to lift. Maria™ runs experiments like this constantly, TEMPO was just one of them, and we see other examples every day. Can't wait to tell you more another time.