Since coding agents became much better at the end of last year, there has been a huge debate about what comes next. If LLMs can transform coding so quickly, why not biology, drug discovery, or healthcare?
Anthropic is now putting a lot of attention into these areas. Dario Amodei has gone much further, arguing that powerful AI could help cure most diseases within five to ten years.
I posted skeptically about this on social media, and I don’t remember ever receiving so many supportive DMs. Many came from people working in biology, running labs, or developing drugs.
I should say clearly that I’m not an expert in drug development. I have a bachelor’s degree in computational biology and a PhD in computational neuroscience, but I don’t currently work directly in biology. For this post, I spoke with several people who run biological labs or work at the intersection of AI and biology.
Of course, nobody really knows what AI systems will be capable of five years from now. But based on what we have today, I think there is a basic problem with the argument that progress in coding will transfer to biology.
The verifier problem
One useful way to think about AI tasks is to ask how easily an answer can be verified.
The fields where scaling and reinforcement learning have produced the most dramatic progress usually have cheap, fast feedback. Code has compilers and unit tests. Formal mathematics has proof checkers. Games have scores and clearly defined rules.
An AI can generate an answer, test it, learn from the result, and repeat this loop millions of times. This is extremely important. It allows the system to search for answers without a human having to check every attempt.
This is why I don’t find the argument “AI can solve hard mathematical problems, so drugs are next” very convincing. The mathematical result may be impressive, but it also shows what happens when you can generate a candidate and check it almost for free.
What happens when checking an answer takes three months? What if it costs hundreds of thousands of dollars? What if the experiment is noisy, and even after it ends you still don’t know whether the result will generalize to humans?
That is much closer to biology.
A drug has to survive several different problems.
A simplified drug-development process has at least four parts: mechanism, molecule, development, and market. A program must succeed at all four. Failure in any one of them can kill the entire project.
Mechanism is the question of which biological process to intervene in.
We are worse at this than people think. High HDL cholesterol is associated with lower cardiovascular risk, yet several drugs that successfully raised HDL failed to improve clinical outcomes. The correlation was real. The intervention was wrong.
Vioxx worked as an anti-inflammatory, but it also increased cardiovascular risk. A drug can do exactly what it was designed to do and still cause serious harm somewhere else in the body. You cannot necessarily discover that by reading more papers or reasoning more carefully.
Molecule is the question of what physical object can implement the intervention.
Liraglutide and semaglutide target the same receptor and are structurally similar, yet one has a half-life of roughly 13 hours while the other lasts about a week. The difference is not simply whether they bind to the receptor. Small molecular modifications change how they bind to albumin, resist degradation, and survive in circulation.
This is why success in binder design should not be confused with success in drug design. Binding matters, but a real drug must also reach the correct tissue, remain stable, avoid unwanted interactions, be manufacturable, and behave safely inside a complete organism.
Development is where most of the time disappears.
Animal experiments take biological time. Some disease models require weeks or months before there is anything to measure. Then the experiment often has to be repeated to make sure the result was not noise.
Human trials are even harder. People must be recruited. Safety must be monitored. Some diseases progress slowly, so researchers need to wait long enough to see whether the treatment changes an actual outcome. According to the FDA, Phase II trials can take several months to two years, while Phase III trials often take one to four years.
You cannot solve all of this by buying more GPUs.
The market also matters, even if people prefer not to discuss it. A known mechanism may be less risky but lead to a crowded market. A novel mechanism may have much greater value, but it also carries a much higher risk of failure. Some diseases have too few patients or too little purchasing power to attract enough investment.
There is no version of this process where everything becomes easy at the same time.
Biology has verifiers, but they are bad ones.
It would be too strong to say that biology has no verifiers. It has experiments.
The problem is that the important experiments are often slow, expensive, noisy, and incomplete.
Near the beginning of the pipeline, researchers can test whether a molecule binds to a target. But after that, the questions become much more difficult. Does it survive in the body? Does it reach the right tissue? Is it toxic? Does it work in an animal? Does the animal model actually predict what will happen in a person?
Eventually, you cannot test whether a compound is safe in humans without giving it to humans. You cannot identify rare side effects without studying enough people for long enough. You cannot determine whether a treatment delays a slowly progressing disease by running a faster simulation.
A clinical trial is not simply a very slow unit test. It is a noisy experiment conducted on a heterogeneous population, subject to ethical and operational constraints. In many cases, you get only a few serious attempts before the money or patent window runs out.
What Anthropic actually showed
None of this means that Anthropic’s recent result is unimportant.
Anthropic reported that Claude managed de novo protein-binder design campaigns and produced successful binders for 14 of 15 evaluated targets. External laboratories then produced and tested the designs. This is a genuinely impressive result. Anthropic also released the prompts, designs, and experimental data.
The company’s research post is actually careful. It says that minibinders are not therapeutics and that a high-affinity binder is only the initial step toward a drug-like molecule.
But this caveat is not a small detail. It is the main point.
Binder design is one of the biological problems most compatible with current AI. You can generate many candidates computationally, rank them, and then test them with a relatively standardized experiment. It looks more like the problems where AI has already succeeded.
That is probably why we are seeing progress there first.
The mistake is assuming that because AI compressed this step, it will compress every later step at the same rate.
The scientific result is real. The questionable part is how the result changes as it moves upward: from designing protein binders to accelerating drug discovery to compressing decades of biological progress to curing most diseases within ten years.
At every step, a caveat disappears.
The economics are different too.
There is also a mismatch between the economics of frontier AI and those of drug development.
Software can be deployed globally, measured immediately, and updated continuously. Drug development requires long capital cycles and often results in many binary failures. A molecule can take years and cost hundreds of millions of dollars before a clinical trial shows it does not work.
This does not mean AI companies should avoid biology. But once a company begins to own and develop drugs, the economics start to look much more like pharma than software.
The promise to cure disease also does more than describe a research program. It provides a moral justification for enormous investments in compute, energy, and AI infrastructure. It makes the concentration of money and power around a few AI companies sound not only profitable yet necessary for humanity.
That doesn’t invalidate the scientific work. It does mean we should demand stronger evidence before accepting the larger story.
What will happen?
I use AI every day, and I am thrilled about projects that use it to search molecular space, design proteins, interpret experiments, and generate hypotheses.
AI will improve biological research. It will likely yield more drug candidates, and some may be much better than those we can design today.
But producing a prospective candidate starts the process. It does not finish it.
My view is that AI will substantially accelerate some parts of drug discovery without proportionally shortening the entire path from an idea to a safe, effective, widely available treatment.
The longest parts of that process are not simply periods when scientists sit around thinking too slowly. Researchers have to manufacture molecules, perturb living systems, recruit patients, observe outcomes, and wait for biology to reveal what happens.
This may change. Better simulations, robotic labs, stronger models, and more predictive animal or cellular systems could shorten some of these stages. AI may eventually become so accurate that it can select the rare molecule that satisfies efficacy, safety, delivery, stability, and manufacturing requirements on the first attempt.
But we have not seen evidence for that yet. And even doing it once would be very different from curing most diseases.
Anthropic’s binder result supports a narrower conclusion. It is less dramatic than curing most diseases, but still important:
AI is becoming extremely good at accelerating the parts of biology that resemble computation. Unfortunately, much of medicine is slow precisely because the decisive parts do not move quickly.


