AI badly needs a dose of skepticism
Some scientists are too eager to believe their own claims
First let me get this part out of the way: AI has made some amazing recent advances. The most obvious one is the ability of large language models (LLMs) like ChatGPT and Claude to engage in conversations about a wide range of topics, and to offer accurate summaries of human knowledge. They also provide a very fast, very smart way to search the internet.
The latest AI systems are also remarkably good at writing code. About a week ago, I used ChatGPT to write a short Perl program, and in a few seconds it spit out code that would have taken me at least half an hour to write. Nothing fancy, but very useful, and pretty impressive.
But the claims being made for AI lately are far, far grander. Case in point: the creators of a number of deep learning models of DNA claim that their models can predict the effects of virtually any mutation in the human genome. The details vary, but similar claims have been published by all the leading AI companies, including Google DeepMind, the Arc Institute, InstaDeep, Anthropic, and many university AI groups including MIT, Stanford, Columbia, and others.
The top scientific journals have been all too eager to go along, especially Nature, which has published some of the most extravagant claims.
I don’t want to get into the technical claims made in the recent papers and press releases – that would take far too long – but a simple summary is this: by feeding huge quantities of DNA sequence to enormous “deep” neural networks that contain billions of adjustable weights, these scientists claim that their models can predict how genes behave in a virtually unlimited range of scenarios. (They call their models “foundation models", so I’ll use that term here, although it is more or less interchangeable with the term large language models, or LLMs.)
One common claim for foundation models is that they can predict the effects of mutations that have never been seen. This would be amazing if true, but there are (at least) two large, fundamental problems here:
The notion that you can use DNA sequence alone to predict how genes will behave is biologically implausible. As every biologist knows, genes’ behaviors change based on context. For example, heart cells do not behave at all like neurons, and neither heart cells nor brain cells behave like skin cells. And yet the DNA in all of these cell types is identical–so, obviously, you need more than the DNA to predict how genes behave.
These claims are largely unfalsifiable. The human genome contains 3 billion base pairs (letters in the DNA alphabet of A, C, G, and T). That means billions of mutations are possible in anyone’s genome, most of which have never been observed. While we might test the effects of a few mutations with a great deal of effort (and we might disprove them in the process), no one is going to do that.
After 15+ years of fighting pseudoscience in my blogs and other writing. I’m very familiar with these problems. Consider acupuncture and homeopathy, two forms of pseudoscience that are widely believed and whose practitioners are deeply invested in convincing people that they work.
Both acupuncture and homeopathy are scientifically implausible, to put it kindly. (I’ve written about them both many times, such as this column from 2016 and this one from 2024.) Homeopathy is perhaps the more wildly misguided practice, but both depend on pre-scientific claims that were never proven, and that contradict a vast amount of modern scientific knowledge. These aren’t close calls.
Proponents also love to make unfalsifiable claims, such as the claim that acupuncture works because it interrupts a vital force flowing along invisible lines of energy throughout the body, called meridians. Yes, that’s one of its central claims! Meridians have never been shown to exist; they’re just fiction.
Now, when it comes to AI claims about the human genome, their implausibility is certainly less obvious than the ridiculous claims of homeopathy or acupuncture, and perhaps that is one of the problems. The AI scientists building these models are experts on machine learning and AI, but they are much less familiar (as evidenced by their own publications) with molecular biology and genetics. So it is likely that they believe their own claims, and they might not understand how implausible they are.
But they ought to understand that it would be nearly impossible to disprove their claims. They claim to predict a vast array of genetic phenomena that could only be tested, in most cases, with slow and expensive experiments. Why would anyone invest the time required to run those experiments? The short answer is that no one will.
Meanwhile the AI train keeps rolling, and Nature keeps publishing more papers with foundation models that claim to understand DNA.
Going about science wrong
A third fundamental problem with the recent claims from AI groups about biomedicine is different: these groups are simply taking a bad approach to science.
What’s happening is that teams of AI scientists are building ever-larger models (and boasting about how many billions of parameters they have), and then claiming that each model can solve even more than the previous one.
Then, model in hand, they go around looking around for problems to solve. Some of them have realized that DNA sequencing is super-efficient and is generating reams of data, so they’ve latched onto this and said “Aha! look at all that data, we can use it to understand life itself.”
This is not how good science gets done. You don’t start with a tool and then look for problems.
As all scientists should know, the first step in doing science is to identify a problem to solve. Then, once you understand something about the problem, you can begin to think about how to solve it, and design experiments, or create new technology, or do whatever else is needed.
Deep learning scientists are doing this backwards: they have a solution in search of a problem. And when you have a beautiful new hammer, suddenly a lot of things look like nails.
Now, going about science backwards doesn’t necessarily mean you will fail, even though it’s a bad idea. A bigger problem that most of the deep learning enthusiasts suffer from is their strong bias in favor of believing that their tools (LLMs or foundation models) are going to work. This colors all their experiments.
Again, not how to do good science. If scientists already believe that AI can predict the effects of every genetic mutation once it has seen enough DNA, then of course they’ll believe its predictions. What they should ask first (as I’ve always taught my own students), before announcing their claims to the world, is: how can we be sure this isn’t wrong? Did we miss something? Is there a less exciting explanation of the data?
(Aside: Buried in the technical details of the recent papers on AI foundation models, one learns that none of their predictions have been independently verified. It might turn out that the models don’t work well at all on new data sets. But that’s not how they’re being sold.)
And here’s another thing about LLMs: AI scientists themselves have no way of “looking under the hood” to tell you what the models have learned. So the only thing they can do is point to performance on some kind of test data and say it’s impressive. None of the papers describing DNA models offer any insights into human biology or genetics.
So I’m still waiting to see a single demonstration that an AI model of DNA can do anything beyond spit out something that was already known.
But wait, there’s more! Just last week, Nature published a paper describing the “AI Scientist,” a deep learning model that does its own original research in … (wait for it) … machine learning! That’s right: according to the paper,
“The AI Scientist creates research ideas, writes code, runs experiments, plots and analyses data, writes the entire scientific manuscript, and performs its own peer review.”
Wow, it seems we don’t need scientists at all any more! I read the paper (reluctantly) and found that their “ultimate” test was to use the AI Scientist to write 3 papers and submit them to a scientific conference on machine learning. They chose a low-quality workshop where 70% of the submitted papers get accepted. Of the 3 papers, 2 were rejected, which the authors reported as a huge success.
Look, everyone knows that workshops get at most glancing peer review. If you can string coherent sentences together–and LLMs are quite good at this–then you can probably get your paper into a workshop with a 70% acceptance rate. Getting 1/3 of your AI-generated papers accepted is not at all impressive.
And yet the authors of The AI Scientist called their work “a milestone in the centuries-long scientific endeavour.” This statement appeared in the paper itself, not in a press release. I am astonished that the editors at Nature failed to remove such a boastful, overstated claim. I guess we’ve come a long way from the 1953 Watson and Crick paper on the discovery of the DNA double helix, which included perhaps the most famous understatement ever: “It has not escaped our notice that the specific pairing we have postulated immediately suggests a possible copying mechanism for the genetic material.”
Perhaps the authors of all of these recent deep learning papers should read that one.



Bingo! Remember when Watson beat Ken Jennings at Jeopardy! and was suddenly going to cure cancer. I am still waiting. We're at the same point in the hype cycle with LLMs/Foundation Models, but there are many, many of those. AlphaFold was very good at predicting protein structures, but a protein is a one-dimensional chain of amino acids. The link between genotype and phenotype (our genetic variants and our traits) is highly nonlinear, multidimensional, and multifactorial. AI models can learn rules (or pretend to), but creating new insights, and testing them, is a far more challenging than using a probabilistic model to string together some words or concepts.
All the glory goes to the predictors, so little goes to the validators!
I work mostly in non-model organisms, and I find it a real challenge when these splashy announcements come out and I get asked when I can integrate them into projects... I want to leverage all the fancy new tools and data available for model organisms, but the amount of validation data is so scarce, I don't really know how I would evaluate if some tool is actually working.
Reminded me of Rachel Thomas' article about the Nature Comms ML Gene Annotation paper which had serious problems with the predictions, and an expert was available to fact-check one that she knew was wrong. What happens as we get fewer deep experts and more generalists?https://rachel.fast.ai/posts/2025-06-04-enzyme-ml-fails/