by datastudy.nl

Monday, September 28, 2026

AI

AI scientific discovery claims face biologist backlash

AI scientific discovery is under fire after Anthropic said Claude found a CRISPR-like DNA pattern that biologists call routine grunt work, not a breakthrough.

Funnel chart showing Claude's biology pipeline narrowing from 200,000 reverse transcriptases to 3,500 candidates, 20 analyzed, and 1 pattern flagged. AI scientific discovery claims under scrutiny.
Claude's enzyme discovery pipeline: 200,000+ reverse transcriptases gathered, 3,500 new candidate systems, 20 compelling candidates, 1 pattern flagged as ART. Source: Anthropic.

Anthropic confirmed earlier this year that it was operating a wet biology lab in the Bay Area where Claude agents sift through DNA databases and human scientists run the physical experiments. Last week the company announced the lab's first result: 950 agents spent 21 hours burning through 210 million tokens and flagged a repeating DNA pattern near a known enzyme. Anthropic called the pattern "reminiscent" of CRISPR. Biologists called it grunt work.

The gap between those two characterizations is the whole story. An AI scientific discovery claim crashed into the reality of how biology actually works, and the collision exposed a fault line that matters for anyone building AI tools for research domains. If you are building agents that search, filter, or synthesize scientific data, the Anthropic episode is a case study in how to oversell a legitimate technical achievement and burn credibility with the domain experts you need most.

What did Claude actually find in the DNA data?

Anthropic gave Claude agents a prompt to search a massive database of DNA sequences for interesting new examples of reverse transcriptases, enzymes that copy RNA into DNA. The agents worked autonomously after the initial prompt. They gathered over 200,000 reverse transcriptases, identified 3,500 new candidate systems, and narrowed those to the 20 most compelling, producing human-readable reports for each. The company says that analysis would take an expert scientist weeks to months. The chart below shows how drastically the pipeline narrowed.

Funnel chart of Claude's biology discovery pipeline: 200,000 reverse transcriptases gathered, narrowed to 3,500 new candidate systems, then 20 compelling candidates analyzed, then 1 pattern flagged (ART). Source: Anthropic.
Claude's enzyme discovery pipeline narrowed 200,000+ reverse transcriptases to 3,500 candidates, 20 compelling candidates, and 1 pattern flagged as ART. Source: Anthropic.

One of those 20 candidates caught the agents' attention: a repeating pattern of DNA sequences sitting next to the gene for an unusual reverse transcriptase found in a jumbo phage, a virus that infects bacteria. Anthropic named the system array-associated reverse transcriptases, or ART. The underlying reverse transcriptase was already known. What Claude appears to have noticed first is the associated array of non-coding DNA sequences and an accessory protein of unknown function. The repeat layout resembles a CRISPR array, which is what makes CRISPR-Cas systems programmable tools.

The announcement leans hard on the CRISPR comparison. The company wrote that the system has characteristics "only ever been found together in a handful of other systems, all of which are programmable and perform operations like cutting, copying, and pasting DNA." The blog post calls the pattern "reminiscent" of the discovery that led to CRISPR, which "has already transformed science and medicine."

The findings were published on Anthropic's blog with an accompanying technical paper but have not appeared in a peer-reviewed journal. Eric Kauderer-Abrams, Anthropic's head of life sciences, acknowledged in an email that the company's researchers "don't yet know" the significance of the finding. He said the company likes to publicize early research findings so the public and other companies can gauge Claude's capabilities.

Why are biologists calling the claim an overclaim?

The pushback was swift and came from people whose opinions carry weight. Lucas Harrington, a biologist and biotech entrepreneur who co-founded multiple gene editing companies, wrote a viral post on X. The CEO of Eli Lilly endorsed it. Harrington's point was blunt: "finding a weird cluster of genes and repeats is often the easy part. The hard part, and where the real discoveries come from, is figuring out what the system actually does."

Steven Salzberg, who heads the Center for Computational Biology at Johns Hopkins University, was sharper. He told Bloomberg that plenty of existing computer tools already find CRISPR-like enzymes, and that other such enzymes have been discovered over the years. "What most biologists don't do is announce an unproven, preliminary finding as if it's the next CRISPR, suggesting that they have something revolutionary," Salzberg said. "That's the behavior of a profit-seeking, attention-seeking company, not a scientist."

The market noticed. Shares of CRISPR Therapeutics and Editas Medicine fell on the day of the announcement, though Editas gained 2.2 percent the next day and CRISPR was little changed. Investors briefly priced in the possibility that AI-accelerated biology could disrupt gene-editing companies, then pulled back as scientists flagged the gap between the announcement and the evidence.

The core complaint targets the framing. Anthropic took a preliminary finding and called it a discovery. In biology, a pattern in DNA data is a starting point. The discovery comes when you figure out what the pattern does, prove it with experiments, and demonstrate a mechanism. Anthropic has done the first step and hinted at the second. The ART array is expressed as distinct short RNAs, which suggests something analogous to CRISPR's programmability might be at play. But "suggests" and "analogous" are doing a lot of heavy lifting for a company that headlined its announcement with the word "discovered."

What does the Mestre accusation mean for AI labs?

Then the story got worse. Mario Rodríguez Mestre, a biologist at the University of Copenhagen, said over the weekend that his team had already discovered the same pattern. The New York Times reported that Mestre, who regularly used Claude in his research, wondered whether Anthropic's team had learned from his conversations with the model. Anthropic denies this. Mestre says he is stopping all use of Claude regardless.

This is the kind of incident that should make any AI lab building research tools lose sleep. If a scientist feeds their unpublished findings into a commercial LLM through normal chat interactions, and the model's maker later announces a similar discovery, the trust contract is broken. Anthropic's denial may be completely true. The appearance of a conflict is enough to make every domain expert who might have used your model as a research assistant think twice.

For builders, the implications are concrete:

  • If your product ingests user conversations for training, you need airtight guarantees about how research-grade inputs are handled, or you lose the expert users who matter most.
  • The line between "the model learned this from public data" and "the model learned this from a user's unpublished work" is going to be litigated repeatedly. Your terms of service and your data pipeline need to survive that scrutiny.
  • One accusation of data appropriation can undo months of capability marketing. The headline lingers even when the denial is sincere.

This connects to a broader pattern. OpenAI's recent claim that its agents cracked a million-dollar math problem drew similar scrutiny, including an accusation from a mathematician that the models may have used his work without credit. The privacy questions around prompt data are becoming central to whether researchers will let AI tools anywhere near their workflows.

How should builders frame AI-assisted research?

The technical achievement is legitimate. Running 950 agents for 21 hours to filter 200,000 enzyme candidates down to 20 worth investigating is real work. A general-purpose LLM doing this kind of large-scale biological data mining at scale is notable. The problem is entirely in the framing.

Kauderer-Abrams said the report is "as much about how the research was performed as it is about the content of the underlying discovery itself." That is the right read of what Anthropic demonstrated. It is the wrong framing for what they announced. The blog post leads with "Claude autonomously discovered a novel enzyme system." That claim is what biologists are rebutting.

If you are building AI tools for scientific or research domains, the Anthropic episode offers a clear playbook:

  • Describe what the system did, not what it might mean. "Claude filtered 200,000 candidates and flagged 20 for review" is a verifiable claim. "Claude discovered a novel enzyme system" invites the rebuttal.
  • Separate the capability demonstration from the scientific claim. Anthropic's real story is that an LLM can do biological database mining at scale. That is a product. The ART pattern is a finding that needs peer review.
  • Get domain experts to validate the framing before you announce. Anthropic ran physical experiments, which is more than most AI companies do. The biologists who pushed back are the kind of experts who should have been consulted on the language before publication.
  • Treat the discovery bar as a feature. Harrington's closing suggestion was to "set the bar high now, so that when an AI actually discovers a fundamentally new biological mechanism, everyone appreciates how big a deal it is." That advice applies to any domain where you are making capability claims.

What happens when the bar is set by marketing

Anthropic and OpenAI are racing to demonstrate that their models can do science, not just chat about it. Dario Amodei and Sam Altman both want the headline that says AI made a discovery. That incentive pushes toward announcement language that overstates what the model achieved, which triggers backlash from domain experts, which makes the next announcement harder to believe even if it is genuinely significant.

The 89 percent of enterprise AI agent pilots that never reach production face a version of the same problem at a smaller scale. A demo that looks impressive in a controlled setting falls apart when it meets real-world complexity. The gap between "it worked in our lab" and "it works in your workflow" is where most AI projects die. Overselling the lab result makes the workflow failure more likely, because it sets expectations the system cannot meet.

For anyone building with AI agents in research, analysis, or discovery domains, the lesson is to calibrate claims to evidence. The Anthropic biology lab did something technically interesting. It also demonstrated exactly how to alienate the community whose validation you need to call the result real. Both outcomes are instructive.

Sources