The 'Research Taste' Trap: Why Inherent's Benchmark Win Isn't the Breakthrough We Think

AI-generated image · Bay Street Wire
Opinion: Inherent claims its Faraday agent outperformed OpenAI and Anthropic at replicating research, but the real test isn't a leaderboard—it's the messy reality of the lab.
In the current AI arms race, we have become accustomed to the 'benchmark theater'—a cycle of releases where startups claim victory based on a specific, narrow metric. The latest contender is Inherent, a London-based lab founded by Google DeepMind alumni. As first reported by TechCrunch, Inherent recently emerged from stealth with a $50 million seed round and a bold claim: its AI agent, Faraday, beat OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.8 at the task of reproducing results from published scientific papers independently.
On paper, this is an impressive feat of efficiency. Inherent’s co-founder and chief scientist Edward Hughes notes that replicating research is a foundational exercise for human PhD students. More strikingly, Faraday achieved this using Qwen 3.6, a model with just 27 billion parameters—a fraction of the size of the frontier-scale systems it supposedly beat.
But as a practitioner, I have to ask: are we measuring scientific intelligence, or are we measuring a model's ability to mimic a PDF?
Inherent is betting on something they call 'research taste'—the instinct to design effective experiments and identify which ones are worth running. To instill this, they are utilizing reinforcement learning to reward outcomes rather than following rigid rules. Hughes told TechCrunch that the goal is to build a collaborative teammate that can independently pursue curiosities and present results for review.
While the ambition is commendable, there is a massive chasm between 'replicating published findings' and 'discovering new scientific knowledge.' The former is a closed-loop problem; the answers already exist in the literature. The latter is a non-linear, chaotic process that involves physical constraints, equipment failure, and the serendipity of the laboratory.
It is telling that Inherent has chosen not to build its own coding tools, instead relying on OpenAI’s GPT-5.5 Codex. While Hughes argues this mirrors how human scientists use existing software, it also highlights a dependency on the very frontier models Inherent claims to have surpassed in 'taste.'
If Inherent truly wants to move beyond the party trick of replication, they need to prove that Faraday can navigate the 'messy' part of science—the parts that aren't documented in a published paper. Beating a larger model on a specific task is a victory for efficiency and parameter optimization, but it isn't a victory for scientific discovery.
Until these agents can handle the unpredictability of real-world lab work, we are simply watching a more efficient version of a student summarizing a textbook. The 'north star' Hughes describes is a noble one, but the path to getting there requires more than just outperforming a benchmark. It requires moving out of the digital sandbox and into the actual lab.

