The SOTA Mirage: Why AI Benchmarking is Currently a PR Game

AI-generated image · Bay Street Wire
Opinion: Until the industry moves away from cherry-picked data and outdated tests, 'state-of-the-art' claims are functionally meaningless.
In the current AI gold rush, 'SOTA' (state-of-the-art) has become the industry's favorite shorthand for superiority. But as a practitioner, I see these claims for what they often are: carefully curated PR wins. As TechCrunch first reported, when AI companies validate their models, they frequently lean on benchmarking metrics that swing in their favor to advertise dominance over competitors. The result is a 'wild west' of evaluation where the goal isn't necessarily accuracy, but marketing.
As TechCrunch reports, the problem is twofold: the tests are outdated, and the players are gaming the system. Many legacy benchmarking systems were not designed to measure the capabilities of modern models. Furthermore, because many of these tests are publicly available, companies can simply train their models against the test materials. In essence, they are cheating on the exam to achieve a high score.
This is why the emergence of companies like Vals is so critical. Founded in 2024 by Rayan Krishnan—a former Palantir intern and Stanford undergraduate who also worked for Microsoft—Vals is attempting to move the needle toward a more transparent, rigorous framework. According to TechCrunch, Krishnan observed that academic benchmarks were failing to keep pace with the rapid frontier advances of new models.
Vals is attempting to solve the 'cheating' problem by refusing to publicly disclose its specific test materials. More importantly, they are shifting the focus from abstract intelligence—such as whether a model can pass a bar exam—to real-world utility. Vals evaluates models on their ability to produce human-quality work in specific domains, including coding, finance, and law. They are even pushing into high-stakes territory like cybersecurity, biosecurity, and the application of the Geneva Convention via the law of armed conflict.
Investors are betting heavily on this shift toward standardization. TechCrunch notes that Vals secured a seed round led by Bloomberg Beta and 8VC, and recently raised a $40 million series A led by Andreessen Horowitz. The company's revenue has reportedly grown to eight times what it was last year, and its staff has tripled from eight to 25 employees.
From my perspective, this move toward third-party, paid verification—which Krishnan compares to the SATs administered by the College Board—is the only way to establish genuine public trust. As companies like Anthropic prepare to go public and others like OpenAI potentially follow, the stakes for accuracy in public filings and investment pitches are too high to rely on cherry-picked internal data.
If we continue to rely on legacy, open-source benchmarks that models can simply memorize, 'SOTA' will remain a meaningless buzzword. We need a framework that prioritizes real-world impact and negative implication analysis over abstract test scores. Until then, treat every 'industry-leading' benchmark with extreme skepticism.

