The Benchmark Problem at the Frontier
For years, model releases have come bundled with leaderboard claims. But as frontier models grow more capable, the traditional benchmark circuit is breaking down — and the stakes for getting model selection wrong have never been higher.
Andreessen Horowitz has announced an investment in Vals, making the case that the AI market urgently needs an independent scorekeeper. The problem the startup is solving is structural, not superficial. Public benchmark datasets get saturated, leak into training corpora, or become explicit optimization targets. A model can top a leaderboard and still fail when deployed on the kind of multi-step, domain-specific work enterprises actually pay for.
Those failures are getting more expensive. Models are no longer just answering questions — they're analyzing financial documents, resolving legal issues, writing and debugging software, and navigating complex enterprise workflows. Agents now operate autonomously over hours and days. A bad model choice can cost thousands of dollars in tokens, plus time and customer trust. That's a procurement risk, not just a technical one.
Real Work, Not Contrived Exams
Vals' methodology is deliberately grounded in actual professional tasks. Rather than asking whether a legal model can pass a bar exam, Vals tests whether it can perform legal research. Rather than a coding quiz, it tests whether a model can build a working application. In finance, it tests whether a model can analyze complex documents — not summarize a textbook chapter about them.
The team works with domain experts across fields to turn real workflows into rigorous benchmarks, then builds automated grading systems that can evaluate final work product to an expert standard. Co-founders Rayan Krishnan and Langston Nashold studied computer science at Stanford and had already been collaborating on real-world measurement problems before starting Vals — giving them credibility with the domain specialists whose cooperation the model requires.
Critically, private test sets are kept private with limited runs — a direct counter-measure to the contamination and gaming that has eroded confidence in public leaderboards. Results can be produced within hours of receiving model access, making the system practical enough to keep pace with frontier model releases.
Benchmarks Are Meant to Be Retired
One of Vals' more counterintuitive design principles: a good benchmark is supposed to become obsolete. When every system scores near-perfectly — a phenomenon the company calls "benchmaxing" — the benchmark has done its job. The correct response is to build a harder one, not to keep citing a saturated score.
In practice, this means continuously retiring evaluations and introducing new tasks as the frontier advances. In May, Vals replaced its CorpFin benchmark with a new Excel Modelling Benchmark after CorpFin stopped providing enough differentiation between leading models. That kind of iterative discipline is rare in a space where vendors often have every incentive to keep a favorable benchmark alive as long as possible. It also means Vals' value compounds over time — the harder it becomes to game their evals, the more their assessments are worth to buyers.
The Moody's Moment for AI
a16z's investment thesis maps Vals onto a familiar pattern in financial and product markets: when information asymmetry between sellers and buyers becomes large enough, trusted third-party measurement institutions emerge to make markets function. Credit markets produced Moody's and S&P; manufactured goods got UL certification; public markets rely on independent auditors. In each case, the institution emerged not because sellers volunteered transparency, but because buyers demanded it.
AI — rapidly becoming one of the largest markets in the technology industry — is overdue for the same dynamic. Vendors grading their own homework is the current norm. Vals is positioning as the answer to that conflict of interest, and a16z's backing is a signal that enterprise buyers are ready to pay for independence.
What This Means for Founders and Buyers
The investment signals a maturing shift in enterprise AI procurement. As agentic AI moves deeper into financial analysis, legal research, and software development, model selection is no longer a dev team decision made on vibes and vendor-supplied demos. It is becoming a procurement and risk decision that requires defensible, third-party evidence.
For startup founders building in regulated verticals — legal tech, fintech, enterprise software — the emergence of an independent evaluation layer raises the bar for how AI performance claims are made and verified. Marketing that leans on saturated academic benchmarks will increasingly look thin next to domain-specific evals that test what the model actually does on the job. Investors and enterprise procurement teams will start asking which independent benchmark a model has been validated against — and "we topped the MMLU leaderboard" will not be a sufficient answer for much longer.



