[funding] · · 4 min read
Vals raises $40M Series A led by Andreessen Horowitz
The startup argues legacy benchmarks are outdated and sells proprietary, industry-specific evaluations to AI labs and federal agencies.
By ByteBulletin Editor · Editor

AI-generated illustration · Z-Image-Turbo, self-hosted
Vals, a San Francisco-based AI benchmarking startup, has raised a $40 million Series A round led by Andreessen Horowitz, according to TechCrunch. The funding follows a seed round led by 8VC and Bloomberg Beta secured last year. The company, founded in 2024 by 25-year-old Rayan Krishnan, positions itself as a solution to a critical gap in the AI industry: the inability of legacy academic benchmarks to accurately measure the capabilities of modern, rapidly evolving models.
Krishnan, who previously interned at Palantir and worked at Microsoft and Stanford’s AI lab, argues that the current benchmarking landscape is broken because many tests are publicly available, allowing companies to train their models specifically to pass those tests. “We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance,” Krishnan told TechCrunch. Vals differentiates itself by keeping its test materials proprietary, preventing models from being overfitted to specific exams.
Proprietary tests and industry-specific metrics
Unlike traditional benchmarks that often measure abstract intelligence or general knowledge—such as passing a bar exam—Vals focuses on real-world task completion in specific verticals. The company evaluates models on their ability to produce human-quality work in domains like law, finance, and coding. Krishnan emphasizes that Vals looks at the “real impacts of the models,” asking whether they can do work that produces a product of the same quality as a human within every domain.
The startup is also expanding its scope into more specialized and high-stakes areas. Krishnan revealed that Vals has developed benchmarks for recursive self-improvement, mental health, cybersecurity, and biosecurity. Notably, the company is working on evaluations related to the law of armed conflict, specifically testing how models apply the Geneva Convention. This approach allows Vals to analyze not just positive outcomes but also potential negative implications if models were deployed without proper safeguards.
Revenue growth and federal expansion
Vals’ business model involves charging AI companies for these evaluations. Krishnan compares the revenue model to how a student pays the College Board to take the SAT, noting that effective measurement helps companies troubleshoot and improve their models over time. The startup has seen significant traction, with revenue currently eight times what it was last year. The company has also tripled its headcount, growing from eight employees at the start of the year to a team of 25. Krishnan stated that the company plans to hire an additional 10 to 15 people and relocate to a larger office.
In a move that signals a broader institutional push, Vals recently launched a program centered around providing model evaluations to federal agencies. This expansion suggests that the company is positioning its benchmarking services as a critical component of government AI oversight and procurement decisions.
The shift toward public accountability
Krishnan sees Vals’ work as foundational to the future of AI business operations and public trust. With major AI companies like Anthropic and potentially OpenAI moving toward public listings, the need for rigorous, standardized evaluation metrics is becoming a central part of how these companies report their capabilities to investors and regulators. “I think as AI models become a core part of the economy and are diffused more broadly, the types of benchmarks and evaluations that we do are going to drive their usage and be a central part of how these companies submit public filings,” Krishnan said.
The rise of Vals reflects a broader industry trend where benchmarking is no longer just a technical exercise but a key driver of PR, investment, and regulatory compliance. As AI models are integrated into every part of society, the demand for verifiable, non-gamed metrics is likely to grow, making Vals’ proprietary approach a significant player in the emerging market for AI validation.
What it means for developers
For developers and AI teams, the rise of Vals highlights a shift in how model performance is validated. Relying solely on public benchmarks may no longer be sufficient for demonstrating true capability, especially in regulated or high-stakes industries. Teams should consider how proprietary, industry-specific evaluations can provide a more accurate picture of their models’ real-world performance. Additionally, the focus on negative implications and safety in benchmarks like those for biosecurity and cybersecurity suggests that developers need to be prepared to address not just what their models can do, but what they might do wrong.
What to watch
- Federal Adoption: Monitor how Vals’ evaluations are integrated into federal agency procurement and oversight processes.
- Competitive Response: Watch how other benchmarking providers react to Vals’ proprietary model and industry-specific focus.
- Public Filings: Track how AI companies referencing Vals’ metrics in their public filings or investor communications.
- New Domains: Keep an eye on Vals’ expansion into new benchmarking areas, such as recursive self-improvement and legal compliance.
Get the signal, not the noise.
One short email when it matters. No recaps of recaps.
SHARE
RELATED

[funding] ·
Insight Partners diversifies AI bets against OpenAI and Anthropic concentration

[funding] ·
Cognition raises $2B at $48B valuation as AI coding market expands

[funding] ·
Mistral AI raises €3B in Europe's largest tech equity round

[funding] ·
Generalist hits $3B valuation as it bets on video-learning robot brains

[funding] ·
Nscale Seeks $3.5B in Pre-IPO Financing Ahead of Potential UK Listing

[funding] ·