The practice of benchmarking has become the primary method for artificial intelligence firms to validate their systems, distinguish themselves from competitors, and generate favorable publicity. However, many legacy testing frameworks are outdated and vulnerable to manipulation, as companies have discovered ways to train their models around public exam questions. Enter Vals, a startup founded in 2024 with a mission to overhaul this flawed system.
Following a seed round led by 8VC and Bloomberg Beta, the company announced last month that it has raised $40 million in a Series A led by Andreessen Horowitz. The investment follows a period of rapid expansion for the firm, which has grown from eight employees at the start of the year to a team of 25.
Rayan Krishnan, the 25-year-old co-founder, established Vals after observing that academic benchmarks were failing to keep pace with the speed of industry advancements. Previously an intern at Palantir and a researcher at Stanford University’s AI lab, Krishnan argued that evaluation systems need to move beyond abstract knowledge tests, such as bar exam simulations, and instead assess practical performance.
Unlike many competitors that release their test data publicly—a practice that can allow for data contamination—Vals keeps its evaluation materials private. The company focuses on measuring how well AI models perform complex tasks within specific sectors, including law, finance, and coding. The goal is to determine if the output quality matches that of a human professional.
“We’re looking at what are the real impacts of the models,” Krishnan said. “Can they do work that produces a product of the same quality as a human within every domain?”
The startup also evaluates potential negative outcomes, analyzing risks such as those found in cybersecurity, biosecurity, and mental health applications. More recently, Vals has expanded into niche areas, including recursive self-improvement and legal compliance with the Geneva Convention under armed conflict scenarios.
While the concept of a company paying for third-party validation may seem counterintuitive, Krishnan compares the revenue model to students paying for SAT exams. Effective measurement helps organizations identify weaknesses and improve their products over time. Consequently, these evaluations are becoming critical factors in enterprise purchasing decisions.
Reporting that its revenue has increased eightfold compared to last year, Vals is currently scaling its operations. The company, headquartered in a historic brick building on San Francisco’s Folsom Street, plans to hire an additional 10 to 15 staff members and move into larger quarters. It has also launched a program to provide model evaluations to federal agencies.
As major AI firms prepare for public offerings, including anticipated listings by Anthropic and OpenAI, Krishnan believes independent benchmarking will become integral to corporate transparency. He predicts that rigorous evaluation standards will drive usage and serve as a central component of public filings and investment prospectuses.
Finally, someone addressing data contamination in AI testing. The industry needs this kind of rigorous, practical evaluation over vanity metrics.
Solid funding, but keeping benchmarks private feels like a conflict of interest. How transparent is their methodology really?