TL;DR
Vals AI just raised $40 million at a $400 million valuation, led by Andreessen Horowitz, per TechFundingNews. Revenue is up 8x since 2025. Customer base doubled. Team tripled in six months. The story is not the capital. It is what the capital signals. Before you buy an AI tool, you need an independent way to verify it works for YOUR use case. Not the vendor's demo. Not their leaderboard. Your actual workload. Vals AI is building the receipts infrastructure for that verification. They are positioning themselves as AI's independent scorekeeper.
The Vendor-Graded Homework Problem
For years, AI evaluation worked like this: OpenAI built GPT-4, OpenAI tested GPT-4, OpenAI published a paper saying GPT-4 was great. Anthropic built Claude, Anthropic tested Claude, Anthropic published results showing Claude was great. Everyone was grading their own homework.
That arrangement worked fine when AI models were research curiosities. It stops working the moment you deploy a $2 million AI platform into your core operations. Labs including Meta, OpenAI, Google, and Amazon ran internal tests, then submitted only their strongest-performing model variants to public leaderboards.
Why Independent Evaluation Matters
Gartner's February 2026 Market Guide for AI Evaluation and Observability Platforms delivers the structural insight: 79% of organizations have deployed AI agents in production. 18% have formal evaluation processes. That 61-point gap is the single most important statistic in AI infrastructure.
Tests verify deterministic outputs. Evals grade nondeterministic systems that require judgment. A calculator test is binary. Grading an essay requires rubrics, calibration between judges, tolerance for variation, and feedback loops when the rubrics themselves are wrong.
The AI evaluation platform market was valued at $1.6 billion in 2025 and is projected to hit $19.8 billion by 2034. Enterprise AI Procurement Validation commanded 34.2% of application revenue.
Vals AI's Specific Bet
Vals Smith is the product that clarifies the bet. Instead of selling you a generic benchmark, Vals AI generates custom coding benchmarks directly from your GitHub repository. The benchmark tests whether an AI coding tool works against YOUR actual codebase, YOUR architectural patterns, YOUR test suite. Not some Stanford dataset.
Founders Rayan Krishnan and Langston Nashold have been running this playbook since their $5 million seed round. The 8x revenue growth suggests the market is paying for exactly that: verification infrastructure that vendors do not control.
What to Ask Your AI Vendor
Before you buy: Who evaluated this tool, and are they independent? What is their methodology? What did they test against, and does that match your use case? Are the results tamper-evident? Can you re-run the evaluation? What is the evaluator's conflict of interest policy?
If a vendor cannot answer those cleanly, you do not yet have the receipts. You have marketing.
Doctrine Connection: Due Diligence Is Non-Negotiable
When I was running investor diligence at AIN, we had a rule: never commit capital based on vendor-provided metrics alone. Independent verification on your use case. Comparable data against alternatives. Third-party assessment of risk. The receipts. AI procurement is no different.
FAQ
Q: Does this mean all AI vendor benchmarks are useless?
No. Vendor benchmarks are useful as a starting point. But they are floor data, not ceiling data. Independent evaluation asks different questions.
Q: Is Vals AI the only player in this space?
No. Scale AI, Weights & Biases, and Arize AI are also building evaluation infrastructure.
Q: Why does not my current AI tool come with independent evaluation?
Vendors have no incentive to fund evaluation that might make them look bad. That is exactly why independent evaluation infrastructure is emerging as a separate category.
*Jeff Barnes has no personal position in any company, fund, or platform named in this article. demg.ai provides marketing education and operator resources, not investment advice.*