Series A & B · Lead investorVertical AI · Infrastructure · Security · Dual-useSan Francisco, CA
Infrastructure · Essay

Evals are the new unit tests, and the new moat

Traditional software is tested by checking that the same input produces the same output. AI systems don’t work that way. The same input can produce different outputs, many of them acceptable, some subtly wrong. That makes the question “does this work?” much harder to answer.

The teams that answer it well ship faster, sell into stricter buyers and build an asset competitors can’t easily copy.

Why evals matter more than models

Most production AI teams switch or upgrade models several times a year. Every change brings the same risk: did this improve the product overall, or did it fix ten cases and quietly break twenty? Without a strong evaluation suite, teams either move slowly or ship regressions to customers. With one, a model upgrade becomes a routine afternoon instead of a month of anxiety.

Models are rented. Evals are owned.

What good evaluation looks like

  • Built from real work. Test sets drawn from production cases, including the hard edge cases, not synthetic examples.
  • Graded by domain judgment. Rubrics written with the experts who used to do the work, so “correct” means what the customer means.
  • Run continuously. Every prompt change, model swap and data update is evaluated before it reaches customers.
  • Tied to outcomes. Offline scores that predict what actually happens in production.

Two kinds of companies we back

The first is infrastructure: companies building evaluation, observability and reliability tooling that enterprises use to run AI in production. Demand here grows with every agent deployed, and the best tools become part of the daily engineering workflow.

The second is vertical AI companies whose evaluation suites are themselves a moat. A company that has spent three years encoding what a correct insurance claim decision looks like, case by case, holds something a general-purpose model provider cannot download. That suite is what lets them adopt every new model faster and more safely than any new entrant.

What we ask in diligence

How many evaluation cases do you have, and where did they come from? How long does it take to qualify a new model? When did your evals last catch a regression before a customer did? Strong teams answer these quickly and specifically. That speed tells us a lot about how they’ll operate at scale.

For founders

Building in this space? We should talk.