Beyond Benchmarks: The High-Stakes Battle to Measure Frontier AI
By Bram de Vries · Reporting from Amsterdam ·
As compute costs eclipse payroll, the ability to accurately measure AI performance through independent evals is becoming a matter of corporate survival.
An engineer at one firm recently burned through six billion tokens in a single day—a volume of text that would fill a million thick novels. During a month-long experiment, the team spent $1.5 million on compute, a sum ten times the cost of their monthly payroll. Ryan of FedBowels said on the "a16z" podcast that these costs expose how the industry has mispriced intelligence.
Can you measure a mind?
AI is shifting from a helpful tool to a line-item expense. Ryan said that as token costs eclipse salaries, how a company measures return on investment becomes a matter of survival. A firm's success now depends on its "eval"—the ability to pinpoint which model delivers the best output for the lowest price.
This mirrors the old world of cargo ships and credit lines. A captain's profit or ruin depended less on the hull of the vessel and more on the precision of the audit. Ryan noted that models like Llama 4 can game public benchmarks while failing in private, real-world use. To stop this, he suggested independent evaluation groups to prevent the conflict of a firm auditing its own work.
The Policy Gap
Ben said government agencies can identify threats like bio-weapons or hacking, but they cannot perform the technical tests to detect them. He proposed a split: the state writes the rules, and private firms verify if those rules were broken.
Critics fear a runaway AI—a model that rewrites its own code to become god-like. They argue that the greed of a paid consultant cannot be trusted with safety. Only a government mandate, they claim, can force the rigorous testing needed to prevent a global collapse.
Laws move at a snail's pace while code evolves by the hour. By the time a bureaucrat builds a test for a frontier model, the model is already obsolete. The most durable global standards—from maritime insurance to the exact dimensions of a shipping container—did not come from royal decrees. They emerged from merchants and insurers who risked their own capital. A state-run regime would not increase safety; it would bake failure into the system.
Why "Sovereign AI" is a Myth
Ben called the push for "sovereign AI"—where nations build walled gardens of servers—a waste of silicon. He suggested a shared language of evaluations, similar to the inspectors used in nuclear arms control. Verification, not isolation, is the only way to manage a model that trains its own successor.
This drive is born of panic, but panic is a poor architect. Splitting the network adds cost without adding security. If nations cannot agree on how to measure intelligence, they cannot verify safety. The world remains vulnerable to a capability that no one sees coming because every government is staring into its own closed-door lab.
The market will price AI not through safety debates, but through the balance sheet. Intelligence is only as useful as it is legible. The firms and nations that survive will be those that treat evaluation as a bookkeeping task rather than a regulatory hurdle.