Google capex cannot buy stable payrolls, Amin Vahdat warns

By Ray Dombrowski · Reporting from Youngstown ·

As Google infrastructure chief Amin Vahdat reveals the constant failures of giant AI clusters on Sequoia Capital, we must audit this capex boom against local payrolls.

On the Sequoia Capital podcast, Google AI infrastructure chief Amin Vahdat laid bare the fragile physics of the tech boom. Silicon Valley cheerleaders talk about a $200 billion capital buildout. Yet Vahdat admitted that at a scale of 100,000 accelerators, the hardware fails multiple times every hour. For anyone who has run a shear on a shop floor, that is an operational nightmare.

The vanity of theoretical throughput Vahdat argued that the industry's favorite metric, FLOPS, is just a vanity metric. Instead, he focuses on "goodput." This means the actual, useful work delivered through constant real-world failures. When a chip in a synchronous 100,000-accelerator cluster fails, the entire system can grind to a halt. The software must hunt for a checkpoint, restore the state, and restart the calculation.

This is a language any millwright from the old Campbell Works would understand. If your line stops multiple times an hour to reset a shear, your theoretical capacity means nothing. What matters is the tonnage that actually rolls out of the gate at the end of the shift. Google is learning that raw scale introduces a long tail of software bugs, compiler errors, and network failures. These errors eat away at efficiency.

The hard ceiling of the power grid Physical limits of this expansion are not found in code, but in the electrical yard. Vahdat pointed out that power is the single most fundamental, binding constraint facing AI infrastructure. A modern training cluster can demand a gigawatt of power. To put that in perspective, that is enough juice to run multiple heavy industrial electric arc furnaces.

Vahdat noted that Google prefers to connect directly to the utility grid rather than generate its own power. But you cannot simply call up a utility and demand a gigawatt of power for delivery tomorrow. It requires years of co-planning, upgrading transmission lines, and building substations. In the meantime, the physical footprint of these data centers is shifting. To avoid massive latencies between buildings, Google must design highly dense, water-cooled racks. These racks can pull hundreds of kilowatts each.

Scoring the capex boom in payroll What does this capex buildout actually leave behind? Vahdat described the lifecycle of these machines. While older TPUs still run at full utilization, the standard depreciation cycle is about six years. When a pod of chips is retired, technicians must pull out the entire physical assembly. They must retrofit the space for the next generation, like the new TPU for inference or training.

We have seen this story before with the Lordstown battery plant and the chip subsidies. We must check this new announcement against what the last round of promises actually delivered to the county employment series. When the Campbell Works shut down on Black Monday, 5,000 jobs vanished in a single morning. A gigawatt data center sounds like a heavy industrial plant. But it will never carry a Local 1112 headcount of 4,500 workers on a permanent payroll. Once the concrete is poured and the cables are plugged in, these automated warehouses do not support broad-based payrolls. They do not sustain the Mahoning Valley like the old mills did.

The tech sector's bet on AI is built on a massive scale, but its foundations are remarkably delicate. The system relies on optical circuit switches to steer light around broken racks in milliseconds. This is a high-wire act, not a stable utility. We must look past the vanity metrics of FLOPS and token generation to audit this massive expenditure. The only metric that matters for these communities is whether the work is still there in five years. It must put names on a permanent local payroll.

Sources

  1. Google's AI Infrastructure Chief, Amin Vahdat, on the Physics & Economics of Frontier AI