Close-up of AI-designed hardware chip layout showing dense circuit traces and processing blocks

A chip design mistake costs millions and months. You don’t get to patch silicon after it ships, which is why hardware engineering has always been one of the slowest, most heavily reviewed disciplines in tech. So when an AI agent produces a working open-source TPU design, the interesting part is not the chip. It’s that AI output has arrived in a domain where being wrong is genuinely expensive.

That changes where your constraint sits. If an agent can draft the work in hours, the bottleneck moves to verification: who checks it, against what standard, and how fast. This post looks at what OpenTPU actually demonstrates about AI-designed hardware, and what it tells you about building verification capacity in quality and manufacturing operations before the volume of AI-generated work outpaces your ability to review it.

Everyone Said Silicon Design Was Too Hard for AI Agents

The argument against AI in chip design was always the same three points. The talent pool is tiny. The tolerances are unforgiving. And the knowledge that matters most never got written down, it lives in the heads of engineers who have taped out a dozen designs and know which corner cases bite.

Those are the exact arguments used on the plant floor. Nobody can run that line like Henk. The spec sheet doesn’t capture what a good weld looks like. Our tolerances are too tight for a model to guess at.

OpenTPU is a working TPU-style accelerator whose hardware description was written by an AI agent and published openly on GitHub. That does not make the argument wrong about tacit knowledge. It makes it irrelevant as a reason to wait.

Engineer reviews an AI-designed hardware chip floorplan with dense routing layers on dual monitors

What OpenTPU Actually Is, and What the AI Actually Wrote

Strip away the framing and OpenTPU is a matrix multiplication accelerator described in Verilog, with simulation and test benches to prove the logic behaves as specified. An AI agent wrote the bulk of the hardware description code under human direction. It runs in simulation. You can read it, fork it, and re-run the tests yourself.

What it is not: a taped-out production part. Nothing here competes with Google’s TPUs on performance, power, or yield. No fab, no mask set, no silicon. That distinction matters, because the gap between working RTL and a shipping chip is where most of the cost and nearly all of the risk live. The verified claim is narrower and more interesting: inspectable, reproducible hardware code in a discipline that was supposed to be closed to this kind of output.

The architecture in one paragraph: systolic arrays and why they suit matrix math

A systolic array is a grid of small identical processing elements. Data flows through the grid in a steady rhythm, each cell multiplying a pair of values, adding the running total, and passing results to its neighbour. No cell fetches from main memory on every operation.

That structure fits matrix multiplication almost perfectly, because the operation is just the same multiply-accumulate repeated across a predictable pattern of indices. The regularity is also why it suits AI-designed hardware. A highly repetitive, well-documented structure gives an agent a tight specification to work against, which is the same reason agents do well on standardised inspection logic and poorly on tribal process knowledge.

Where the human engineer stayed in the loop

The human set the architecture, defined the interfaces, and decided what correct looked like before a line of Verilog existed. The agent filled in the implementation. Direction came first, generation second.

Then came the part that actually determined whether any of it counted: running the test benches and reading the simulation output. The agent produced volume. The engineer owned the standard the volume was judged against. Nobody skipped that step, and the project would mean nothing if they had.

The Real Lesson: Generation Got Cheap, Verification Did Not

Writing the hardware description was the fast part. Proving it behaved correctly across every input pattern, timing assumption, and edge case took the real effort, and that work still ran at human speed because someone had to decide what “correct” meant before a single test could be written.

That inversion is the part worth carrying into your own operation. An agent can draft a work instruction, a control plan, a PLC routine, or an inspection script in minutes. None of that output has value until something independent confirms it holds. The drafting cost collapsed. The confirming cost did not move at all.

Which means the question in your next AI conversation is not “can the model do this task.” It is “do we have a way to tell when it got it wrong, and how long does that check take.”

Why strong test coverage is now an AI adoption multiplier

Teams with written acceptance criteria, automated checks, and documented pass/fail thresholds absorb AI output almost immediately. The agent produces a candidate, the harness grades it, and the human reviews the failures instead of reading everything line by line. Throughput goes up because the review scales without adding reviewers.

Teams where quality lives entirely in senior judgement stall on contact. Every AI-generated artefact lands on the same two or three experienced people, who now have more to read than before and no faster way to read it. The agent made the queue longer, not shorter. This is the most common reason AI pilots in manufacturing quietly die: nobody built the grader.

So treat verification capacity as the thing you invest in first. Write down what good looks like for the processes you want to automate. Convert tacit acceptance into explicit criteria, even crude ones. Build the checks before you build the generation. Every hour spent making quality measurable raises the ceiling on how much AI-produced work your team can safely consume, and that ceiling, not model capability, is what determines your return.

Engineer reviewing failed testbench output beside generated Verilog for AI-designed hardware

Translating This to Quality, Operations, and Manufacturing Work

Look at what your specialists actually produce. SPC rule sets, CAPA investigation reports, PLC logic, inspection routines, IQ/OQ/PQ protocols, FMEA tables. Every one of those is a structured artefact built by someone expensive, with a defined notion of correct and a real cost when it’s wrong. That profile is the same one hardware description code has.

Which means the same split applies. Drafting is cheap now. Confirming the draft is sound is where your people still matter.

A three-question filter for deciding what an agent can draft

Before you point an agent at any of it, run the artefact through three questions.

  • Is “correct” written down somewhere? If correctness lives only in a senior engineer’s judgement, you have no test bench. Start with artefacts governed by a standard, a spec, or a validated template.
  • Can you verify it faster than you can write it? A control plan you can check against a drawing in ten minutes is a good candidate. A CAPA root cause analysis that needs three days of plant investigation is not.
  • What happens if a bad draft gets through? Scrap and rework is recoverable. A regulator finding or a line crash is not. Pilot where the failure mode is cheap.

Artefacts that pass all three go first. Everything else waits until your verification layer is good enough to catch what the agent gets wrong.

What the ROI looks like: engineer hours redirected, not headcount removed

Do not model this as a headcount reduction. Your quality engineer’s value was never in typing up the protocol. It was in knowing which tests actually prove the thing works, and that skill becomes more valuable, not less.

The return shows up as reclaimed hours. A validation engineer who spends 60% of the week on documentation drafting and 40% on review flips that ratio. Same person, same salary, more throughput on the work only they can do. Measure it as cycle time on deliverables and backlog burn-down, not as FTEs saved.

Ready to find AI opportunities in your business?
Book a Free AI Opportunity Audit. It is a 30-minute call where we map the highest-value automations in your operation.

The Expert-Domain Threshold Is Moving Faster Than Hiring Plans

Most capability plans assume specialist scarcity is a fixed constraint. You budget for a validation engineer in 2027, a controls specialist the year after, and plan the roadmap around when those people land. That assumption is now the weakest part of the plan.

The threshold for what agents can draft has been moving one domain at a time, and it does not announce itself. It shows up as a discipline everyone called untouchable suddenly producing reviewable output. Hardware description was supposed to be years away from that. It wasn’t.

So plan for shorter shelf life. Treat any three-year hiring plan built purely on headcount as provisional, and ask what the same budget buys if drafting capacity stops being the thing you’re short of. The people you hire should be the ones who can judge output, not just produce it.

The first move: write down how you’d prove the output is correct

Pick one specialist task where correctness is objectively testable. A gauge R&R calculation, a batch release checklist, a PLC interlock routine with known pass and fail conditions. Then write the test before you write the prompt. What inputs must it handle, what outputs are unacceptable, who signs off, and what evidence satisfies them.

This is unglamorous work and most teams skip it, which is why their pilots stall at “the output looked fine.” If your acceptance criteria live only in a senior engineer’s judgement, you cannot scale review, and review is the constraint. Writing the criteria down converts tacit expertise into something an auditor and an agent can both work against.

Two quarters is enough time to do this properly. Quarter one: document verification criteria for three to five artefact types and build the evaluation harness, meaning the fixed set of test cases every draft must survive. Quarter two: run an agent against that harness on the single easiest task and measure how much reviewer time you actually save.

Source: github.com

Leave a Reply