Skip to main content

Can AI Agents Design the Chip They Run On? Inside openTPU

An open-source project used AI agents to tune a small inference chip design for an FPGA. The results are promising, but most of the numbers come from the author.

By Mohammad Muneer Ahmed
Published: Oct 07, 2026
5 mins read
👁️ 47 Unique Views
Can AI Agents Design the Chip They Run On? Inside openTPU
The scale of inference: Optimized for multimodal workloads.
Premium Insight

Why It Matters

openTPU is open source under the Apache 2.0 licence, and most of it can be run and tested on a laptop with Python and Verilator, without owning an FPGA card. That makes it a free way for Indian students and small teams to study how an AI chip works, from the software down to the hardware.

A new open-source project asks a bold question: can AI agents design the chip they run on? openTPU, posted on GitHub by FeSens, is a small AI accelerator built with heavy help from AI agents. It covers the whole stack, from the hardware design to the compiler. It targets an older Kintex-7 FPGA card. The honest answer to the question is "partly." The details show where.

What openTPU Actually Is

The repository holds the hardware design in SystemVerilog, an instruction set, a simulator, a small kernel language with its compiler, and the software that drives a PCIe card. The machine is deliberately simple. One sequencer sends an instruction each cycle to a few units: one moves data, one multiplies, one does vector math. There is no cache and no hidden scheduling, so a trace shows where every cycle goes.

It runs small open models, such as Qwen3-0.6B, LFM2.5-230M and Qwen3.5-0.8B. Newer README versions list models up to about 4 billion parameters. The card has two DDR3 memory channels with a peak of 17.1 GB/s. That matters because decoding is limited by memory, not math. The README says the card reads more than 80% of that peak while generating text. High-end AI GPUs use much faster memory.

How the AI Part Worked

The AI work is an automated hill climb called a tournament. Each round, a few LLM agents read one hardware component and propose a change. Another agent writes it. A candidate must pass lint checks, bit-exact tests against the simulator, a cycle-count check and performance tests. It is kept only if it is smaller at the same estimated speed, or faster at the same size.

The first overnight run tried 96 changes across six components and kept 38. With some manual fixes, the estimate for the accelerator logic went from 41 MHz to 106 MHz, and from 132K to 82K lookup tables. Two cautions apply. These are yosys estimates, not results from the vendor's Vivado tool, and they leave out the PCIe and DDR3 controllers. And the "manual fixes" are a reminder that people were involved. One Hacker News commenter put it well: an experienced person pointing an LLM in a tasteful direction.

Claims Versus Evidence

An earlier README said openTPU had only run in simulation and had not run on the FPGA. Later versions report measurements on a physical card. Qwen3-0.6B decodes at roughly 22 tokens per second in 8-bit form, and about 34 with 4-bit weights. The README also says the card produces the same tokens as the simulator, bit for bit. The repo's design spec is dated September 23, and the README reports card measurements by September 28.

Treat these as the author's measurements. No independent reproduction was found in this research. The numbers have also been revised between README versions. In an early simulator test, the 8-bit model matched Hugging Face's original on only 2 of 8 prompts for 16 tokens. The README blames rounding from 8-bit storage, which is plausible, but it is still a reminder to check the claims.

There is a second gap. "The chip that runs their own inference" sounds like agents running on their own hardware. The README does not say that. The card runs small open models, roughly 0.2 to 4 billion parameters, not the agents that wrote the design. That is a research prototype, not a product anyone can buy.

What Could Come Next

In the near term, the README's own list is modest: faster prefill (processing the prompt) with a wider matrix unit, a faster clock, and the last few percent of memory bandwidth. Because decoding is memory-bound, a faster clock mostly helps prefill.

The long-term question is bigger. An FPGA can be reprogrammed to test a new design. Custom silicon cannot be changed once it is made. Whether agents can do this for a real chip, with fabrication costs and verification at scale, is a prediction, not a finding. What openTPU does show is a pattern. It works because the checks are strong. Every change had to match a simulator bit for bit before it counted. That is the same lesson as formal proofs: AI output becomes useful when something other than the AI can check it.

openTPU is a proof of concept, not a competitor to GPUs. As a learning tool and a test of agent-driven hardware design, it is worth watching.

Found this analysis insightful?

Share with colleagues, engineers, and your network.

Link copied to clipboard!