- announcement
- grasshopper
Introducing XPU Grasshopper, an AI chip co-designed with AI for the era of recursive self-improvement
Grasshopper is co-designed with autonomous AI agents, and in 13 weeks it went from 0.06 to 51 tokens per second, 816x faster. The devkit is open for pre-order.
Today we are announcing XPU Grasshopper, our XPU architecture running AI models end to end, and opening pre-orders for its devkit, XPU Grasshopper (KU040).
A conventional chip takes a hundred engineers and years of handoffs from architecture to RTL to software. Grasshopper takes a handful of engineers and autonomous AI agents that co-design the hardware and the software together. The agents rewrite kernels, rework the hardware overlay, schedule and run the FPGA builds themselves, and test each new image on the full model. Whatever makes the model faster stays, and becomes the starting point for the next build.
That took Grasshopper from zero to running models in weeks. The first tokens came out on June 30, at 0.06 tokens per second. On October 2 the same model ran at 51 tokens per second, 816x faster in 13 weeks.
Fig. 1. gemma-3-1b-it generating live on XPU Grasshopper.
On gemma-3-1b-it, four tiles on an FPGA run at 51 tok/s using 89% of their peak memory bandwidth. For reference, an NVIDIA Jetson Orin Nano Super runs the same model at 41 tok/s. The devkit puts the same architecture on your desk, in a limited edition of 100 numbered boards.
The name is inspired by the SpaceX Grasshopper, the test rocket that learned to take off and land again years before SpaceX landed its first orbital booster. XPU Grasshopper is that first hop for polymorphic computing.
The Grasshopper architecture
In AI, almost all of the energy goes to moving data around, and very little to the math itself. Moving a byte costs far more than computing with it, so Grasshopper keeps the data next to the math. It is a mesh of Mutation Tile Units (MTUs), each with its own scratchpad right beside its compute, and DRAM sits at the edges of the mesh.
The dataflow reconfigures for each model instead of freezing one paradigm into silicon. It is also simple enough for AI agents to program, which is what makes the co-design loop possible.
The build we benchmark is four tiles in a 2x2 mesh on an AMD VU47P. Gen 1 silicon is the same architecture with 196 tiles.
XPU Grasshopper (VU47P), today
| Mutation Tile Units (MTU) | 4, 2x2 mesh |
| Compute per tile | 28.8 GOPS INT8 |
| Compute, total | 115 GOPS INT8 |
| BRAM | 512 KiB/MTU, 2 MiB total |
| DRAM | 16 GB, 4 ports |
| Interconnect | PCIe 4.0 x8 |
| Clock | 225 MHz |
| FPGA | AMD VU47P |
Gen 1 silicon (design)
| Mutation Tile Units (MTU) | 196, 14x14 mesh |
| Compute per tile | 2.05 TOPS INT8 |
| Compute, total | 401 TOPS INT8 |
| SRAM | 2 MiB/MTU, 392 MiB total |
| DRAM | 48 GB |
| Interconnect | PCIe x16, QSFP-DD |
| Clock | >= 1 GHz |
| Node | TSMC N7 |
Fig. 2. XPU Grasshopper (VU47P) as it runs today, next to the Gen 1 silicon design.
Co-designed with AI
The curve below is that loop running on gemma-3-1b-it at INT8, batch 1. It is the same model the whole way, measured end to end on the 2x2 mesh, and every point is the best run of that day.
816x
faster in 13 weeks, from 0.06 to 51 tok/s
Each run starts from the best build so far, so the curve only ever steps up.
- 1First tokens, one tile
- 22x2 mesh, DRAM online
- 3AI finds a hardware optimization, 12 to 21 tok/s
- 4AI kernel optimizations, 21 to 25 tok/s
- 5AI finds a second hardware optimization, 28 to 35 tok/s
Fig. 3. gemma-3-1b-it at INT8, batch 1. The full model, one user, end to end on XPU Grasshopper (VU47P).
What moved the curve
- June 30. First tokens, on a single tile.
- August 19. First run on the 2x2 mesh, with DRAM online.
- September 7. The AI finds a hardware optimization, and the next run goes from 12 to 21 tok/s.
- September 8. Its kernel optimizations take that image from 21 to 25 tok/s.
- September 14. A second hardware find takes the same kernels from 28 to 35 tok/s.
- September 14 to October 2. The loop keeps running, from 35 to 51 tok/s.
Our engineers set the targets and the architecture. Within them, the AI proposed these changes, built them, measured them, and kept the ones that made the model faster. This is an early form of recursive self-improvement, with AI improving the hardware that AI runs on.
Measured end to end
When a single user is decoding, every machine is limited by memory, because each token reads every weight once. So next to tokens per second we also show how much of each machine's peak memory bandwidth it actually uses (MBU).
gemma-3-1b-it on XPU Grasshopper (VU47P). For scale: Jetson Orin Nano Super1
tok/s
Higher is better
Weight bandwidth
Higher is better
MBU
Higher is better
Fig. 4. gemma-3-1b-it, one user, batch 1. Our bar is the solid one.
All workloads
| Workload | Precision | tok/s | TTFT | FPGA core power |
|---|---|---|---|---|
| gemma-3-1b-it | PrecisionINT8 | tok/s51.0 | TTFT~180 ms | FPGA core power~21 W² |
Fig. 5. Measured end to end.
XPU Grasshopper (KU040)
The devkit is a PCIe card with two AMD KU040 FPGAs linked by SerDes, running the Grasshopper 2x2. Everything above the hardware ships open source, from the SDK down to the runtime, the driver, and libxpu. You own your intelligence. The runtime opens the card at /dev/xpu/N, and from there the whole device is yours to program. And since it is an FPGA board underneath, you can use it as one too.
It is a limited edition of 100 numbered boards. Pre-orders are 20% off, reserved with a fully refundable deposit, and come with batch #0 of early API and SSH access to the cluster. Boards ship in February 2027.
What is next
Next is a cluster of XPUs you can use over an API, then Gen 1 silicon, and then Monolith. Monolith connects many XPUs into one machine that behaves as one chip, and runs the whole loop on the same silicon, from agents to experience generation to training.
The loop that made Grasshopper 816x faster keeps running, and every turn improves the hardware that the next turn runs on.
Gen 1 for clusters and enterprise
Enterprise and cluster inquiries are for Gen 1 silicon, not the devkit. Tell us your power budget and what you want to run, and we will get back to you.
Questions or partnerships? Write to us at contact@zscc.ai.
Method
- 1XPU Grasshopper (VU47P) was measured on October 2, 2026, at INT8, reading 1.0 GB per token. The Jetson Orin Nano Super numbers come from a public llama.cpp benchmark at 25 W with clocks locked, running Q4_K_M at 0.81 GB per token. Both use a 128-token prompt and 256 generated tokens, report the median of 20 requests, and leave out startup. MBU is achieved bandwidth divided by each machine's peak. Source: Tiny LLM Benchmark: Jetson Orin Nano Super 8GB (yuvrajsingh.io, May 2026), raw data.
- 2FPGA core power is the VCCINT rail under sustained load, about 14 W before the request. It leaves out the other board rails, regulator losses, and host power.