[23]entry 23 of 29
AI Inference Accelerator
this entry is written twice, in full
- 1.in brief
int8 neural-net inference in SystemVerilog
- 2.in general
MNIST classification in hardware, where a wrong answer is a wiring problem.
- 3.in particular
An int8 MNIST classifier in SystemVerilog that agrees bit for bit with a NumPy model of its own integer pipeline, in 1,701 cycles an image.
- 1.in brief
a small neural network built as a circuit rather than a program
- 2.in general
Handwritten-digit recognition in hardware, where a wrong answer is a wiring problem.
- 3.in particular
A digit-recognising circuit that gives exactly the same answer, down to the last bit, as a software copy of its own arithmetic, in 1,701 clock ticks a picture.
note
A two-layer, int8-quantized MLP (784 → 16 → 10) built from scratch: eight pipelined multiply-accumulate lanes, ReLU, and an int32 → int8 requantize-and-saturate stage between the layers, simulated in Verilator2 over 1,000 real MNIST test images. It classifies 90.1% of them at an average of 1,701 cycles each, about 7.5× fewer than a serial single-MAC design, and a NumPy5 model that reimplements the RTL’s exact integer arithmetic agrees with it on every output, bit for bit; the PyTorch6 float baseline reaches 90.8% and agrees on 98.7%. Bit-exact rather than merely accurate is the point, because that is what caught a silent int32 → int8 truncation between layers, a signed/unsigned comparison in a testbench, a non-monotonic pixel quantizer and a float/int8 training mismatch, any of which could have hidden inside a respectable accuracy figure.
note
A small neural network built from scratch as a circuit, working in 8-bit whole numbers rather than decimals because a chip this size has no room for anything more. Eight multiply-and-add units work side by side, and between the two layers the large running totals are squeezed back down to 8 bits without overflowing. Simulated in exact detail on a thousand real handwritten digits, it gets 90.1% right, at an average of 1,701 clock ticks each, about seven and a half times fewer than doing one multiplication at a time. A software copy that repeats the circuit’s exact whole-number arithmetic agrees with it on every single answer, down to the last bit, while ordinary machine-learning software using decimals gets 90.8% right and agrees on 98.7%. Agreeing exactly, rather than merely scoring well, is the point: it is what caught a total being silently cut short between the layers, a sign mix-up in a test, a pixel conversion that got its order wrong and a training mismatch between decimals and whole numbers, any of which could have hidden behind a respectable score.