Transformer, pixel by pixel
A 128×72 greyscale buffer scaled up in hard pixels — roughly 700 rectangle fills a frame. Each attention head is a 3×3 weight matrix, the blanked upper triangle is the causal mask, and the bars at the far right are the softmax choosing the next token.