FPSA · ARC-AGI-1 · qualitative diagnostic

What the model actually draws

Real inference from the trained checkpoint below on real ARC-1 official-test puzzles — every grid on this page is an unedited model output, not a mock-up. Built to answer one question: when the aggregate board-exact accuracy is 0.68%, what does that failure actually look like, cell by cell?

Checkpoint arc1_blended_d320_L4_scaling_ema_s1_3868132/best.pt (val-selected, EMA weights) · official reported metrics: cell 80.15%, board 0.68% (n=4000 fixed subsample, seed 12345).

0%
100%

How to read this