Real inference from the trained checkpoint below on real ARC-1 official-test puzzles — every grid on this page is an unedited model output, not a mock-up. Built to answer one question: when the aggregate board-exact accuracy is 0.68%, what does that failure actually look like, cell by cell?
Checkpoint arc1_blended_d320_L4_scaling_ema_s1_3868132/best.pt
(val-selected, EMA weights) · official reported metrics: cell 80.15%, board 0.68%
(n=4000 fixed subsample, seed 12345).
evaluate() compares against
Target — cropped to the target's own bounding box, since ARC output grids can be a
different size than the input. The fourth panel only appears when the shapes differ: it
shows what the model drew over the input's footprint, so you can see whether it
even attempted to shrink the grid.