The model was taking the exam blindfolded

ARC-AGI-3, released by the ARC Prize Foundation in March, is a set of 25 interactive pixel mini-games with 183 levels and no instructions. Players figure out the rules by trial and error. The scoring is strict: the model must clear levels taking no more steps than the roughly 500 first-time human novices who set the baseline.

The official interface handed models a 64x64 grid of numbers per frame. Characters, mechanisms, paths, everything a human sees, compressed into 4,096 cells. Claude Opus 5 scored 40.68 through that interface. The exam meant to measure general intelligence was partly measuring how well models read spreadsheets.

Eyes, a photo album, and four sentences

VISTA replaced the grid with 512x512 screenshots and archived every frame in original form, each with a serial number. The model can inspect any frame, zoom into corners, read exact pixel values, and compare frames side by side. Flipping through the album costs nothing; only in-game actions count as steps. It also keeps two notebooks: GUIDE.md for inferred rules, WORKING.md for scratch. The prompt for each game is four sentences. Nothing is trained.

Just swapping numbers for pictures took GPT-5.6 Sol from 13.33 to 47.32, clearing 9 games instead of 1. It also used fewer tokens: 30.7 million per game versus 71.9 million for the text grid. A single 64x64 numeric-grid frame costs about 4,000 tokens; a 512x512 image needs about 308.

The full system: Claude Opus 5 at a perfect 100 across all 25 games and 183 levels, in 7,302 steps, or 0.43 times the human average. GPT-5.6 Sol at 98.27. Even a 320B open-weight model that managed 1.89 through the official interface reached 66.93 with VISTA.

More was not better. Raising context from 200K to 780K tokens dropped the score from 99 to 93.9 while tripling token burn. Enlarging images 16-fold dropped it to 88.3. The team designed for minimalism, and the numbers back them up.

Four footnotes before you quote the 100

First: public games only. The authors are honest that their models were trained after the public games were released, and custom harnesses currently cannot run against ARC-AGI-3's private set. The contamination question is open.

Second: Nvidia got there first. In August, Nvidia's AVO harness scored 100 percent on the same public set with Claude Opus 5, and used 6,624 actions, about 9.3 percent fewer than VISTA's 7,302.

Third: the photo album is free to browse. Humans cannot replay their visual memory losslessly; VISTA can, and none of that inspection counts as a step. That is a scoring choice, and it flatters the number.

Fourth: this is the second perfect score for the same model under a different harness. That is the actual signal. Summer's top solutions had the model write roughly 4,000 lines of Python per game to build simulators; VISTA's checkers notes ran to three sentences. Whatever the benchmark measures now, it is not the model.

What the 100 actually measures

It is not an AGI milestone. It is a diagnosis of how the industry has been evaluating agents: through a straw, then blaming the model for the view. Give the model the screen and a memory, and the frontier flagship does in three sentences what took programmers a summer.

The useful takeaway for anyone building agents is uncomfortable. The bottleneck was never intelligence; it was the interface. ARC Prize co-founder Mike Knoop put it as more "learning" happening outside the model's weights. Before buying a smarter model, check what your harness is withholding from the one you have.

And the caveat stays attached: a public-set perfect score is a demonstration, not a measurement. Treat it like a demo tape. The private set is where the claim gets tested, and nobody's harness is allowed in there yet.

Sources

  1. [1] HTX Insights — “Claude Gets a Perfect Score, GPT Gets 99: He Kaiming's Team Aces the AGI Exam” (Oct 5, 2026)Read source
  2. [2] AI Daily Digest (dev.to) — “VISTA Aces ARC-AGI-3” (Oct 5, 2026)Read source
  3. [3] OnMSFT — “NVIDIA's AVO Agent Scores 100% on ARC-AGI-3 With Claude Opus 5” (Aug 2026)Read source