ARC-AGI-3 is a test where an AI plays puzzle games with no instructions. It has to figure out the goal by trying things. Humans score 100% on it, according to ARC Prize. The best score on the contest's Kaggle leaderboard is now close to 56%. When the first milestone closed on June 30, the winning entry scored about 1%. That is a huge climb in a few months. So what changed? Part of the answer is the code wrapped around a small model. But the model itself also got newer, so the code cannot take all the credit. And some of the headline numbers need care.
What Actually Moved
ARC Prize runs the contest on Kaggle. Teams get no internet access during evaluation, so they cannot call ChatGPT or Claude. They run open models on the competition's own computers. The ARC-AGI-3 track offers $850K. Final submissions are due November 2, and final awards are scheduled for December 4.
Here is the documented climb. At the first milestone, Tufa Labs won with 1.21%, using a small open model, Qwen 3.6 27B. At the second, which closed on September 30, ARC Prize named Daniel Franzen the winner at 27.9%. The runners-up scored 23.8% and 22.5%. At the time of writing, the leaderboard showed Tufa Labs at 55.89 and Yi-Chia Chen at 48.59.
Two cautions. The leaderboard uses about half the test data, and the final results use the other half, so rankings may change. Also, Tufa Labs did not open-source for the second milestone, so its latest approach is not public. Prize rules require winners to open-source their code, so it would have to share it to win a final prize.
What a Harness Is and Why It Matters
A harness is the code around a model. It decides what the model sees, what it remembers and what it can run. Tufa Labs' winning "Duck" harness lets a small open model write and run Python code in a live session, like a programmer poking at each game. To keep playing without running out of memory, it drops the oldest messages and keeps only the system prompt and recent history.
The second-milestone winners show how much scaffolding can add. Reports say all three built on the Duck harness and an open Qwen model. Franzen's version added faster serving software, smarter compute scheduling, and clearer board images and action records for the agent. His notebook ran for more than eight hours on a single GPU. But the base model also changed. Milestone 1 used Qwen 3.6 27B, while Franzen used a newer Qwen 3.8 model. So the gain cannot be credited to the harness alone.
Speed matters because every run must fit inside the contest's compute limits. The score also rewards efficiency, not just finishing: ARC-AGI-3 compares how many actions an agent needs with how many humans need.
After the first milestone, Tufa Labs said its gains came from multimodality and better base models, not hand-built tools, and described its harness as lightweight and generic. That fits what the second milestone shows: the model and the harness moved together.
Claims Versus Verified Results at the Top
Harness effects show up at the frontier too. ARC Prize reports two kinds of results for commercial models. The standard harness is a neutral setup. It carries forward only the notes the model chooses to keep. The provider adapter is built by the model's maker. It keeps the model's hidden reasoning state between requests and compresses long histories.
On September 30, ARC Prize listed OpenAI's GPT-6.1 Sol at 52.7% on the standard harness and 96.4% on the provider adapter. For GPT-6 Astra, ARC Prize lists 62.7% against 99.9%. Both numbers are verified, but they measure different setups. ARC Prize labels them separately, so they should not be compared directly. Here the model stayed the same and only the harness changed, which shows how much the setup can matter.
Kaggle entries and frontier models are also different races. Kaggle entries are research systems built for a contest, with no internet and open models. Frontier results are collected separately through commercial APIs.
What Comes Next
In the near term, the final results will use the hidden half of the data after November 2, and prize winners must open-source their code. That will show how much of the climb holds up on unseen games.
The long-term question is bigger. ARC-AGI-2 went from single-digit scores at launch to above 50% by December 2025. Whether this approach leads to general intelligence is unknown. François Chollet, who created ARC, has been careful here. Asked on X whether solving ARC-AGI-3 would mean AGI, he answered, "We're not making this claim."