GPT-5.6 Sol Becomes the First Model to Score Above Zero on ARC-AGI-3
ARC-AGI-3 drops agents into interactive game environments with no instructions, no stated goals, and no rules — the agent discovers everything through action and observation, exactly like a human encountering an unknown game. Every previous frontier model scored effectively 0%. GPT-5.6 Sol scored 7.8% on the semi-private evaluation — a small number that represents the first time an LLM has shown interactive, goal-discovering behavior in a genuinely unseen environment.
Why it matters
💻 Developer · If you're building agents for genuinely novel, unspecified environments (not just well-defined tool-calling), this is the first real evidence a frontier model can do it at all — worth tracking closely.
📦 Product · Interactive goal-discovery capability, even at 7.8%, opens a category of agent product (navigating truly unfamiliar software/interfaces) that wasn't previously feasible at any level.
🎨 Design · This kind of exploratory, no-instructions reasoning is relevant if you're designing AI features meant to operate in unfamiliar or user-customized environments.
📈 Business · The gap between 0% and 7.8% is qualitative, not incremental — worth understanding for any long-term bet on general-purpose autonomous agents.
🤔 Just Curious · A new, extremely hard AI test drops the AI into a video-game-like environment with zero instructions — no rules, no goals explained. Every AI model scored basically 0% until now; OpenAI's newest model just became the first to actually figure out how to play at all.
Sources: ARC-AGI-3 result replay