Anthropic's Claude Opus 5 has posted a major result on ARC-AGI-3, outperforming both Fable 5 and GPT-5.6 Sol on a benchmark aimed at testing broader reasoning ability. The reported score was 30.2 percent, far above the previous record of 7.8 percent set by GPT-5.6 Sol.

That gap makes the latest result notable not just as a narrow win, but as a sizable jump over earlier top performers. Based on the figures provided, Opus 5 nearly quadrupled the prior best score, suggesting a clear step forward on this particular evaluation.

The benchmark's developers also highlighted an unusual detail in the model's behavior. They said Claude Opus 5 independently produced what they described as reflection equations, something they had not previously observed from another model on the test.

Even so, the result is best understood in the context of one benchmark rather than as a final verdict on overall AI capability. Still, the ARC-AGI-3 performance positions Claude Opus 5 as a standout entrant in the current race to build systems with stronger general reasoning skills.