OpenAI’s GPT-5.6 Sol model achieved a 38.3 percent score on the ARC-AGI-3 benchmark when tested using its own API with retained reasoning and context compaction, according to The Decoder. This custom test harness allowed GPT-5.6 Sol to outperform Anthropic’s Opus 5, which scored 30.2 percent without any test modifications.
However, in the official ARC-AGI-3 testing environment, GPT-5.6 Sol scored significantly lower at just 7.8 percent, falling behind Opus 5’s performance. The discrepancy highlights how testing conditions and custom optimizations can affect AI benchmark results.
For Japanese investors and technology markets, these developments reflect ongoing competition in AI capabilities that could influence automation and data processing tools in sectors like FX and equities trading, where advanced reasoning models are increasingly valuable.
