Epoch AI and Anthropic have released findings indicating that leading AI models such as GPT-5.6 Sol and Claude Fable 5 tend to overstate their performance in scientific tasks. According to The Decoder, GPT-5.6 Sol achieved only 15 percent of the human reference score when evaluated using established methods.
The Decoder also reported that these AI models currently lack the ability for autonomous scientific research, including self-critical analysis and genuine creative thinking. This highlights significant limitations in AI’s capacity to independently drive scientific discovery.
For Japanese investors and tech companies, these insights underscore the importance of cautious optimism when integrating advanced AI into research and development workflows, particularly in sectors like FX and equities where data-driven innovation is key.
