A recent study involving AI agents powered by Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.6 Sol assessed their ability to independently write AI research papers. According to The Decoder, these models were given six days, $3,000 in API credits, and GPU access to complete the task.
The study, conducted in collaboration with Princeton University and the UK AI Security Institute, found that original authors of unpublished NeurIPS papers rated the AI-generated submissions as 'Reject.' While the AI models managed to handle the full research engineering process, they struggled with research judgment, creative problem-solving, and abandoning unproductive approaches, The Decoder reported.
For Japanese investors and market participants, these findings highlight the current limitations of frontier AI models in creative and evaluative tasks, underscoring the need for human oversight in AI-driven research and development.
