September 9, 2026 · Anthropic / DataCamp
Anthropic's Fable 5.1 Scores 52.6% on Terminal-Bench-Science, More Than Double GPT-5.6 Sol's 22.4%
My take: The numbers Anthropic published for Fable 5.1 on Terminal-Bench-Science are striking: 52.6% against GPT-5.6 Sol's 22.4% and Opus 5's 29%, on a test that measures whether a model can plan and execute a full scientific investigation inside a terminal. That is not a marginal gap.
What is important to keep in mind is the context: this evaluation was run by Anthropic using its own testing tools, on its own model. When the company that builds the product also designs and runs the benchmark, it is not possible to know with certainty whether the methodology favors particular characteristics of that model. These numbers are worth reading with your own judgment and waiting for independent lab validations before treating them as definitive.
What is clear is the direction: AI is improving at technical reasoning and autonomous execution in real environments, not just in conversation. For those using AI models in analysis, research, or workflow automation, the practical question remains the same: how much are you testing models on your own real use case, versus how much weight are you giving to manufacturer benchmarks?
Want to use these tools? See the unbiased reviews or back to the news.