July 26, 2026 · MarkTechPost / Sakana AI
Sakana AI Claims 86.9% on CyberGym; Independent Tests Put Top Models at About 20%
My take: Sakana AI published results this month for its new cybersecurity model, Fugu-Cyber, with attention-grabbing numbers: 86.9% on CyberGym, a benchmark administered by UC Berkeley, and 72.1% on CTI-REALM. The problem is that CyberGym's own creators found that the best models on the market clear only around 20% under independent evaluation conditions, leaving a gap of nearly 67 points between Sakana's claimed score and what third parties have verified.
This is not unique to Sakana; it is the pattern that makes evaluating AI tools difficult. When a company runs its own benchmarks, on its own models, using its own methodologies, the resulting number is not necessarily a lie, but it is not the same class of evidence as an independent evaluation. As judge and party at once, we cannot know with certainty whether the test design is neutral. That distinction matters when you are deciding whether to adopt or recommend an AI tool for critical work.
Before trusting any benchmark a company publishes about itself, it is worth asking: is there an independent third-party evaluation that confirms those numbers?
Want to use these tools? See the unbiased reviews or back to the news.