
Sakana AI has released Fugu-Cyber (model ID is fugu-cyber-v1.0), a cybersecurity-specialized addition to its Fugu orchestration family. It is not just a new frontier model. It is a third endpoint on the Fugu orchestrator, tuned for security reasoning. Sakana launched that orchestrator a month earlier. Sakana reports a success rate of 86.9% on CyberGym and 72.1% on CTI-REALM. It describes those results as comparable to cyber-focused frontier models such as GPT-5.5-Cyber and Claude Mythos Prev
Will Fugu-Cyber achieve a higher CyberGym score than 86.9% on any public leaderboard by October 2026?
Resolves by Oct 31, 2026
Sakana AI released Fugu-Cyber, a cybersecurity-focused orchestration model that coordinates multiple specialist AI models to solve security tasks. The model reported performance scores of 86.9% on CyberGym, a benchmark testing vulnerability discovery and proof-of-concept generation, and 72.1% on CTI-REALM, a benchmark testing threat detection rule creation from security reports. These scores represent marginal improvements over comparable frontier models, with access restricted through manual application review, acceptable use policies limiting defensive use, and pricing at a 20% premium over the base Fugu model. The orchestration approach uses specialized roles for different stages of security analysis, with verification emphasized as the critical component before any patches are proposed.

Datalab has released Marker 2, a full rewrite of its open source document conversion pipeline. Marker converts PDF, image, PPTX, DOCX, XLSX, HTML, and EPUB files into markdown, JSON, HTML, or chunks. The Datalab team rebuilt it around three components shipped over the preceding months: Surya OCR 2, a 20M-param fast layout model, and a rebuilt pdftext that is 3× faster than the previous one. The main result comes from olmOCR-bench, a third-party benchmark from Allen AI. Marker 2’s balan

Open speech recognition stopped being a Whisper monoculture some time in the last twelve months. In March 2026 Cohere released Transcribe, a 2B Apache 2.0 model that took the top of the Hugging Face Open ASR Leaderboard at 5.42% average word error rate. Five weeks later IBM shipped Granite Speech 4.1 2B at 5.33%. Since then ARK-ASR-3B and MOSS-Transcribe-preview-2B have posted lower numbers still. The top of that leaderboard is now separated by less than one WER point. That has a specific co

"I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," one expert told TechCrunch.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven