Yes, it really happened — Fugu Ultra v2 has beaten Claude Fable 5.1 on coding, chart reading and multi-model orchestration benchmarks.
A lab that wasn’t on anyone’s shortlist just leapfrogged the two names the industry treats as favourites.
If your whole stack is welded to one vendor, that should scare you — and then make you greedy.
See the original announcement on X 👇
— @nikkei View the post on X →
Fugu Ultra v2 vs Fable 5.1: What Just Happened
Sakana shipped Fugu Ultra v2 today, and the claim is blunt: it tops Claude Fable 5.1 and GPT-6 Astra on coding, chart reading and multi-model orchestration.
In plain English, it writes better code, reads data charts more accurately, and coordinates teams of models better than the two flagships everyone defaults to.
It also beat GPT-6 Astra on the same sweep, which matters more than the headline suggests.
When two supposedly untouchable models lose to the same challenger on the same day, you stop having a two-horse race and start having a market.
I care most about the orchestration number.
Orchestration means the model doesn’t just answer a prompt — it splits the job, picks the right model for each piece, runs them in parallel and stitches the results back together.
That’s glue work I’ve been doing by hand for two years.
One honesty note before the hype carries you off: these are the lab’s own day-one benchmarks.
Treat them as a strong signal, not a signed contract.
I never move a workflow on launch-day numbers, and you shouldn’t either.
What Fugu Ultra v2 Wins On
Coding is the headline, and it’s the one that stings the incumbents most.
Beating Fable 5.1 on coding means agentic work — multi-file refactors, test writing, pull requests that survive review — not just clever one-shot snippets.
For operators, coding wins are table stakes now, because every frontier release claims them.
Chart reading sounds dull until you’ve watched a frontier model confidently misread a bar chart.
Most models still squint at anything that isn’t a clean CSV.
A model that reads messy real-world charts accurately deletes a whole category of human double-checking from your day.
But orchestration is the quiet killer feature.
The most expensive part of my stack was never the models — it was the plumbing between them.
Fugu Ultra v2 moves that plumbing into the model layer itself.
Take those three wins together and you get the exact job description of an operator: build, analyse, delegate.
That’s why this sweep matters more than a single leaderboard jump.
The Moat Is Gone — and That’s Good News for You
Here’s the part that should reshape how you plan your quarter.
The lab behind Fugu Ultra v2 wasn’t one of the three names you’d have guessed.
For years the working assumption was simple: frontier models come from a handful of giants, and everyone else follows at a discount.
That assumption just died.
The moat is gone — what’s left is your workflow.
The moat was never really the models.
It was the belief that nobody else could build them, and that belief set the prices, the pace and the terms.
When a challenger you’d never shortlisted can match the leader overnight, lock-in stops being a strategy and starts being a tax.
Recognise the pattern, because it won’t be the last one: talent is everywhere, compute is rented, and recipes travel fast.
Your leverage comes back the moment you act like a buyer instead of a tenant.
Expect pricing pressure, faster release cycles and sharper terms across the board.
Competitive markets are dull for vendors and wonderful for operators.
Old Way vs New Way: An Operator’s Day
This is where the trend actually lands on your desk, so let me make it concrete.
The old way, my day looked like this: I picked one vendor, hoped its coding model stayed competitive, and became the human glue between every tool that couldn’t talk to the others.
Every chart got screenshotted, pasted, squinted at and re-checked.
Every multi-model job meant scripts, waiting and manual merges.
And when my vendor slipped behind, I waited months for the next release, because switching meant re-plumbing everything.
The new way flips each of those steps.
An orchestrator routes each task to whatever model wins it that week, reads the charts itself, and merges the output before you see it.
You review decisions instead of assembling artefacts.
| Old way — single-vendor era | New way — Fugu Ultra v2 era |
|---|---|
| One coding model, whatever your vendor ships | Best model per task, routed automatically |
| Manual screenshot-and-squint chart analysis | Charts read and interpreted inside the run |
| You glue models together with scripts | Orchestration built into the model layer |
| Locked to one roadmap when it slips behind | Swap the losing model the day it loses |
| Switch cost: two weekends of re-plumbing | Switch cost: one afternoon behind a router |
Read that last row twice.
The old way taxed every improvement; the new way compounds it.
In my own week, the human-glue hours — model swaps, paste relays, merge fixes — ran to the best part of a working day.
That’s the budget an orchestrator hands back to you on day one.
Same skills, same calendar, radically different leverage.
How to Act on Fugu Ultra v2 Today
You don’t need permission or a migration budget — here’s the exact play I’d run this week.
- Rerun your own eval, not theirs: take five real tasks from last week and run them through Fugu Ultra v2 blind against your current model
- Put it behind a flag: route a slice of your coding traffic to it and compare pass rates on your actual codebase
- Test orchestration hardest: hand it one job that normally needs two models plus you as the relay
- Break the chart reader: feed it your messiest screenshots, not clean benchmark graphics
- Wait for independent evals before any paid commitment, because day-one claims deserve day-two proof
- Rewrite your model-routing checklist so any new name can slot in within a day
Notice what that list doesn’t say: rip out your current stack.
Champions get promoted on evidence, not press releases.
Give the eval to whoever will actually run the model, not whoever bought the last one.
The goal is to make switching boring — an afternoon’s work instead of a quarter’s project.
Optimise for that once, and you never have to care this much about a launch day again.
FAQ
Did Fugu Ultra v2 really beat Claude Fable 5.1?
On the lab’s own published benchmarks, yes — coding, chart reading and multi-model orchestration all went its way.
Those are day-one claims, so I treat them as provisional until independent results land.
The right response is cheap verification, not blind belief or reflex dismissal.
What is multi-model orchestration, exactly?
It’s one model acting as manager: it breaks a job into parts, assigns each part to the best model, runs them and merges the results.
Until now, most of us did that management by hand with scripts and copy-paste.
It’s the least glamorous benchmark on the list and, for operators, the most valuable.
Should I switch my coding stack today?
Don’t rip anything out on launch day.
Run the five-task eval behind a flag, keep your current model as the default, and promote only what wins on your own work.
Switching should feel like an afternoon now, not a migration — that’s the real headline.
Why does a small lab beating the frontier matter to me?
Because it kills the last excuse for lock-in.
If frontier quality can come from anywhere, you get to shop on price, speed and terms instead of brand loyalty.
Your leverage as an operator just went up — whichever model you run this week, Fugu Ultra v2 included.
Also on our network: juliangoldie.com · juliangoldie.co.uk