- NO.
- 012
- DATE
- Updated 2026-07-20
- READ
- ~13 min
- KIND
- Notes
- STATUS
- Reviewed
AI Leverage and the Bread Paradox
Everyone can bake bread, and bakeries have lasted five thousand years. On Sakana Fugu and why buyers pay for convenience and accountability.
Four pieces of AI content crossed my feed recently, apparently about different things:
- Sakana AI released Sakana Fugu, packaging a multi-agent system behind an OpenAI-compatible model interface.
- Dan Koe's viral essay on restarting your life and escaping old ways of working, re-translated in Chinese circles as "how to survive mass AI displacement."
- A Chinese post about a 19-year-old who built an AI speed-camera system with Claude and sold it to the Hong Kong government for $550,000 in a month. I could not find an independently verifiable source for it.
- Joan Westenberg's "The Bread Paradox," answering the opposite anxiety — will SaaS die now that AI makes software cheap to write — with the observation that everyone can bake bread and bakeries have lasted five thousand years.
Read together, the interesting part is not which post went furthest. It is that they point at one shift: AI leverage is moving from "knowing how to use a model" to "knowing how to design a system." And the bread paradox supplies the economic floor under that shift — what buyers pay for was never the ability to produce the thing. It is the service of not having to think about it.
"Designing a system" is easy to misread, though. A system is not a few APIs chained together, and it is not a prompt that makes a model earn money on its own. A useful system has to verify its results, be deliverable to someone else, carry risk, and survive maintenance.
That is my basic read of this material: Fugu shows a real direction at the infrastructure layer; Dan Koe and Smith show individuals reorganizing their own work; and the $550,000 speed-camera story, absent further evidence, is better treated as a piece of engagement content that compresses a real trend into an overnight-riches myth.
Boundaries first: this endorses no company, tool, or post. Fugu's direction is worth discussing, and its effects still have to be verified on your own tasks. Personal narratives on X provide observational material, not evidence. What is worth keeping is not whether a story was exciting, but that it forces a distinction between three things: prototype capability, verifiable systems, and trustworthy delivery.
What Fugu sells is orchestration, not a model
Sakana Fugu's headline is direct: Multi-Agent System as a Model.
That line matters more than the benchmarks on the page. It does not merely claim a higher score. It says: you do not have to manage a pile of models, agents, routes, and prompts — Fugu compresses that complexity behind one model interface. To a developer it looks like an OpenAI-compatible API. Inside, it is a system that selects, switches between, and coordinates expert models.
That is where the shape of AI products is changing.
We used to ask:
- GPT or Claude?
- Which model writes better code?
- How should I prompt this task?
Fugu wants the questions to become:
- What result do you want?
- Which set of models should collaborate on it?
- How does the system trade off quality, latency, cost, and compliance?
It turns "model choice" into a capability the product hides. The user buys the outcome, not the routing.
That is not a small change. Many software products end up here: the underlying choices get more complex while the interface gets simpler. Cloud hides servers, payment gateways hide banking rails, search engines hide ranking. Now model orchestration is being hidden behind an API too. For users that lowers the barrier to entry. For buyers and deliverers, the responsibility to verify does not disappear with it.
The bread paradox: a business that should have died, and lasted five millennia
Westenberg makes the same product logic legible through a much older example.
He owns a bread machine. It sits unused. Flour and yeast are cheap, the method has been public knowledge for five thousand years, and the machine kneads and bakes on its own. He still buys bread every week. Because "can make" and "worth making yourself" are different questions: planning, shopping, maintaining the machine, accepting a shorter shelf life — those mental costs add up past the cost of picking up a loaf at the store. Economics calls it the make-or-buy decision, and for most people, buy wins almost every time.
History is on the buyer's side: commercial bakeries operated along the Nile in ancient Egypt, Rome had bakers' guilds and mechanical kneading, medieval London had a bakers' company policing quality. Every generation fully possessed the knowledge of bread-making, and every generation kept buying bread. Today Americans consume roughly ten million loaves of mass-produced bread a day while bread machines gather dust in countless kitchens.
Westenberg's key line: a bakery never sold flour and a recipe. It sells convenience, consistency, accountability, and somebody to blame when something goes wrong.
Look back at Fugu and it is doing exactly that. Models, routing, prompts, coordination strategy — all the flour and recipes go into the back kitchen, and one API sits on the counter. You are not buying the right to use a recipe; you are buying not having to think about it. So "AI lets anyone build their own multi-agent system" is not a threat to Fugu, in the same way "everyone can bake bread" never threatened bakeries. The real competition happens on another axis: whose bread is more consistent, and who is answerable when it is not.
Read the benchmarks, don't swallow them
Fugu's page reports strong numbers across SWE Bench Pro, TerminalBench, LiveCodeBench, Humanity's Last Exam, GPQA-D, SciCode, and long-context reasoning, with Fugu Ultra first or near-first on many.
Worth reading, not worth taking at face value, for three reasons.
First, the page itself notes that some baselines use scores reported by the model providers. Evaluation conditions, toolchains, scaffolding, and sampling settings are hard to hold identical across vendors.
Second, Fugu's routing is not public. The FAQ says plainly that which models are used and how they are coordinated is proprietary and not exposed. Understandable as a product decision, and it also means an outside user cannot reproduce the same chain of judgment.
Third, what decides usability in production is not only a score. It is response time, stability on long tasks, cost predictability, how it degrades on failure, and whether data handling meets your compliance requirements.
So the stable reading: Fugu's benchmarks establish that multi-model orchestration may be an effective direction. They do not establish that it beats calling a single frontier model directly in your particular business.
A credible trend signal, not a purchase order that skips verification.
The get-rich story omits the hard part
Now the post about the 19-year-old, Claude, a speed-camera system, and $550,000 from the Hong Kong government. I treat it as a sample of how such stories spread, not as a confirmed commercial case.
The narrative is standard: young founder, low cost, short timeline, large client, government procurement, on-site demo, immediate payment. Nearly every element serves one emotion — AI has suddenly given an ordinary person the productive capacity that used to require a company.
The emotion is not entirely false. A model like Claude really does let one person reach a prototype faster. Someone with a bit of engineering, product sense, and communication skill really can build what used to need a small team.
The problem is what the story leaves out.
A traffic enforcement system is not "a camera that reads speed and plates." It also involves:
- Measurement accuracy and who is accountable for calibration.
- Whether video evidence is admissible for enforcement.
- Where owner identity data comes from and who may query it.
- Whether automatic ticketing satisfies administrative process.
- How misreads, appeals, audit, and log retention are handled.
- Whether government procurement can be bypassed by one USB-stick demo.
- Why a $550,000 payment would appear as an on-the-spot cheque.
Claude writing the prototype is not the hard part. The hard part is getting a system into the real world and having an organization with legal liability adopt, accept, and maintain it.
In the language of the bread paradox: the story claims someone sold a bakery's worth of value on the strength of being able to bake. But the buyer was never paying for baking ability. They were paying for someone to get up at 2 a.m. when the system is down, for logs that exist when the auditor arrives, for an accountable party when a ticket is issued in error. Taking on those obligations is the actual good that $550,000 corresponds to. The prototype is the bread; procurement is buying the bakery.
So I do not read stories like this as a roadmap. They work better as a warning in reverse: the get-rich narrative spreads because it captures part of a real trend while folding source verification, demand validation, trustworthy delivery, compliance, sales, and maintenance into the phrase "used Claude."
That is dangerous for a reader, because it suggests a prompt is a business model, a prototype is a product, and a demo is a closed deal.
The genuinely useful part of Dan Koe-style content
Dan Koe's writing, and Smith's Chinese rendering of it, address the other end: how an individual faces AI displacement.
This genre mixes self-management, creator economy, skill recombination, and anti-employment rhetoric. Its weakness is a tendency toward hype. Its strength is that it grabs a real point: if AI makes ordinary execution labor cheaper, an individual has to move from "the person who completes tasks" to "the person who designs systems."
The point is not a headline about escaping wage slavery. It is a shift in how work is organized:
- From waiting to be assigned tasks, to defining the problem.
- From producing deliverables, to designing reusable processes.
- From a single skill, to a combination of skills.
- From a résumé asserting ability, to public work and real cases demonstrating it.
Which is structurally the same move Fugu makes.
What Fugu does at the infrastructure layer is coordinate several models into a stronger system. What an individual has to do at the career layer is coordinate writing, code, research, sales, delivery, and retrospectives into a system that keeps producing.
Only the scale differs. One is a company's product; the other is a personal workflow.
Which software dies, which survives
The bread paradox also answers a practical question: once AI lowers the cost of producing software, what gets eliminated?
Westenberg's judgment is that the genuinely endangered products are thin single-function tools an AI can replicate from one sentence — PDF converters, meeting-notes generators. Their nature is "price one feature as if it were a company," and when the marginal cost of the feature approaches zero, that pricing cannot hold.
But mature platforms like Notion or Jira stopped selling code long ago: integration ecosystems, compliance certifications, a stability record, the usage habits an organization accumulates, and most of all — someone accountable when it breaks. He also cites a number worth remembering: AI-generated code carries roughly 1.7× the serious defects of human-written code. Building in-house means bearing not only the generation cost but every subsequent day of maintenance, security, and knowledge loss. When the person who wrote that internal tool leaves, who picks it up?
The same line applies to individuals. If what you offer is "I can build a feature with AI," you are on the thin-tool side, competing with zero cost. If what you offer is "hand me this part of the process, and come to me when it breaks," you are on the bakery side — and the stronger AI gets, the cheaper your raw materials become.
Where the leverage actually is
Compressed into one line:
Real AI leverage is not "which model I used." It is whether I can connect model capability into a system someone can trust.
That system has at least five layers.
One: a specific scenario. Do not start from "I want to make money with AI." Start from a real, narrow, verifiable problem — job-description parsing, code review, customer record cleanup, first-pass research screening, document migration, internal knowledge Q&A.
Two: verifiable steps. AI can extract, generate, compare, and summarize, and each step should have an acceptance criterion. Whether the result is right cannot rest on the model's own say-so.
Three: human accountability. AI produces prototypes, drafts, and candidate options; a person makes the final call. The closer you get to law, finance, medicine, hiring, enforcement, and public services, the less the responsibility can be pushed onto a model.
Four: delivery and trust. Why would anyone rely on your system? How is data handled? What happens on failure? Who maintains it? Where are the logs? How are permissions managed? None of it is glamorous, and it decides whether a prototype becomes a product.
Five: reuse. A one-off prompt is not a moat. Reusable data structures, checklists, workflows, evaluation sets, documentation, and delivery templates are what make capability compound.
Which is why I stay skeptical of "sold to a government for $550,000 in a month." Not because AI cannot produce the prototype, but because the real world does not pay for prototypes. It pays for risk handled, responsibility carried, process cleared, and results that reproduce.
What to do as an individual
If you already use AI but are stuck at the tool layer, the next step is not collecting more prompts. It is five things.
1. Pick one specific scenario. Not a general assistant. Something you know well, that repeats, where the pain is obvious. The more specific, the easier to verify whether AI actually helped.
2. Break the work into checkable steps. Turn "have AI do this task" into input, processing, output, acceptance. At each step ask: how would I notice it was wrong? Who confirms? Is there a reference example?
3. Prototype with AI, accept the result yourself. AI gets you to a first version faster; acceptance cannot be delegated to it. Anything going onto a résumé, to a client, to your manager, or into a user's decision gets human review.
4. Fill in distribution, sales, and compliance. Most people stall not because they cannot build a demo, but because they cannot find real demand, explain the value, handle data and permissions, or convince anyone it will keep running.
5. Consolidate into a reusable system. Write the steps that worked into documents, templates, scripts, checklists, and small tools. Real efficiency compounds; it does not come from starting a fresh chat every time.
Taken together, those five turn you from someone who bought a bread machine into someone who runs a bakery. Anyone can afford the machine — prompts and tools were never the moat. The bakery's moat is that people are willing to outsource this to you for the long run: you are consistent, predictable, and you own it when it breaks.
Conclusion
Products like Fugu show AI capability evolving toward system orchestration. The overnight-riches stories show how distorted the popular imagination of that capability has become. Dan Koe-style content shows individuals searching for a new way to organize work. And the bread paradox shows the economics underneath has not changed in five thousand years: people will always pay to outsource complexity to someone they trust.
The reminder is clear enough:
Do not treat AI as a wishing machine. Do not treat it as just a chat box either.
The accurate view: AI is lowering the cost of prototypes while raising the relative value of judgment, verification, delivery, and trust. The more people who can produce a demo, the scarcer the thing that is not a demo — the ability to turn one into a system others will depend on for years. The more people who can bake, the more a good bakery is worth.
That is what AI leverage actually is.
Comments
Comments are powered by GitHub Discussions. Sign in with GitHub to comment. Open the matching Discussion