- NO.
- 019
- DATE
- READ
- ~10 min
- KIND
- Notes
- STATUS
- Reviewed
Cheap Generation, Expensive Verification
Claude Code, agent swarms and AI bug reports all raise the supply of candidates. The scarce input is verifying results, ranking them, earning trust.
Read a few recent announcements together and they look like one production chain closing.
Anthropic shipped Claude Sonnet 5, emphasizing that agentic capability which used to require pricier models now sits in a lower price band, and lands directly in Claude Code. Simon Willison observed that coding agents have pushed the cost of trial and error in device reverse engineering and home automation down to "worth a casual attempt." Cursor's agent swarm experiments push generation volume further still: different model mixes reach comparable quality at very different costs, and the system needs version control and conflict handling redesigned just to absorb the commit rate.
On the other side, GitHub restructured its bug bounty programme. The stated reason was not a shortage of submissions. It was that AI-assisted reports increased the number of low-quality or hard-to-verify candidates along with everything else. GitHub now uses a Signal threshold, invitation-only projects, and trusted-researcher relationships to point limited human attention at the reports most likely to hold.
All four describe one mechanism: falling generation cost does not automatically increase valid results in proportion; it first increases candidates. Once candidate supply exceeds human verification capacity, the scarce input moves from "writing the first draft" to verification, filtering, and trust.
This post is not about whether AI replaces programmers, and not another argument that code needs tests. The question is narrower: as Claude Code and GitHub weld generation, submission, and collaboration into a faster and faster line, which stages that used to look like cleanup start deciding total throughput?
What gets generated is a candidate, not a conclusion
The easiest mistake cheap generation invites is counting candidates as output.
Claude Code can read a repository, change code, run tests, and submit a patch that looks complete. An agent swarm can split one large task into many parallel branches. AI can produce vulnerability analysis, remediation suggestions, and reproduction steps from public code. All of that is real.
But what it produces first is candidates:
- A candidate patch may have fixed the symptom, or changed some other boundary condition.
- A candidate vulnerability report may be a real issue, or may not reproduce, or may fall outside scope.
- A candidate analysis may cite facts, or may have written correlation as causation.
- A candidate article may be well structured while carrying no independent judgment and no stated limits.
Candidates are not waste. Without candidates there is nothing to choose between. But counting them as completed work makes an organization look increasingly productive on paper while getting increasingly busy in reality.
The reason is arithmetic:
total review cost = number of candidates × verification cost per candidate
AI may reduce part of the per-candidate cost — auto-generating tests, checking format, finding similar issues. As long as candidate volume grows faster, total review cost still rises. Worse, low-quality candidates do not remove themselves. They land in issues, PRs, inboxes, and report queues, consuming exactly the people with the most experience.
GitHub's problem moved from "will anyone report" to "who do we read first"
The bug bounty change is a clean sample of the mechanism.
A traditional bounty encourages more people to find and submit issues. When submissions are scarce that incentive is right: one more report is one more chance at a real vulnerability. Once AI lowered the cost of producing a report, volume stopped being the bottleneck. Whether a report reproduces, falls in scope, carries sufficient evidence, and shows an understanding of the project's context now determines whether the review queue functions at all.
Introducing Signal and a VIP programme does two things:
- Ranks candidate signals using historical quality and behavior.
- Institutionalizes the working relationship between high-signal researchers and maintainers.
It does not mean newcomers are not worth reading, and reputation does not replace technical verification. Reputation only helps decide what to verify first when verification capacity is finite. The report still has to reproduce, the impact still has to be assessed, the fix still has to be accepted.
What actually changed is attention allocation. The system used to optimize for "more people submitting." Now it must simultaneously optimize for "do not let cheap generation drown the verifiers."
The same shift shows up in code repositories, content moderation, job applications, academic submissions, sales leads — every entrance that lets AI produce candidates in bulk. The cheaper the entrance, the more the downstream needs a ranking mechanism.
Verification is not QA at the end; it is the production
In the handmade-production image, verification is the inspection station at the end of the line: the real work is done, you spot-check and ship.
With agents involved, that ordering fails. Verification has to start when the task is defined.
On this blog, an AI writing the code is not the finish line. A change passes through at least:
goal and boundaries
→ a reviewable diff
→ schema / link / content rules
→ type check and build
→ browser behavior tests
→ production verification
→ human confirmation for irreversible decisions
None of those layers is "looking for mistakes." Together they define what counts as the product: URLs are not silently changed, the two language versions do not cross-contaminate, downloads match their SHA-256, and static routes actually work on Cloudflare rather than only looking right locally.
I described AGENTS.md, the content lock, and the publishing gates in building a site with an AI agent, and I hit the counter-example — code claiming a fix while real behavior never changed — in one screenshot, a cascade of fixes. Together they point one way: verification is not stamping approval on a generated result. It is converting a candidate into a delivery someone can be accountable for.
Delete the verification layer and faster generation simply means errors reach production faster.
Filtering is not deleting junk; it is protecting judgment
"There is too much AI content, so we need filtering" sounds like an information management problem. It is a capacity allocation problem.
An experienced person can seriously review a limited number of complex items per day. Making them open low-quality candidates one by one — even at a few minutes each — squeezes out the work that actually needed their judgment. The goal of filtering is not zero false negatives. It is putting expensive attention where value, uncertainty, or risk is highest.
A mature filtering layer uses three kinds of signal:
- Machine-verifiable: schema, tests, build, formatting, duplicates, scope, reproduction steps.
- Task context: does it answer the original question, does it touch protected fields, does it carry business or security impact?
- History: the submitter's past accuracy, whether their fixes survived review, whether they record failures and limits honestly.
The first two automate well; the third is the basis of trust. The order cannot invert: a good track record does not let this submission skip its evidence, and all-green automated checks do not make the business judgment correct.
Good filtering is not a machine deciding for a person. It is a machine clearing mechanical problems, organizing evidence, and flagging risk, so the human judgment goes where it cannot be replaced.
Why trust becomes a means of production
Trust is usually read as brand, fame, or a vague feeling. Placed inside a production chain, it turns out to be a set of accumulable evidence:
- Does this result reproduce?
- Are the sources, assumptions, and non-conclusions stated?
- Is the change history traceable?
- When it broke, was there an owner, a rollback path, and a correction record?
- Does this submitter's past work still hold up on re-examination?
That evidence lowers the verification cost of the next collaboration. A maintainer reads a particular researcher first not because verification became unnecessary, but because that person has repeatedly delivered reproducible, correctly scoped, clearly communicated results. A team hands agents more work not because a model became unconditionally trustworthy, but because tests, permissions, review, and rollback in the repository make errors visible, stoppable, and recoverable.
So trust is not a substitute for verification. It is the compressed result of a long verification record.
GitHub matters here, but GitHub does not manufacture trust. Commits, PRs, reviews, checks, releases, and issues are containers. The value is the evidence left inside them: who changed what, why, which checks passed, what remains uncertain, and how it was corrected when it broke.
Who this is useful to
The mechanism reaches past developers.
If you use Claude Code
Do not only optimize for how much it writes per run. Define acceptance criteria first, and have the agent produce reproduction steps, tests, and a risk note alongside the patch. Automate the high-frequency low-risk checks; keep releases, deletions, permissions, and outbound messages as human control points.
If you maintain an open-source project
An open entrance does not require an undifferentiated queue. Demand a minimal reproduction, a scope statement, and passing automated checks. Track high-quality contributors, but never let reputation bypass evidence. When measuring automation, count how much invalid review it removed, not only how many submissions it added.
For content and research teams
Separate raw facts, mechanism judgments, who it applies to, and what does not follow. AI can produce summaries and candidate arguments; editors should prioritize sources, counter-examples, omitted conditions, and downstream consequences. Volume published is not the output. Judgments that still hold after verification are.
For recruiters and buyers
When résumés, portfolios, and proposals can all be generated in bulk, polished text loses discriminating power. What gains value: traceable work, stated boundaries, evidence of process, a maintenance record over time, and how a candidate handles having been wrong once.
What this does not imply
"Verification, filtering, and trust are worth more" is easy to abuse into new bureaucracy, so the limits need stating.
It does not imply that more gates are safer. An approval with no matching risk only manufactures waiting.
It does not imply that only veterans and famous names deserve trust. If historical reputation becomes the only entrance, a newcomer can never accumulate a first credible record. The system still needs a lane for low-history, high-evidence candidates.
It does not imply that generation has no value. Precisely because generation is cheap, reverse engineering, automation, and small tools that were not worth attempting have become viable. The ROI shift Willison describes is a real gain.
And it does not imply that a benchmark or a vendor experiment extrapolates to your production environment. The Anthropic and Cursor material establishes direction; performance, cost, and stability still have to be verified against your own tasks, data, and failure conditions.
A working method for the cheap-generation era
Six rules I would keep:
- Write the acceptance criteria before generating. Without them, more candidates only make judgment harder.
- Deliver the result with its evidence. Code with tests and a reproduction, analysis with sources and counter-examples, data with its definition and cut-off date.
- Automate the mechanical filter first. Format, schema, scope, duplicates, and build should never consume expert attention.
- Allocate human judgment by risk. Releases, permissions, security, legal, financial, and outbound communication get human confirmation first.
- Record misjudgments and corrections. Showing only successes manufactures false trust; a reliable system has to allow reversals.
- Measure pass rate and survival rate. Not only how much was generated, but how much passed verification and still held after shipping.
None of this makes review disappear. It converts review from undifferentiated human reading into a production process with evidence, ranking, and boundaries.
Conclusion
In AI leverage and the bread paradox I read "trustworthy delivery" as the part beyond the prototype that someone actually pays for. Claude Code, agent swarms, and GitHub's bounty restructuring make that concrete: trustworthy delivery is not a service promise. It is built out of verification, filtering, and a long traceable trust record.
Generation tools will keep getting cheaper. Candidate patches, vulnerability reports, content drafts, and business proposals will keep multiplying. What limits throughput next is not who can generate one more, but who can answer:
- Which of these results is real?
- Which deserves to be read first?
- Which can carry production consequences?
- If the judgment was wrong, who notices, stops it, and corrects it?
The people and systems that answer those reliably hold the means of production in an era of cheap generation.
Comments
Comments are powered by GitHub Discussions. Sign in with GitHub to comment. Open the matching Discussion