- NO.
- 028
- DATE
- READ
- ~4 min
- KIND
- Notes
- STATUS
- Reviewed
When the Model Becomes a Swappable Part
If models get easier to swap, what is actually worth accumulating in an agent? The model supplies capability; the harness makes it engineerable.
I have been looking at DeepSeek Harness.
My first reaction was not "another coding agent." It was a different question:
If models keep getting easier to swap, what in an agent is actually worth accumulating over the long run?
For a while now the attention has gone to model capability. Which one codes better, which has the longer context, which is cheaper.
Those matter.
But once a model enters a real engineering environment, the outcome stops being decided by the model alone. It also depends on:
- Which tools it can reach.
- Which project rules it knows.
- How it persists execution state.
- When it has to ask a human.
- Whether it can recover and be traced after a mistake.
Those are exactly the concerns of a harness.
An agent is not just a model
A simple way to hold it:
Agent = Model + Harness
The model does the reasoning.
The harness lets that reasoning reach the real world.
Files, shell, git, tests, permissions, sandbox, session, skills, subagents — none of these live in the model weights. They are the runtime built around the model.
Which is why the same model can behave completely differently inside two agent products. The difference does not necessarily come from the model. It comes from the engineering layer wrapped around it.
The model may be becoming a swappable part
As the harness layer matures, a workflow stops being bound to one model product. It looks more like this:
My agent environment
|
+-- Model A
+-- Model B
+-- Model C
+-- Tools
+-- Skills
+-- Sandbox
+-- Workflow
+-- Verification
The model supplies capability. Project rules, toolchain, verification flow, and permission boundaries become the assets that accumulate.
That is close to what I recorded in building a site with an AI agent: the toolchain behind this bilingual blog. What made me willing to let an agent modify a long-lived project was not that the model got clever enough. It was that AGENTS.md, the schema, the publishing flow, and the verification boundaries existed first.
Verification keeps getting more important
Once generation gets cheap, the bottleneck stops being "can we produce a result" and becomes "how do we decide whether this result deserves to enter the system."
I wrote about that in when generation got cheap, verification and trust became the means of production. Agents intensify the trend rather than resolving it.
A mature agent system needs more than generation:
generate
↓
check
↓
test
↓
review
↓
ship
Without those stages, a faster agent simply pushes errors into production faster.
Traceability is easier to undervalue than intelligence
The part of DeepSeek Harness I keep returning to is how it handles sessions and execution traces.
The worst problem with a complex agent is not occasional failure. It is being unable to answer, after a failure:
- Where did it start going off?
- Which tool call caused it?
- Would another model reproduce this?
- What is the sane point to resume from?
Traditional software needs logs and tracing. A long-running agent needs them for the same reasons.
If agents move from chat tools to infrastructure, execution traces become as important as system logs.
Permission boundaries do not disappear as models improve
Plenty of people expect agents to eventually work fully autonomously.
But the stronger the capability, the more the permission design matters. The target state is not:
The agent is smart, so give it every permission.
It is:
read project allow
modify workspace allow
delete resources confirm
production deploy confirm
sensitive data least privilege
Which is why sandboxing, approvals, and policy — none of them glamorous — keep gaining weight.
Writing hooks for coding agents: from rule to gate recently left me with the same impression: a reliable agent is not one without limits. It is one where the limits are part of the system.
But a harness does not make a model smarter
A boundary worth stating.
A harness improves context management, tool calling, execution flow, permission control, and observability.
It does not raise the model's capability ceiling. A question the model does not understand will not become understood by adding plugins.
So I prefer to read a harness as:
wrapping unstable model capability into a manageable engineering system.
My read
I am not migrating any workflow because DeepSeek Harness exists. It still looks like an infrastructure direction worth watching and experimenting with, not a mature replacement.
But the question it raises matters: in long-term human–AI collaboration, what is the thing that actually accumulates?
Probably not a model. Models change.
What stays is more likely to be project knowledge, tool connections, verification rules, permission boundaries, recorded failures, and workflow. Together, those are an agent's harness.
If that holds, the competition is not only about who has the stronger model. It is also about who has the better system for making a model work reliably.
Comments
Comments are powered by GitHub Discussions. Sign in with GitHub to comment. Open the matching Discussion