Back to blog
NO.
028
DATE
READ
~4 min
KIND
Notes
STATUS
Reviewed

TAGS: AI collaboration Architecture

When the Model Becomes a Swappable Part

If models get easier to swap, what is actually worth accumulating in an agent? The model supplies capability; the harness makes it engineerable.

I have been looking at DeepSeek Harness.

My first reaction was not "another coding agent." It was a different question:

If models keep getting easier to swap, what in an agent is actually worth accumulating over the long run?

For a while now the attention has gone to model capability. Which one codes better, which has the longer context, which is cheaper.

Those matter.

But once a model enters a real engineering environment, the outcome stops being decided by the model alone. It also depends on:

  • Which tools it can reach.
  • Which project rules it knows.
  • How it persists execution state.
  • When it has to ask a human.
  • Whether it can recover and be traced after a mistake.

Those are exactly the concerns of a harness.

An agent is not just a model

A simple way to hold it:

Agent = Model + Harness

The model does the reasoning.

The harness lets that reasoning reach the real world.

Files, shell, git, tests, permissions, sandbox, session, skills, subagents — none of these live in the model weights. They are the runtime built around the model.

Which is why the same model can behave completely differently inside two agent products. The difference does not necessarily come from the model. It comes from the engineering layer wrapped around it.

The model may be becoming a swappable part

As the harness layer matures, a workflow stops being bound to one model product. It looks more like this:

My agent environment
        |
        +-- Model A
        +-- Model B
        +-- Model C

        +-- Tools
        +-- Skills
        +-- Sandbox
        +-- Workflow
        +-- Verification

The model supplies capability. Project rules, toolchain, verification flow, and permission boundaries become the assets that accumulate.

That is close to what I recorded in building a site with an AI agent: the toolchain behind this bilingual blog. What made me willing to let an agent modify a long-lived project was not that the model got clever enough. It was that AGENTS.md, the schema, the publishing flow, and the verification boundaries existed first.

Verification keeps getting more important

Once generation gets cheap, the bottleneck stops being "can we produce a result" and becomes "how do we decide whether this result deserves to enter the system."

I wrote about that in when generation got cheap, verification and trust became the means of production. Agents intensify the trend rather than resolving it.

A mature agent system needs more than generation:

generate
 ↓
check
 ↓
test
 ↓
review
 ↓
ship

Without those stages, a faster agent simply pushes errors into production faster.

Traceability is easier to undervalue than intelligence

The part of DeepSeek Harness I keep returning to is how it handles sessions and execution traces.

The worst problem with a complex agent is not occasional failure. It is being unable to answer, after a failure:

  • Where did it start going off?
  • Which tool call caused it?
  • Would another model reproduce this?
  • What is the sane point to resume from?

Traditional software needs logs and tracing. A long-running agent needs them for the same reasons.

If agents move from chat tools to infrastructure, execution traces become as important as system logs.

Permission boundaries do not disappear as models improve

Plenty of people expect agents to eventually work fully autonomously.

But the stronger the capability, the more the permission design matters. The target state is not:

The agent is smart, so give it every permission.

It is:

read project        allow
modify workspace    allow
delete resources    confirm
production deploy   confirm
sensitive data      least privilege

Which is why sandboxing, approvals, and policy — none of them glamorous — keep gaining weight.

Writing hooks for coding agents: from rule to gate recently left me with the same impression: a reliable agent is not one without limits. It is one where the limits are part of the system.

But a harness does not make a model smarter

A boundary worth stating.

A harness improves context management, tool calling, execution flow, permission control, and observability.

It does not raise the model's capability ceiling. A question the model does not understand will not become understood by adding plugins.

So I prefer to read a harness as:

wrapping unstable model capability into a manageable engineering system.

My read

I am not migrating any workflow because DeepSeek Harness exists. It still looks like an infrastructure direction worth watching and experimenting with, not a mature replacement.

But the question it raises matters: in long-term human–AI collaboration, what is the thing that actually accumulates?

Probably not a model. Models change.

What stays is more likely to be project knowledge, tool connections, verification rules, permission boundaries, recorded failures, and workflow. Together, those are an agent's harness.

If that holds, the competition is not only about who has the stronger model. It is also about who has the better system for making a model work reliably.

Comments →

CC BY-NC-SA 4.0

Comments

Comments are powered by GitHub Discussions. Sign in with GitHub to comment. Open the matching Discussion