- NO.
- 022
- DATE
- READ
- ~10 min
- KIND
- Notes
- STATUS
- Reviewed
What "Accountable for Results" Actually Means
Not a guarantee of correctness, and not "I tried hard." Four levels of evidence behind a claim, and the hard boundaries you should not cross alone.
"Be accountable for the result." Everyone agrees with the phrase and almost nobody defines it.
It cannot mean "guarantee the result is correct." Tests only verify the scenarios you thought of, monitoring only sees the metrics you chose to record, and even an expert cannot fully verify a sufficiently complex system. If accountability meant guaranteed correctness, no accountable person would exist.
But it cannot mean "I did my best" either. Effort is an attitude, not a standard, and it cannot be checked.
The real answer hides in a distinction most people never make.
1. Which level is your claim standing on?
For anything you deliver, you can make four claims of increasing strength:
Level 1: it ran. No errors, the page loaded. The floor.
Level 2: it passed my tests. In the scenarios I thought of, it behaves as expected.
Level 3: it met a real user's need. Not just yours — someone who actually uses it confirmed it is the right thing.
Level 4: it deserves to be relied on by real users. It handles irreversible real data, affects real people, and has held up under that weight.
Most arguments about someone being "irresponsible" do not start with a wrong result. They start with someone whose evidence supports level 1 while their words claim level 4.
That is where false precision comes from. "Response time must be under two seconds." "Test coverage must reach 90%." "Error rate must stay below 0.1%." If those numbers do not come from real users, an existing baseline, or a technical measurement, they are a level-4 building on a level-1 foundation.
So the first step toward accountability is not verifying harder. It is being able to state three things at any moment: which level does my evidence support, which level am I claiming out loud, and do those two match?
That is not modesty. It is calibration.
2. Unclear requirements are normal; faking precision is the trap
"Define the acceptance criteria clearly, then let the AI build it" carries a hidden premise: that whoever writes the requirement already understands the business logic, the technical limits, the exception paths, and the performance constraints.
Experienced developers cannot do that in one pass either. Someone new to programming knows only "I want to import a CSV." They typically will not think of: is the file encoding consistent? What if one row is malformed? What about duplicates? What if the import is interrupted halfway? Could it overwrite existing data? How does the user learn which records failed?
Those questions do not arrive by thinking harder. They come from experience, tests, user feedback, and incidents.
So "the definition is incomplete" is not a reason to stop. Acceptance criteria were never written once; they are discovered. You start with "the page must not visibly freeze during import," find through testing that 500 rows are fine while 5,000 hang, and gradually convert a vague feeling into a measurable threshold. The sane process is not "perfect definition → AI implements → final acceptance." It is "vague goal → minimal implementation → observe problems → add criteria → implement again → verify again."
The danger is not vagueness. It is faking precision while vague. A beginner should learn to separate four kinds of requirement:
- Known requirement: it genuinely must not overwrite existing data.
- Provisional requirement: let's try to support 5,000 rows.
- Open question: I do not yet know how many rows users typically import.
- Hypothesis to test: batching may reduce the freeze.
A vague but honest description is always more trustworthy than a precise number with nothing behind it.
3. Let ignorance hit a wall, rather than demanding self-knowledge
"You must know what you don't know" is correct, and executable only within limits.
For known categories of risk, self-knowledge works: you know you have never written a payment system, you know you do not understand database permission models, you know this is your first time handling private user data. Those known unknowns should be faced directly; they need no mechanism to surface them.
For genuine unknowns — problems you do not know exist — self-knowledge fails. You cannot enumerate a list that is not in your head. Treating "I must know what I don't know" as the only goal produces anxiety, not action.
What you need is a different response for each: for known ignorance, admit it and get help; for unknown ignorance, build channels that let problems surface on their own.
There are four channels.
Let the machine enumerate. Have the AI list exception scenarios, edge conditions, and risks. The list will not be complete, but it converts "things you never considered" into a candidate list you can review item by item. Do not let it start writing code first. Ask: "Given this requirement, list the normal flow, the exception flows, data risks, performance risks, and suggested test methods. Do not modify code; explain it to me first."
Let the world test it. Use real samples, not AI-generated ones. The world does not cooperate with your assumptions. Three files from real situations are more honest than ten fabricated test fixtures.
Let time stretch it. Stage the rollout, ramp a small percentage, run against a copy first. Break one big irreversible bet into several small correctable ones.
Let someone else see it. A second model reviewing, a knowledgeable friend checking, a real user trying it. Your blind spots do not vanish, but another person's are in different places.
The principle is not "I know everything." It is "my ignorance will hit a wall." A wall is more reliable than self-awareness.
4. Start from five questions
If you do not know how to define a requirement, answer these five. They do not need perfect answers, but they need honest ones.
What do I want to see in the normal case? Describe one concrete example: I select a CSV containing names and emails, click import, the page shows how many succeeded, and the new records appear in the contact list. That is the baseline happy path.
What must absolutely never happen? It must not delete existing contacts. It must not upload keys or passwords to an external service. It must not discard the whole file because one row failed. It must not overwrite data without warning. Prohibitions are always easier to define than performance numbers — and more important.
What happens when the input is abnormal? Empty file, missing column, duplicate records, encoding errors, oversized file, interrupted import. You can have the AI produce the candidate list, but you decide which cases are worth handling. AI can widen your field of view; it cannot decide which risks matter to you.
How will I prove it works? Test with three different real samples. Compare record counts before and after. Deliberately insert one bad row. Close the page mid-import. Read the git diff. Check browser logs and network requests. Compare against the old version's result. "I ran it once and nothing errored" is evidence — level-1 evidence. Do not dress it up as adequate verification.
If I got it wrong, can I recover? Commit before changing. Back up before touching the database. Test on a small dataset. Never run against the only copy of production data. Keep the old version. Write down the rollback steps. Being wrong is not shameful; being wrong with no way back is the dangerous part.
5. The examiner and the grader cannot be the same source
If one AI understands your requirement, writes the code, designs the tests, and then explains why the tests pass, the code and the tests very likely share a blind spot.
The fix is not "a second AI will be more accurate" — it can replicate the same bias. The fix is to make the witness a different source from the examiner:
- Test samples come from the real world, not from the AI.
- The comparison is against the old version or a manual result, not the AI's recollection.
- A real user tries it, rather than the author demonstrating it.
- For payments, permissions, privacy, and deletion, get review from someone experienced.
An independent witness does not need to be perfect. It needs to be independent. One independent channel for discovering errors beats ten checks from the same source.
6. Hard boundaries: what you should not do alone
Accountability includes knowing when not to carry something alone.
You can use AI for local tools, personal sites, disposable experiments, and data-processing scripts with backups.
You should not ship the following on the strength of "it ran," alone:
- Payment systems.
- Permission and authentication systems.
- Mass deletion or data migration programs.
- Processing of private user data.
- Medical, legal, or financial decision systems.
- Irreversible operations on a production database.
Not because you did not try hard enough. In these domains the distance between level-1 and level-4 evidence — domain knowledge, compliance requirements, the sheer number of edge conditions — is not something personal care plus AI assistance can cross.
Sometimes the responsible choice is to reduce scope, find expert help, or not build it. Knowing when not to is itself accountability.
7. What responsibility really is: letting others check you
The most overlooked point.
Responsibility is not only a personal quality — "I am careful, I am thorough." Its original sense is responsible to someone. It is a relational structure, not an internal state.
Concretely: git history lets people see what you changed. Assumptions written in comments let people see what you were thinking. A verification script lets people reproduce your test. A sentence saying "I am not sure this part is reliable" tells whoever comes next where to look harder.
You cannot be accountable for a result nobody can see.
So the most underrated skill for a beginner is not being more careful or being smarter. It is exposing yourself to review. Show the thing you just made to someone. Write down how you verified it. Say "I am not sure" honestly. Those acts get closer to the substance of responsibility than any individual technique.
Some will say that since results cannot be guaranteed, you can only be accountable for decisions. True as far as it goes, and it is not a disclaimer. If the mistake was a signpost you could have read in three minutes and ignored, that is not "the result was out of my control." That is a faulty decision, and it should be examined. Auditable decisions make accountability easier, not harder — because there is a list to check item by item instead of a vague "the outcome was bad."
8. An honest ending
Even after all of the above, you may still miss the thing that actually mattered.
Tests only cover what you thought of. Monitoring only records what you chose. A second model may repeat the first one's bias.
There may be no cheap, universal, complete verification method that lets someone without domain knowledge safely operate arbitrarily complex systems. AI lowered generation cost, tooling lowered part of the verification cost, but domain knowledge, risk judgment, and real-world feedback remain scarce. AI pushed the capability boundary outward; it did not remove it.
I wrote the other side of this in cheap generation, expensive verification: at the organizational level, the scarce capacity is verification and filtering. This post is the same constraint landing on one person — the scarce thing is the evidence behind your sentence.
So the final advice is one line: match the project's risk, scope, and irreversibility to your current level of understanding.
And progress is not jumping from "it ran" straight to "I can prove it is correct" — three levels of evidence sit in between. Progress is being able to say which level you are standing on, what evidence you paid for the next one, and whether the claim leaving your mouth matches the evidence in your hand.
Those four things — it ran, it passed my tests, it met one person's need, it deserves to be relied on — should never again be treated as the same sentence.
Comments
Comments are powered by GitHub Discussions. Sign in with GitHub to comment. Open the matching Discussion