Skip to content
Blog

Don't Trust an Agent That Says Done

Oct 1, 2026 4 min read ai, engineering, testing

Coding agents grade their own homework. kanman doesn't let them: an acceptance spec is written before the code, must fail before work starts, can't be edited by the agent, and has to pass on a clean checkout before anyone is asked to review.

“All tests pass. The feature is complete.”

Every coding agent says this. Most of the time it’s even true, in the narrowest sense. The tests pass. The tests the agent wrote. For the code the agent wrote. Based on the agent’s reading of a ticket that said “make the export faster”.

That’s not proof. That’s a student grading their own exam and handing you the score.

Where “done” goes wrong

I ran an AI team on a real product for ten weeks before I rebuilt kanman around what I learned. The failures that hurt were never syntax errors. CI catches those.

The ones that hurt looked like this:

  • The agent mocked the exact thing the test was supposed to check. Green run, nothing tested.
  • The agent “fixed” a failing test by changing the assertion.
  • The feature worked on the agent’s machine, with state left over from three earlier attempts, and nowhere else.
  • The PR did what the ticket said, and not what the person who wrote the ticket meant.

Each one passed review at least once, because reviewers trust green checks. Of course they do. That’s what green checks are for.

The outcome gate

So kanman doesn’t ask the agent whether it’s done. The gate does, and the gate is not an agent.

It works in four moves.

1. The spec comes first. Every story kanman writes during intake gets a Demonstrate block: a concrete, runnable way to show the outcome. Not “unit tests pass”, but “export 10,000 rows and the download starts within two seconds”. That block becomes an acceptance spec, before anyone writes a line of feature code.

2. Red before Ready. A story only moves to Ready when its spec exists, runs, and fails. A spec that passes before the work starts tests nothing. A spec that mocks its own target is rejected. No spec, no Ready.

3. The agent can’t touch the spec. The spec lives on a protected path. If a commit from the coding agent changes it, the gate stops the story. The agent can make the spec pass. It can’t make the spec easier.

4. Green on a clean room, or no review. When the agent says it’s done, the spec runs again on a fresh checkout, against a freshly provisioned environment. No leftover state, no “works on my machine”. If it’s red, the story goes back to work. If it’s green, the pull request opens. No green run, no review.

Then the pull request comes with an evidence pack: each acceptance criterion with its result, the spec run, the trace and screenshots from the run, which model did the work, and what it cost. Your reviewer reads evidence instead of reconstructing it.

“Our repo has no acceptance tests”

Most repos I see don’t. That’s fine.

You don’t need a test suite to start, you need a manifest: a small .kanman/acceptance.json that tells kanman how to start an environment and how to run a spec. If your repo doesn’t have one, kanman opens a pull request that adds it, and you review it like any other change. The acceptance manifest guide walks through it.

Until then, there is the Trial preset. It relaxes the spec requirement for the first pilot week and falls back to CI, a test plan and an independent reviewer run. And the evidence pack says so, in plain words: no outcome proof exists for this change. I’d rather tell you that than dress up a green CI badge as proof.

Why this is the part that matters

There’s a reason this sits in the middle of kanman and not in a settings page.

An AI teammate that ships unverified work doesn’t save your team time. It moves the work from writing code to doubting code. Your senior people end up reviewing every line twice, because they can’t tell which PRs are real.

An AI teammate that shows evidence changes the review. You stop asking “does this work?” and start asking “is this what we wanted?”. That second question is the one humans are good at.

Agents that prove their work beat agents that demo. If you want the details, the outcome gate is documented step by step.

Want an AI teammate that proves its work before it asks for your review? kanman - Pilot from €5,000 for 6 weeks, Team €990 per AI team per month.
Marco Kerwitz
Author

Marco Kerwitz

Founder of kanman.ai

Meet kanman

kanman is an AI teammate for engineering teams. It takes requirements, writes the stories, ships the code through Claude Code or Codex and proves each change against your acceptance criteria.

  • Works inside your Jira, GitHub or GitLab.
  • Your policy decides what it may touch, spend and merge.
  • Every pull request comes with an evidence pack.
Book a pilot

Pilot from €5,000 for 6 weeks, Team €990 per AI team per month.