“All tests pass. The feature is complete.”
Every coding agent says this. Most of the time it’s even true, in the narrowest sense. The tests pass. The tests the agent wrote. For the code the agent wrote. Based on the agent’s reading of a ticket that said “make the export faster”.
That’s not proof. That’s a student grading their own exam and handing you the score.
Where “done” goes wrong
I ran an AI team on a real product for ten weeks before I rebuilt kanman around what I learned. The failures that hurt were never syntax errors. CI catches those.
The ones that hurt looked like this:
- The agent mocked the exact thing the test was supposed to check. Green run, nothing tested.
- The agent “fixed” a failing test by changing the assertion.
- The feature worked on the agent’s machine, with state left over from three earlier attempts, and nowhere else.
- The PR did what the ticket said, and not what the person who wrote the ticket meant.
Each one passed review at least once, because reviewers trust green checks. Of course they do. That’s what green checks are for.
The outcome gate
So kanman doesn’t ask the agent whether it’s done. The gate does, and the gate is not an agent.
It works in four moves.
1. The spec comes first. Every story kanman writes during intake gets a Demonstrate block: a concrete, runnable way to show the outcome. Not “unit tests pass”, but “export 10,000 rows and the download starts within two seconds”. That block becomes an acceptance spec, before anyone writes a line of feature code.
2. Red before Ready. A story only moves to Ready when its spec exists, runs, and fails. A spec that passes before the work starts tests nothing. A spec that mocks its own target is rejected. No spec, no Ready.
3. The agent can’t touch the spec. The spec lives on a protected path. If a commit from the coding agent changes it, the gate stops the story. The agent can make the spec pass. It can’t make the spec easier.
4. Green on a clean room, or no review. When the agent says it’s done, the spec runs again on a fresh checkout, against a freshly provisioned environment. No leftover state, no “works on my machine”. If it’s red, the story goes back to work. If it’s green, the pull request opens. No green run, no review.
Then the pull request comes with an evidence pack: each acceptance criterion with its result, the spec run, the trace and screenshots from the run, which model did the work, and what it cost. Your reviewer reads evidence instead of reconstructing it.
“Our repo has no acceptance tests”
Most repos I see don’t. That’s fine.
You don’t need a test suite to start, you need a manifest: a small .kanman/acceptance.json that tells kanman how to start an environment and how to run a spec. If your repo doesn’t have one, kanman opens a pull request that adds it, and you review it like any other change. The acceptance manifest guide walks through it.
Until then, there is the Trial preset. It relaxes the spec requirement for the first pilot week and falls back to CI, a test plan and an independent reviewer run. And the evidence pack says so, in plain words: no outcome proof exists for this change. I’d rather tell you that than dress up a green CI badge as proof.
Why this is the part that matters
There’s a reason this sits in the middle of kanman and not in a settings page.
An AI teammate that ships unverified work doesn’t save your team time. It moves the work from writing code to doubting code. Your senior people end up reviewing every line twice, because they can’t tell which PRs are real.
An AI teammate that shows evidence changes the review. You stop asking “does this work?” and start asking “is this what we wanted?”. That second question is the one humans are good at.
Agents that prove their work beat agents that demo. If you want the details, the outcome gate is documented step by step.
Want an AI teammate that proves its work before it asks for your review? kanman - Pilot from €5,000 for 6 weeks, Team €990 per AI team per month.
Marco Kerwitz
Founder of kanman.ai