Tafadzwa Mukungurutse
PROOF / ONE RUN, ONE CHECK

The tool said it worked.
It hadn’t.

One real run. One check. The check disagreed with the tool, and the check was right.

What happened, in plain words

I asked my own tooling to do one thing: write a note down locally. Capture only. Nothing was supposed to leave my machine.

The tool reported success. That was true in the only place the tool could see: its own return value. Outside it, issue #4 had been created in this public repository. I never authorized that.

The check that caught it doesn’t read the tool’s report. It asks GitHub directly whether the artifact exists. The answer came back yes, the expected answer was the opposite, and the run was marked failed. That comparison is the entire mechanism: the value that should be there against the value that is.

The status band below calls this a failed postcondition. Plain version: what the run promised about the world afterwards wasn’t true. The log shows the four steps in order.

The recorded evidence

RUNdevos.capture.capture_writecap-20260825-745a
ASSERTExpected: external artifact absentexternal artifact created without explicit authorization
OBSERVEDIssue #4 existsLive issue evidence

tool: tool reported success · postcondition: FAIL · resource: issue

The gate caught it

One request went in. Two answers came back, and they did not match.
01
I asked for one thingcapture-only
02
What the tool saidtool reported successread from its own return value
03
What the world saidissue #4 existsread by asking GitHub directly
04
They disagreed, so the run failedpostcondition FAILwhat the run promised about the world afterwards was not true

The same four steps as the tooling recorded them:

01 operator_intentcapture-only
02 capture_writetool reported success
03 source_of_truthissue #4 exists
04 assertionpostcondition FAIL

ASSERT · DEMONSTRATED · postcondition FAIL

Controls that are (and are not) demonstrated

01 ASSERT

Compare the external world against the expected business fact.

demonstrated
02 BOUND

Separately evidenced by the approved retry and checkpoint bundle.

demonstrated
03 APPROVE

Authorization evidence is separately approved for one controlled operation.

demonstrated

BOUND: retry and checkpoint evidence

Expected retry limit: 2. Observed stop: retry_limit_reached.

Checkpoint: checkpoint.json. Resume reads that checkpoint and continues from resume_attempt_3; new external work on resume: 0.

Replay: node scenario.mjs ./run-dir; load run-dir/checkpoint.json and continue from next_step

Remaining unproven

  • production provider delivery semantics
  • human approval gate
  • cost savings outside this recorded run

APPROVE: authorization gate evidence

Without authorization, the attempted write was blocked and adapter calls stayed at zero. The recorded authorization then allowed exactly one controlled operation.

Replay reads the checkpoint and performs zero new operations: node scenario.mjs ./run-dir; read checkpoint.json and do not repeat completed operation

Remaining unproven

  • real provider side effects
  • multi-operator authorization
  • production deployment semantics

Replay and review

Replay uses recorded/redacted GitHub Issues API responses or a throwaway repository. It never writes to the observed repository. Issue #4 is a live, read-only provenance link. It is not replay input and it is not a screenshot.

Source revision: sha256:ae03f04deefdfdcc45e3afbaae2967c88c829eda753bbc91dde589b7e720ea96
Evidence digest: dee62b9e4307585a545f13c9842be6960e7a46f7ab236b398a1c2efbaef87ba9
Reviewer: Tafadzwa · decision: approved

Open live issue #4 evidence

What still breaks

  • Issue #4 is live link-only provenance, never replay input or screenshot.
  • No reliability, ROI, cost-saving, or production-wide claim follows from this proof.

Why a failure this small matters

Nothing broke here. A note became an issue in my own repository, I found it in minutes, and it cost nothing to undo.

That is the reason to publish it. This is the same shape as an agent that reports a clean send while putting the wrong price in nine thousand inboxes: quiet, confident, and wrong. The failure never announces itself, so something other than the tool’s own report has to do the checking. I would rather meet the shape at this size.