The tool said it worked.
It hadn’t.
One real run. One check. The check disagreed with the tool, and the check was right.
What happened, in plain words
I asked my own tooling to do one thing: write a note down locally. Capture only. Nothing was supposed to leave my machine.
The tool reported success. That was true in the only place the tool could see: its own return value. Outside it, issue #4 had been created in this public repository. I never authorized that.
The check that caught it doesn’t read the tool’s report. It asks GitHub directly whether the artifact exists. The answer came back yes, the expected answer was the opposite, and the run was marked failed. That comparison is the entire mechanism: the value that should be there against the value that is.
The status band below calls this a failed postcondition. Plain version: what the run promised about the world afterwards wasn’t true. The log shows the four steps in order.
The recorded evidence
tool: tool reported success · postcondition: FAIL · resource: issue
The gate caught it
The same four steps as the tooling recorded them:
ASSERT · DEMONSTRATED · postcondition FAIL
Controls that are (and are not) demonstrated
Compare the external world against the expected business fact.
demonstratedSeparately evidenced by the approved retry and checkpoint bundle.
demonstratedAuthorization evidence is separately approved for one controlled operation.
demonstratedBOUND: retry and checkpoint evidence
Expected retry limit: 2. Observed stop: retry_limit_reached.
Checkpoint: checkpoint.json. Resume reads that checkpoint and continues from resume_attempt_3; new external work on resume: 0.
Replay: node scenario.mjs ./run-dir; load run-dir/checkpoint.json and continue from next_step
Remaining unproven
- production provider delivery semantics
- human approval gate
- cost savings outside this recorded run
APPROVE: authorization gate evidence
Without authorization, the attempted write was blocked and adapter calls stayed at zero. The recorded authorization then allowed exactly one controlled operation.
Replay reads the checkpoint and performs zero new operations: node scenario.mjs ./run-dir; read checkpoint.json and do not repeat completed operation
Remaining unproven
- real provider side effects
- multi-operator authorization
- production deployment semantics
Replay and review
Replay uses recorded/redacted GitHub Issues API responses or a throwaway repository. It never writes to the observed repository. Issue #4 is a live, read-only provenance link. It is not replay input and it is not a screenshot.
Open live issue #4 evidenceWhat still breaks
- Issue #4 is live link-only provenance, never replay input or screenshot.
- No reliability, ROI, cost-saving, or production-wide claim follows from this proof.
Why a failure this small matters
Nothing broke here. A note became an issue in my own repository, I found it in minutes, and it cost nothing to undo.
That is the reason to publish it. This is the same shape as an agent that reports a clean send while putting the wrong price in nine thousand inboxes: quiet, confident, and wrong. The failure never announces itself, so something other than the tool’s own report has to do the checking. I would rather meet the shape at this size.