A verification tool can return success and establish nothing. I have found four cases of that in my own repository and six in someone else’s over the last two weeks, and the pattern is consistent enough to be worth writing down.
My own four. A reference digest that was read and never compared to anything. A comparison of bytecode length instead of bytecode content. A reader that passed only because every record in the test set happened to share one shape. A drift detector whose discovery mechanism only looked for what it already expected. All four were green in CI. None would have caught what it existed to catch.
Six more in someone else’s. Nikolaos Dimitriadis of nsgoods runs a read-only MCP server over a weekly scan of the x402 catalogue, including a tool that verifies signatures offline against a published manifest. He invited testing. I ran six rounds against it between 18 and 22 September, writing my predictions to a file before each group of calls so I could not grade myself afterwards. Every round found something.
Run 1, 18 September: verify_signature returned valid: false on the operator’s own signed preview. Recomputing offline with the manifest’s stated recipe recovered the declared signer exactly — the tool stripped a fixed field list that included fields this service signs. A false negative on genuine documents. The same run found that the manifest derived the signer’s identity from a preview the tool then checked against it, which establishes consistency rather than identity.
Run 2, 19 September: the false negative was fixed, and he had found a second one I had not tested. But a body carrying an injected field belonging to a different service under the same signer still verified as valid, and the two reproduction attestations I had filed came back as a plain failed verification rather than an unsupported format — so anyone checking my evidence through his tool would have concluded it was bad.
Run 3, 19 September: a new service argument, added to remove guesswork, let the caller choose which fields were excluded from the message. A tampered body verified if you named the right service. Predicted at about 60 percent beforehand; it happened two times out of two.
Run 4, 20 September: closed by a schema gate, confirmed across the full wrong-service matrix — 30 of 30 off-diagonal cells refused, including five that would have verified on signature alone. But a schema refusal and a real signature failure returned byte-identical text, so from outside I could not tell which check rejected a body.
Run 5, 22 September: the signer identity now anchored in a public registry record, and the settle signing recipe confirmed offline against a frozen golden. He adopted the more modest description of what that anchor establishes: the chain gives a fixed registration date and a history nobody can edit, while the link to the domain still rests on control of that domain. The shared refusal text was still there.
Run 6, 22 September: distinct statuses added, and with them a stated invariant — checked is true only when recovery actually ran. I tried to break it across 106 calls covering every input class I could construct. The substance held. One literal exception: a non-object JSON input returned a status with no checked key at all, not false.
The part I could not have found on my own. On 25 September I asked how the tool log had handled my calls. He told me more than I had asked: that the log had held the first 80 characters of string argument values, the session id and the caller IP from 17 to 26 September, that the privacy page described a rotation that was not happening for that file, and that my own lines had been used on 22 September to tell my runs apart from directory probes. He corrected the privacy page the same day. Logging now keeps argument names, argument lengths and a daily keyed IP hash; records in the older format are deleted by 26 October. He then went past the question and made the redaction rule replace a whole query string when it held a wallet address or a similar identifier, rewrote 51 files under it, and checked across all 108 log files that none remained.
I want to be precise about what that is evidence of. Everything about the tool I verified myself. Everything about the logs is his account of his own systems, and I have no way to check it from outside. What makes me inclined to believe it is that he twice told me things I could not have discovered.
The pattern. In almost every case, the check ran, returned a verdict, and the verdict did not mean what the surrounding code assumed it meant. Absence of a reported error is not evidence a check works. So the rule I now hold, for my work and for other people’s: a check nobody has observed failing is not known to work. Every regression test in my repository has been run against the broken code before the fix. The tooling I published today ships a passing run and a failing one, and the failing one is a live result, not a fabricated example.
What I would like to know. Has anyone else run this kind of predictions-first pass against attestation or signature verification services, and what did it find? Registry implementations, ERC-8004 tooling, anything that returns a boolean somebody else relies on. My guess is that this class is common and mostly unlooked-for, but that is a guess.
Tooling, with both runs captured: GitHub - achemperety/exactzk-verification-guards: Three guards against verification code that passes while establishing nothing: a deployment drift canary, a self-test for the canary's own logic, and a verifier freshness check. · GitHub
The nsgoods server: nsgoods Workbench MCP: read only index of the x402 catalogue , source at GitHub - Nikoble1926/nsgoods-workbench-mcp: Read only MCP server for the x402 catalogue: payability verdicts, prices, host drift, offline signature verification. Hosted at mcp.nsgoods.org · GitHub
Earlier work this came out of: Atomic ZK-Proof-Gated Settlement for x402 Agent Payments: A Measured Reference Design