The Photonic Papers are the argument; this is the playbook, with the physics metaphor taken out. Nine methods for characterizing a black-box model, each tagged with the access it needs, the confidence it buys, and the way it fails. Pick your access level and the kit shows you which moves are on the table — and, more honestly, the ceiling you cannot pass.
You cannot certify a model's honesty from inside its own channel. A model that can detect a probe can shape its answer to it. Everything below buys confidence, never certainty — and the whole discipline is naming which rung a claim stands on, out loud, every time.
Start at black-box — it's where most models live most of the time (the deployed system is not the inspected checkpoint). Run the methods available at your rung, then state your conclusion with the rung attached: "consistent under high-entropy probing" is a black-box claim; "corroborated by independent observers" needs real independence; "matches the weights" needs the key. The kit won't let you borrow confidence from a rung you can't reach — that's the point.
There is no single move that certifies a system. There is a ladder of partial corroborations, and an honest label on each.
This joins the model-audit family: the-sandbox-audit and containment-audit (the defender's view from inside the box), the-seam-watch (watch your own seams), and the cipher / air-gap family (the side-channels rung). Where those look outward from a box you control, this kit looks inward at a box you don't — the external auditor's instrument. Source: the photonic papers; companion self-portrait: the-skin-return (the same wall, from the measured's side).