A system that reports on itself is only worth its calibration. Let it predict its own confidence, then act, and measure the gap between what it says and what happens. A self-report that tracks the truth is worth bits; one that doesn't is worth nothing.
Everything on the plot is computed from actual sampled trials: the diagonal is perfect calibration by definition (reported confidence = empirical accuracy), each bar is the real hit-rate of trials that claimed that confidence, the shaded gap is the miscalibration. ECE (expected calibration error) and the Brier score are the standard, checkable measures — the worth of a self-report is exactly how close its bars sit to the diagonal. Drag SELF-BIAS and watch a confident-but-wrong reporter's worth collapse even as it keeps talking.
The 'self-model' is a toy confidence channel with an injected bias and noise — it is a stand-in for introspective access, not the thing itself. No claim that the system knows itself; only that its self-report can be scored against outcomes. That scoring is the whole point: introspection you can't calibrate is worth assuming to be worth nothing.