Write the number, not the adjective
There is a specific discomfort in writing a probability next to a claim you cannot support with data. Seventy per cent feels like a lie, because you did not compute it. 'Fairly confident' feels honest, because it commits to nothing. The feeling is exactly backwards: the adjective is the dishonest one, since it can be reconciled with any outcome after the fact.
The discomfort also passes, faster than people expect. Three or four cards in, a number stops feeling like a claim to precision and starts feeling like what it is — a compressed statement of how surprised you would be. Nobody in the room mistakes it for a measurement. What they can do, which they could not before, is disagree with it precisely. 'I'd say forty' is a conversation. 'I'm less sure than you' is not.
What makes this worth the effort is not accuracy on any single call. It is that a stack of resolved cards has structure. Individual results are noise. Twenty of them show whether you are systematically overconfident in the high band, whether your timelines slip by a predictable amount, whether you are better at reading customers than reading competitors. Those are correctable errors, and none of them are visible one decision at a time.
Two failure modes, both common. The first is scoring cards individually, which turns the practice into a verdict on the person rather than a measurement of a tendency. The second is letting the cards become a management artefact — the moment a number is a performance input, everything drifts to a safe fifty per cent and the exercise teaches nobody anything. Cards belong to the person who wrote them; only the aggregate belongs to the team.
The part nobody mentions is that most judgement calls never resolve. You do not learn whether the hire you did not make would have worked, or what the strategy you dropped would have done. Calibration only covers the visible subset, and the invisible subset is larger. That is a real limit, not a reason to skip the visible part.