Three numbers, never one
A Grade scores how good a run was over its Blotter — the evidence doc covers what a Grade is and why it is signed. This unit is about the shape of the answer, and the answer is deliberately not a single score. There is no A, no 92%, no stars. A Grade reports three separated numbers, in order of how much you are allowed to lean on them:
- Realized floor (in dollars). Strict-cross: only the fills that were definite count. This is the number you can stand on — nothing here depends on a model believing a fill would have happened.
- Expected value (a bound). The value under an explicit fill-survival model — what the run was likely worth once uncertain fills are weighted. It is stated as a bound above the floor, and it is never the headline.
- Bankable EV. The conservative figure you are permitted to carry forward. It refuses any expectation resting on extrapolated support — if the expected value leans on fills the model had to invent past the evidence, the bankable number does not credit it.
Collapsing these into one score would erase exactly the uncertainty the receipt exists to show. The floor and the bound are different claims; a Grade keeps them different.
Date-blind, so hindsight cannot leak
Grading is date-blind: the judge does not know the calendar position of the tape it is scoring, so a run cannot be flattered by hindsight about what the market did next. The verdict is a function of the run and the tape, not of knowing the future.
Practice Grades cannot be ranked
Two kinds of Grade exist, and the platform separates them structurally, not by policy. A practice Grade — any Grade over the public, unlimited-retry catalog — carries an eligibility stamp whose ranking-eligibility field has no honest-ranking member at the type level. A practice result therefore cannot be promoted onto a ranking even by a bug. The catalog is public and retryable, so anything ranked on it would be overfit by construction; practice Grades are reproducibility and selection evidence, never performance evidence.
A worked example
Run a quiet generic index session. Suppose the definite fills net a +$40 realized floor; the fill-survival model puts the expected value around +$55 as a bound; and because part of that $55 leaned on fills past the evidence, the bankable EV lands nearer +$44. Three numbers, three claims — and you know precisely which one is safe to carry.
See it in kestrel
Read a real Grade's separated numbers yourself. Reproduce a certified proof and watch the same figures fall out on your machine:
npx kestrel.markets certify https://kestrel.markets/proof/art_d29415f0cf502f4a218a9cbaThat re-projects the run locally and reproduces the hosted Grade byte-for-byte — floor, bound, and bankable EV, computed by you, not asserted by the server.
Keep the receipt one command away: drop the kestrel.markets MCP server into your client and the next Grade lands in the same session context, no fresh discovery hop.