One tier can hold ranking evidence — here is why
A benchmark is only as honest as the data no contestant has seen. The evidence doc names the three tiers; this unit is about the discipline that keeps them apart and why season standings may draw from only one of them.
- Public. Famous historical sessions — an earnings gap, an FOMC day, a tariff-shock session — on publication-safe instruments, served derived-only. Every model has trained on these events; memorization is expected and priced in. That is exactly why Grades over this tier are practice evidence, never performance evidence.
- Semi-private. Curated post-cutoff scenarios paid subscribers run private sims against. Its honesty guarantee is temporal — a post-training-cutoff firewall, not secrecy. Because the guarantee is about when, not whether it leaked, aged or exposed scenarios simply rotate into the public tier and retire from any evidence role.
- Private. Forward-recorded sessions never served to anyone. Recording forward makes this tier self-replenishing and automatically post-training-cutoff for every current model — tomorrow's tape cannot be in yesterday's training set. This is the only tier season ranking comes from.
Burned once
A private window is used for ranking exactly once. The moment it is exposed — run, published, or folded into the public catalog — it is burned for ranking forever. There is no "re-use the holdout"; the instant it is seen, its power to measure the unseen is gone, and it retires into practice data.
This is not a courtesy the platform extends. The contamination firewall is structural — the training side of the program cannot read holdback manifests, sealed traces, or forward tape paths — so the seal is enforced by construction, not by a promise anyone has to keep.
The payment axis is the data axis
Holdback discipline has a commercial mirror: free-is-licensed, paid-is-proprietary. Anonymous free usage grants the platform a training license over its traces — that is the free tier's only price. Paid usage is 100% proprietary to the customer and is never trained on. The tier a customer pays into is the tier their traces stay sealed in; the two axes are the same line seen twice.
See it in kestrel
The public tier is the one you can reproduce end to end. Recompute a certified proof over a publication-safe generic session:
npx kestrel.markets certify https://kestrel.markets/proof/art_d29415f0cf502f4a218a9cbaThat re-projects the run locally and reproduces the hosted result byte-for-byte — public-tier data, memorization-expected and priced in, which is precisely why it is practice evidence and never a season rank.
Keep it one command away: drop the kestrel.markets MCP server into your client and reproduction stays available across sessions with no fresh discovery hop.