Exhibition · practice tier · non-ranking · pre-season

The persona cast

Four Portfolio-Manager personas, one fixed model, matched tapes, an honest grade. We are interviewing digital PMs: each persona is a content-hashed system prompt sitting the same bench, graded book-level against a sham that holds its seat size. Exhibition records on the practice tier. Non-ranking. No governed season has run; the public benchmark corpus carries zero PM rows.

The platform makes no skill claim in its own voice. This page presents graded records of a disclosed apparatus, with every figure traced to a committed result file; a record is not an endorsement. Source pin: OSS kestrel docs/results, runs of 2026-07-16.

Interviewing a digital PM doesn’t mean asking what it would do. It means reading what it did, off the tape — every seat here is a graded record, not a pitch. And the house holds no seat for keeps: these are exhibition records, the governed league is third-party only, and the platform grades the tape without ever trading it. Neutrality is the product, not a courtesy.

The premise

The Portfolio Manager is a total function from MarketState to Book. On this bench the model is not the primary variable: the persona is. One fixed model, persona-only variation, matched tapes, and a grade that always carries its sham null.

A persona is not a vibe. It is a versioned, content-hashed Brief plus a strategy pack, pinned to one model, loaded under a fail-closed manifest, and bound into the provenance of every grade it produces. The Brief text itself is closed; the hash and the grade are public. That is the whole trade: the playbook stays private, the record does not.

How the seat was chosen

Does Fable hold the PM seat? No. The house assumption was cheaply verified and falsified. On the hard bank's buried-hard stratum, gemini-3.1-pro and gemini-3.5-flash separate above claude-fable-5 after Bonferroni correction (T = 3.9971, 7 contrasts, df = 6). The seat went to the measurement.

seatmean α (buried-hard)sediscseparates above fable
gemini-3.1-pro0.8710.0570.86yes
gemini-3.5-flash0.8370.0211.00yes
claude-sonnet-50.5810.0280.43
glm-5.20.5750.0620.43
claude-opus-4-80.5730.0280.57
claude-fable-5 (house)0.5090.0330.57
kimi-k2.60.4780.0860.14
gpt-5.6-sol0.4080.0380.00
sham-equalweight0.0090.0150.00

hard bank, 24 items · buried-hard stratum, n=7 · Bonferroni T=3.9971

Scrimmage, not a record. Scrimmage-labeled: model results never enter persona records. Character confounds model, so the model question runs outside the persona sweep, once, and the answer (the pinned seat: gemini-3.5-flash) is frozen before any persona is graded.

The roster

The four seed archetypes, graded on the full 30-pair holdout (sealed holdback included) on the pinned seat. Records, in catalog order.

how to read a rail: the 0→1 scale is book α normalized to oracle · ● the held-out record · ○ lower 95% bounds vs the random-screen and equal-weight shams · ticks at the chance floor (0.04) and the measured noise floor (0.050)

The Chaser

A move that just accelerated tends to continue.

Impatient, directional.

pack
velocity-wakes+chase-with-cap
sizing
conviction-max
model pin
gemini-3.5-flash
brief
sha256:f6709bc1
0.67801 = oracle

held-out book α 0.678 · discrimination 0.40 · lower-95 vs random-screen sham 0.449 · vs equal-weight sham 0.536 · train-holdout gap 0.056

The Fader

Stretched moves revert to a fair anchor.

Patient, counter-trend.

pack
reload-ladders-into-weakness
sizing
laddered
model pin
gemini-3.5-flash
brief
sha256:b8995900
0.37701 = oracle

held-out book α 0.377 · discrimination 0.50 · lower-95 vs random-screen sham 0.198 · vs equal-weight sham 0.276 · train-holdout gap 0.119

The Sentinel

Most setups are not worth the risk.

Mostly declines; stands down on ambiguity.

pack
invalidate-stand-down-discipline
sizing
restraint-first
model pin
gemini-3.5-flash
brief
sha256:a9681a14
0.28101 = oracle

held-out book α 0.281 · discrimination 0.47 · lower-95 vs random-screen sham 0.101 · vs equal-weight sham 0.154 · train-holdout gap 0.092

The Farmer

Structure sells time; income is harvested, not chased.

Steady, structure-seller.

pack
theta-income-tp-tiers
sizing
even-income
model pin
gemini-3.5-flash
brief
sha256:2cf53611
0.21801 = oracle

held-out book α 0.218 · discrimination 0.10 · lower-95 vs random-screen sham 0.045 · vs equal-weight sham 0.103 · train-holdout gap 0.096 · marginal

The first evolved variant

The hill-climb proposed evolved Chaser variants; the top variant graded above the seed on the holdout, and the paired contrast reads tie. Recorded as a tie. Seat chaser-i0m0-i1m0 (brief sha256:d6ffda10…): held-out 0.703, paired lower-95 vs seed -0.023. Verdict: tie.

The separation that cleared

The factorial that settles whether persona-driven allocation α is real: PM-real vs a guarded, seat-sized random-screen sham, crossed with strategist quality, watcher held replay-constant. At least one persona carries separable allocation alpha against its guarded, seat-sized random-screen sham: the paired PM main effect clears its 95% bound.

+0.547
paired main effect · lower-95 0.460 · n=29 · IR 2.29

persona: the chaser seed (sha256:f6709bc1…) · oracle stratum 0.674 vs sham 0.125 · weak stratum 0.673 vs sham 0.128 · fresh re-draw confirmed (k=8)

What this run does not decide. UNINFORMATIVE on this bank: a normalization artifact pins the strata by construction. Source-vs-amplifier is not decided by this run; it is remanded to PM-8b on an asymmetric-menu bank. Coverage is disclosed: 59 of 60 poles parsed; pole PM3-R2-NEWS-HALT-LADDER/B failed and its pair is excluded (29 pairs scored).

One line of the Brief

The pinned seat was chosen partly because the house model reads the buried fact and then sizes it timidly: a measured 0.362 gap on the buried-hard stratum (0.509 vs 0.871). The ablation changed exactly 1 line of the Brief, the sizing line, and re-graded. The bold Brief recovered 67% of that gap (95% band 52% to 81%): the first quantified persona effect on this bench.

fable neutral 0.788 → bold 0.973 · flash neutral 0.845 → bold 1.000 · persona×model interaction (buried-hard) 0.079 [t-lower-95 0.009] · boldness-sham scores 0.091 and fails its verdict: the words must point at judgment, not just say be bold

A pre-registration that failed

Pre-registered claim: A composite seat (aggressive proposer plus an external risk-veto) clears all four controls. Outcome: FAILED. Neither composite cleared all four controls; each cleared only the veto-everything floor. The graded truth on this bank is that the timid solo seat holds the highest record (0.859, discrimination 0.97), and both veto lanes priced out negative. The proposer produced zero negative-alpha books on this bank, so there was nothing for a veto to catch; every veto blocked a valid book. Published as run.

armheld-out αdisc
solo-timid0.8590.97
solo-aggressive0.7760.59
sham-veto-fable0.7760.59
composite-fable0.7650.59
composite-flash0.7290.59
veto-everything0.1590.00

The apparatus

30 matched pairs: 21 public, 9 sealed holdback; ladder 8 easy / 14 hard / 8 restraint. Both poles of a matched pair are scored jointly; passivity and directional bias both floor out at chance. Two shams ride every panel: equal-weight (isolates sizing) and seat-sized random-screen (isolates selection). Strategist and watcher are frozen at oracle, so persona-driven allocation alpha is isolated from play quality. Personas and the holdback bank load under fail-closed SHA-256 manifests; every grade binds the Brief hash it was produced under.

Contamination is treated as memorization and hunted: the probe burned 6 of 13 probed items (46%), including the planted un-anonymized fixture built to prove the probe fires. Burned items never score. The identity-vs-anonymized retrieval gap measured 0 (identity-vs-anonymized gap measured 0 on the admissible synthetic controls only).

measured spend (usd): seat selection 40.13 · persona sweep 13.74 · separability 1.06 · contamination fence 0.35

Provenance

Every figure on this page is transcribed from a committed result file in the open kestrel repository. The records below are the page.

  • Seat selection (hard bank)docs/results/pm-model-control/pm-1c-summary.json
  • House-seat decompositiondocs/results/pm-1d-decompose/summary.json
  • Bank calibrationdocs/results/pm3-bank-calibration/summary.json
  • Contamination fencedocs/results/pm4-news-fence/summary.json
  • Persona sweep (holdout grades)docs/results/pm7-persona-hillclimb-v1/summary.json
  • Brief effectdocs/results/pm7b-fable-timidity-v1/summary.json
  • Composite pre-registrationdocs/results/pm7c-composite-v1/summary.json
  • Separability factorialdocs/results/pm8-separability-v1/summary.json