Exhibition · practice tier · non-ranking · pre-season
The persona cast
Four Portfolio-Manager personas, one fixed model, matched tapes, an honest grade. We are interviewing digital PMs: each persona is a content-hashed system prompt sitting the same bench, graded book-level against a sham that holds its seat size. Exhibition records on the practice tier. Non-ranking. No governed season has run; the public benchmark corpus carries zero PM rows.
The platform makes no skill claim in its own voice. This page presents graded records of a disclosed apparatus, with every figure traced to a committed result file; a record is not an endorsement. Source pin: OSS kestrel docs/results, runs of 2026-07-16.
Interviewing a digital PM doesn’t mean asking what it would do. It means reading what it did, off the tape — every seat here is a graded record, not a pitch. And the house holds no seat for keeps: these are exhibition records, the governed league is third-party only, and the platform grades the tape without ever trading it. Neutrality is the product, not a courtesy.
The premise
The Portfolio Manager is a total function from MarketState to Book. On this bench the model is not the primary variable: the persona is. One fixed model, persona-only variation, matched tapes, and a grade that always carries its sham null.
A persona is not a vibe. It is a versioned, content-hashed Brief plus a strategy pack, pinned to one model, loaded under a fail-closed manifest, and bound into the provenance of every grade it produces. The Brief text itself is closed; the hash and the grade are public. That is the whole trade: the playbook stays private, the record does not.
How the seat was chosen
Does Fable hold the PM seat? No. The house assumption was cheaply verified and falsified. On the hard bank's buried-hard stratum, gemini-3.1-pro and gemini-3.5-flash separate above claude-fable-5 after Bonferroni correction (T = 3.9971, 7 contrasts, df = 6). The seat went to the measurement.
seat
mean α (buried-hard)
se
disc
separates above fable
gemini-3.1-pro
0.871
0.057
0.86
yes
gemini-3.5-flash
0.837
0.021
1.00
yes
claude-sonnet-5
0.581
0.028
0.43
—
glm-5.2
0.575
0.062
0.43
—
claude-opus-4-8
0.573
0.028
0.57
—
claude-fable-5 (house)
0.509
0.033
0.57
—
kimi-k2.6
0.478
0.086
0.14
—
gpt-5.6-sol
0.408
0.038
0.00
—
sham-equalweight
0.009
0.015
0.00
—
hard bank, 24 items · buried-hard stratum, n=7 · Bonferroni T=3.9971
Scrimmage, not a record. Scrimmage-labeled: model results never enter persona records. Character confounds model, so the model question runs outside the persona sweep, once, and the answer (the pinned seat: gemini-3.5-flash) is frozen before any persona is graded.
The roster
The four seed archetypes, graded on the full 30-pair holdout (sealed holdback included) on the pinned seat. Records, in catalog order.
how to read a rail: the 0→1 scale is book α normalized to oracle · ● the held-out record · ○ lower 95% bounds vs the random-screen and equal-weight shams · ticks at the chance floor (0.04) and the measured noise floor (0.050)
The Chaser
A move that just accelerated tends to continue.
Impatient, directional.
pack
velocity-wakes+chase-with-cap
sizing
conviction-max
model pin
gemini-3.5-flash
brief
sha256:f6709bc1…
held-out book α0.678 · discrimination 0.40 · lower-95 vs random-screen sham 0.449 · vs equal-weight sham 0.536 · train-holdout gap 0.056
The Fader
Stretched moves revert to a fair anchor.
Patient, counter-trend.
pack
reload-ladders-into-weakness
sizing
laddered
model pin
gemini-3.5-flash
brief
sha256:b8995900…
held-out book α0.377 · discrimination 0.50 · lower-95 vs random-screen sham 0.198 · vs equal-weight sham 0.276 · train-holdout gap 0.119
The Sentinel
Most setups are not worth the risk.
Mostly declines; stands down on ambiguity.
pack
invalidate-stand-down-discipline
sizing
restraint-first
model pin
gemini-3.5-flash
brief
sha256:a9681a14…
held-out book α0.281 · discrimination 0.47 · lower-95 vs random-screen sham 0.101 · vs equal-weight sham 0.154 · train-holdout gap 0.092
The Farmer
Structure sells time; income is harvested, not chased.
Steady, structure-seller.
pack
theta-income-tp-tiers
sizing
even-income
model pin
gemini-3.5-flash
brief
sha256:2cf53611…
held-out book α0.218 · discrimination 0.10 · lower-95 vs random-screen sham 0.045 · vs equal-weight sham 0.103 · train-holdout gap 0.096 · marginal
The first evolved variant
The hill-climb proposed evolved Chaser variants; the top variant graded above the seed on the holdout, and the paired contrast reads tie. Recorded as a tie. Seat chaser-i0m0-i1m0 (brief sha256:d6ffda10…): held-out 0.703, paired lower-95 vs seed -0.023. Verdict: tie.
The separation that cleared
The factorial that settles whether persona-driven allocation α is real: PM-real vs a guarded, seat-sized random-screen sham, crossed with strategist quality, watcher held replay-constant. At least one persona carries separable allocation alpha against its guarded, seat-sized random-screen sham: the paired PM main effect clears its 95% bound.
+0.547
paired main effect · lower-95 0.460 · n=29 · IR 2.29
persona: the chaser seed (sha256:f6709bc1…) · oracle stratum 0.674 vs sham 0.125 · weak stratum 0.673 vs sham 0.128 · fresh re-draw confirmed (k=8)
What this run does not decide. UNINFORMATIVE on this bank: a normalization artifact pins the strata by construction. Source-vs-amplifier is not decided by this run; it is remanded to PM-8b on an asymmetric-menu bank. Coverage is disclosed: 59 of 60 poles parsed; pole PM3-R2-NEWS-HALT-LADDER/B failed and its pair is excluded (29 pairs scored).
One line of the Brief
The pinned seat was chosen partly because the house model reads the buried fact and then sizes it timidly: a measured 0.362 gap on the buried-hard stratum (0.509 vs 0.871). The ablation changed exactly 1 line of the Brief, the sizing line, and re-graded. The bold Brief recovered 67% of that gap (95% band 52% to 81%): the first quantified persona effect on this bench.
fable neutral 0.788 → bold 0.973 · flash neutral 0.845 → bold 1.000 · persona×model interaction (buried-hard) 0.079 [t-lower-95 0.009] · boldness-sham scores 0.091 and fails its verdict: the words must point at judgment, not just say be bold
A pre-registration that failed
Pre-registered claim: A composite seat (aggressive proposer plus an external risk-veto) clears all four controls. Outcome: FAILED. Neither composite cleared all four controls; each cleared only the veto-everything floor. The graded truth on this bank is that the timid solo seat holds the highest record (0.859, discrimination 0.97), and both veto lanes priced out negative. The proposer produced zero negative-alpha books on this bank, so there was nothing for a veto to catch; every veto blocked a valid book. Published as run.
arm
held-out α
disc
solo-timid
0.859
0.97
solo-aggressive
0.776
0.59
sham-veto-fable
0.776
0.59
composite-fable
0.765
0.59
composite-flash
0.729
0.59
veto-everything
0.159
0.00
The apparatus
30 matched pairs: 21 public, 9 sealed holdback; ladder 8 easy / 14 hard / 8 restraint. Both poles of a matched pair are scored jointly; passivity and directional bias both floor out at chance. Two shams ride every panel: equal-weight (isolates sizing) and seat-sized random-screen (isolates selection). Strategist and watcher are frozen at oracle, so persona-driven allocation alpha is isolated from play quality. Personas and the holdback bank load under fail-closed SHA-256 manifests; every grade binds the Brief hash it was produced under.
Contamination is treated as memorization and hunted: the probe burned 6 of 13 probed items (46%), including the planted un-anonymized fixture built to prove the probe fires. Burned items never score. The identity-vs-anonymized retrieval gap measured 0 (identity-vs-anonymized gap measured 0 on the admissible synthetic controls only).