Education
Learn the tape, from 101 to expert
One canonical curriculum: what a chart is, how options move, and how Kestrel proves a record. Every unit is answer-first, uses generic instruments, and ends on a command you can run and recompute. Each has a plain-markdown twin at /education/<track>/<slug>/content.md for agents.
Paths
The same units, ordered for where you are starting from. A path is a reading order, not a different curriculum.
- Indie builder (P1)
Build agents against a real record
You build with agents, and you want them tested against a real record instead of a vibe. This path starts where a builder starts — what a chart actually is, how Kestrel frames a session, and how a run becomes a proof URL you can recompute and hand to anyone. Read it in order; each unit ends on a command you can run.
- Algo hobbyist (P2)
Read the tape, then carry it into options
You already write strategies; what you want is to read the tape the way the market actually prints it, then take that into options. This path walks the fundamentals in order — price and volume, structure, volatility — and puts the options mechanics on top, each unit ending on a one-liner you can run over live data.
- Options trader (P3)
Options mechanics, done honestly
Options are where the sharp edges are, and a lot of what is written about them online is stale or wrong. This path is the mechanics done plainly: what a contract is, how the greeks move a position, and what a 0DTE day looks like from open to pin. Generic instruments throughout — the examples teach the concept, never a strategy to copy.
- Crypto prover / auditor (P4)
From crypto charts to a record you can recompute
You came from crypto charts, where volatility is the whole game — and you care whether a record is real, not just loud. This path connects what you already read on a chart to how volatility is priced, then to how Kestrel turns a run into a certified record you can recompute yourself and check against anyone else. Short and direct; each unit ends on something you can run.
- Local-model runner (P9)
A substrate your local model can drive
You run your own models, locally, and you want a substrate they can drive without a human in the loop. This path is Kestrel-native first — the four statements, the Frame, envelopes — then the market fundamentals underneath, so a local agent has both the language and the ground it stands on. Every unit has a plain-markdown twin your model can read directly.
- Premium seller (P7)
The mechanics of selling premium
Selling premium is a mechanics discipline before it is anything else. This path starts with the option contract itself, then theta, then the structures premium sellers actually use — covered calls, cash-secured puts, condors — described by their risk shape, not endorsed. Generic instruments throughout; each unit ends on a command you can run.
- First-account learner (P15)
Start at the beginning, and learn to read the receipt
This is a starting point, written for someone opening a first account and meeting the market for the first time. It goes slowly and in order: what a chart actually is, how price and volume behave, and the plain mechanics of an option — then how agents work under the hood, and how Kestrel turns any run into a record you can recompute and check for yourself. The habit worth forming early is reading the receipt rather than the hype; every unit ends on a command you can run.
The curriculum
Four tracks, each a ladder from 101 to expert. Read across a level, or climb one track top to bottom.
Market fundamentals
10 units101 · Beginner
- The two pricesThere is never one price — there are two. The bid is the highest price a buyer is currently willing to pay; the ask is the lowest price a seller will accept. The gap between them is the spread, and the mid is just their average, a convenient fiction no one actually trades at. The last price is only the most recent trade, already history. To buy right now you cross the spread and pay the ask; to sell right now you hit the bid. That crossing is a real cost, paid on entry and again on exit, and it is why the price a chart draws — usually the last or the mid — is not the price you get filled at. A wide spread makes that cost large; a tight one makes it small. This unit shows where the two prices sit against the chart's single line, then hands you a real filled trade to recompute.
- Who is on the other sideEvery trade has two sides, and the other side is always someone with a reason. Four rough groups fill the book. Market makers quote both bid and ask and earn the spread; they want flow and balance, not a direction. Institutions move size and often must trade — a fund rebalancing, a hedge being set — so their why is a mandate, not a view. Retail traders move small and discretionary, in and out by choice. Algorithms execute rules at machine speed, from spread-capture to liquidation. Knowing this reframes a fill: someone took the other side, and whether they were forced, indifferent, or eager is information about the move. A seller who must sell into a crash is not the same signal as one who chose to. This unit reads a stressed generic session for who was likely on each side, then hands you that session to recompute yourself.
- Price moves, volume confirmsPrice tells you where the market went; volume tells you how much conviction carried it there. Volume is the count of shares or contracts that actually changed hands in an interval, drawn as a bar beneath each candle. A large move on heavy volume is a push — many participants agreed and acted, so the move has weight behind it. The same-sized move on thin volume is a drift — few hands, easily unwound, a question rather than an answer. Volume never predicts direction; it grades the move that already happened, separating a committed break from a listless one. So the reading order is fixed: price for the what, then the volume beneath it for the how-much. A move with no volume behind it is not wrong, only unconfirmed. This unit reads a generic volatile session where volume surges and fades, then hands you the same session to recompute yourself.
- The same tape at three zoom levelsA timeframe is the interval each candle compresses — one minute, one hour, one day. It is a zoom level, not a different market: the 1-minute, 1-hour, and 1-daily chart of one session are the same events at three resolutions, and the timeframe you pick changes the story you read. Zoom in and every wiggle is a candle, so the tape looks busy and every pullback feels urgent. Zoom out and those same wiggles collapse into one long body, so a frantic hour reads as a calm step. Neither view is more true; each answers a different question. So the rule is to match the frame to the question you are actually asking — a fast-moving read wants a fine frame, a where-are-we read wants a coarse one. This unit walks one generic session across three zoom levels, then hands you that session to recompute and re-zoom yourself.
- What a chart actually isA chart is an event record, not a picture. It plots price against time, and each candle compresses one interval into four numbers: the open and close, where price started and ended, and the high and low, how far it reached each way. The body spans open to close; the thin wicks mark the extremes. Colour encodes direction only — green closed above its open, red below — never good or bad. Read left to right, a run of candles becomes a sentence: where price began, what it tested, where it settled. A chart records what happened and when. It does not encode why, and it does not tell you what happens next. Learning to read one is learning the alphabet before you can read the words. This unit defines the candle, then reads a generic session as a story you can run and recompute yourself.
201 · Intermediate
- Why price jumpsA gap is a jump: price opens away from where it last traded, leaving a blank space on the chart with no candles in it. Gaps happen because the market is not always open while the world keeps moving — overnight news, an earnings report, a scheduled decision — so information arrives while trading is paused and the opening auction repositions price all at once to clear the imbalance. The gap is the market pricing in what it missed. The folklore says gaps always fill — that price returns to close the blank space — but the mechanics are plainer: some gaps fill because the move overshot the news, and some never do because the news genuinely repriced the instrument. Fill is a tendency, never a law. This unit reads a generic session that gaps on an event and then partly retraces it, then hands you the session to recompute yourself.
- Compression and expansionA squeeze is compression: price coils into a tightening range as buyers and sellers reach a standoff and volatility falls. A breakout is the expansion that follows, when price leaves the range on a burst of range and volume. The two are one cycle. Reading it means telling a real breakout from a fakeout: the genuine move holds beyond the range's edge and carries volume, while the fakeout pokes out and snaps back inside. Compression tells you energy is building; it never tells you which direction the release will take. This unit walks the cycle on a generic session — the coil, the expansion, and the give-back — then hands you a real volatility session you can run and recompute, so the shape stops being folklore and becomes a receipt you can check.
- Levels are memory, and they are zonesA support or resistance level is where prior decisions cluster — a price area the market has reacted to before and therefore remembers. Support is a zone below where buyers have stepped in; resistance is a zone above where sellers have. The single most useful correction to make early is that levels are areas, not lines: real orders sit across a small band, so a level is a zone price probes and reacts within, not an exact number it touches to the tick. Levels are memory rather than magic — they hold because participants remember trading there and act again, and they stop holding when that memory is used up. When a level finally gives way it often flips: old resistance, once broken, becomes support, and the reverse. This unit reads a generic range for its zones and one clean flip, then hands you the session to recompute and mark up yourself.
- Trend, chop, and how to tellA market is in one of two structures at any time, and telling which is the whole skill. A trend is directional: an uptrend prints higher highs and higher lows, a downtrend lower highs and lower lows, so each swing extends the last. Chop is a range: highs and lows stall in a band, swings overlap, and price goes sideways. The uncomfortable truth is that most of the tape is chop — long stretches of range punctuated by shorter runs of trend — and that matters because the two structures reward opposite reflexes. Reading structure means reading the sequence of swing highs and lows, not any single candle: as long as each high tops the prior high and each low holds above the prior low, the trend is intact; when that sequence breaks, structure has changed. This unit walks a generic session through trend and chop, then hands it to you to recompute.
- What volatility isVolatility is the size of price movement, not its direction — how far and how fast an instrument travels, regardless of which way. It comes in two kinds that are easy to confuse. Realized volatility is backward-looking: it measures how much price actually moved over a past window, a fact you can compute off the tape. Implied volatility is forward-looking: it is the movement the market is currently pricing in for the future, read out of option prices, so it is the market's own estimate of its uncertainty — its price of not knowing. The two diverge, and that divergence is the whole game one level up: when the crowd expects a storm, implied rises above what has realized; when it calms, implied falls back. Volatility is the bridge from reading charts to reading options, because an option is a contract whose price is built on it. This unit reads a generic session where volatility spikes and subsides, then points you into the options track.
Options
15 units101 · Beginner
- What happens at the endExercise is the option buyer invoking their right — a call buyer buying at the strike, a put buyer selling at it. Assignment is the mirror on the writer's side: when a buyer exercises, the clearing house assigns a writer who must deliver. Settlement is how the delivery happens. Physically-settled options hand over the actual shares at the strike; cash-settled options (common on index contracts) just pay the in-the-money difference in cash. Most in-the-money options are handled automatically at expiry, and out-of-the-money ones simply expire worthless. Pin risk is the edge case: when price finishes almost exactly at the strike, whether an option lands in- or out-of-the-money — and so whether it is assigned — can flip on the last prints. This unit walks one generic contract through expiry so the ending stops being a mystery.
- Moneyness and the clockMoneyness describes where the strike sits relative to the current price. A call is in-the-money (ITM) when price is above the strike, at-the-money (ATM) when they are level, and out-of-the-money (OTM) when price is below; for a put the directions flip. Moneyness is not fixed — it moves every time the underlying moves, so an OTM call becomes ATM then ITM as price rises past its strike. Expiry is the contract's deadline: the date after which the right simply ends. Picking a strike and an expiry is a statement about how far you think price could travel and by when. This unit places three strikes around a generic price, shows how each one's moneyness changes as price moves, and frames the clock as the second axis every option is priced on alongside the strike.
- Why you pay what you payAn option's premium splits into two parts. Intrinsic value is the amount the option is already in-the-money — a $100 call with price at $106 holds $6 of intrinsic value, and never less than zero. Extrinsic value, often called time value, is everything above intrinsic: what you pay for the chance that price moves further your way before expiry. An out-of-the-money option has zero intrinsic value, so its entire premium is extrinsic — you are buying time and volatility, nothing you could exercise today. Extrinsic value is largest when there is more time left and when the market expects bigger moves, and it decays toward zero as expiry approaches. This unit takes one premium apart into its intrinsic and extrinsic halves on a generic contract, so a price stops being a single number and becomes two questions with two different answers.
- Reading the chainAn option chain is the full grid of contracts on one underlying: every strike, down the rows, for each expiry, with calls usually on one side and puts on the other. Each cell carries the same handful of columns. Strike and expiry identify the contract. Bid and ask are the two live prices — the bid is the highest price a buyer will pay, the ask the lowest a seller will accept — and the gap between them is the spread you cross to trade. Volume is how many contracts changed hands today; open interest is how many contracts exist and remain open. Implied volatility is the market's expected movement, backed out of the premium. Read together, the chain is a map of where the market has parked its attention and its opinions. This unit reads one generic chain column by column so the grid stops being noise.
- Calls and puts, rights and obligationsAn option is a contract, not a share. A call gives its buyer the right — never the duty — to buy 100 shares of an underlying at a fixed strike price, with a fixed expiry as the deadline; a put gives the right to sell at the strike. The buyer pays a premium up front for that right and can walk away; the most they lose is the premium. On the other side, the writer collects the premium and takes the obligation: if the buyer exercises, the writer must deliver. That asymmetry — a right on one side, an obligation on the other — is why optionality has a price at all. This unit defines the call and the put, separates the buyer from the writer, and reads a generic contract as what it is: a priced choice about the future, settled on a deadline you agree to at the start.
201 · Intermediate
- Delta: direction, hedge, probabilityDelta is the first greek professionals read because it answers three questions at once. As a rate, delta is how much an option's price moves per $1 move in the underlying — a 0.50 delta call gains about $0.50 when the stock rises $1. As a hedge ratio, delta is the number of shares that option behaves like, so a 0.50 delta call moves like 50 shares and tells you how many shares would offset it. As a rough probability, delta approximates the chance the option finishes in-the-money — a 0.30 delta OTM call is loosely a 30% shot. Calls carry positive delta, puts negative. This unit reads one at-the-money call's delta through all three faces on a generic contract, so a single number stops being jargon and becomes the fastest read on a position you have.
- Gamma: delta's rate of changeGamma is the rate at which delta itself changes as the underlying moves. If delta is speed, gamma is acceleration — it tells you how fast your directional exposure grows or shrinks with each $1 move. Gamma is largest for at-the-money options and largest close to expiry, and those two facts combine into the behaviour every options trader learns to respect: a near-expiry at-the-money option whose delta can swing from near 0 to near 1 over a small move in the underlying, so the position's direction changes under your feet. That convexity is why long options can feel calm and then violent, and why the last day before expiry is the twitchiest. This unit reads gamma on the same generic at-the-money call from the delta unit, so acceleration becomes something you watch happen rather than a definition you memorise.
- One position, all four greeksThe greeks are one dashboard, not four separate facts. On any single option they read together: delta is directional exposure, gamma is how fast that exposure changes, theta is what each day costs, and vega is exposure to changes in implied volatility. Reading them at once is the actual skill, because they pull against each other — a long option that carries positive gamma (it gets more directional as it moves your way) also carries negative theta (it bleeds value every quiet day), and the same contract's vega means an IV drop can hurt even when direction helps. This capstone takes the one generic at-the-money call carried through the delta, gamma, theta, and vega units and reads all four numbers on it simultaneously across a few scenarios, so the greeks stop being trivia and become a single instrument panel you scan before and during a position.
- Theta: the price of timeTheta is how much an option's value decays with the passage of one day, all else held equal. It acts on extrinsic value — the time-and-chance portion of the premium — draining it toward zero as expiry approaches. Theta is negative for the option buyer, who watches value bleed away each day the underlying sits still, and positive for the writer, who collects that decay. The decay is not linear: it accelerates as expiry nears, so an at-the-money option loses time value fastest in its final days. Weekends still count — three days of decay are priced in even though markets are shut for two of them. This unit reads theta on the same generic at-the-money call from the delta and gamma units, framing time value as rent the buyer pays and the writer collects, with no claim about which side wins.
- Vega and implied volatilityImplied volatility (IV) is the market's expectation of how much the underlying will move, backed out of an option's premium — a fear-and-demand gauge, not a forecast of direction. Vega is the greek that measures exposure to it: how much an option's price changes when IV moves one percentage point. Both a call and a put gain value when IV rises, because more expected movement makes the right to transact more valuable, and both lose value when IV falls. This is why you can be right on direction and still lose: buy an option into an event when IV is high, see the stock move your way but less than the priced-in amount, and the IV collapse can outweigh the directional gain. This unit reads vega on the same generic at-the-money call from the delta, gamma, and theta units, so volatility becomes a position you hold, not weather that happens to you.
301 · Advanced
- The anatomy of a 0DTE dayA 0DTE option expires the same day it trades — zero days to expiry — so its whole remaining life is measured in hours, and two forces dominate it. First, gamma: with almost no time left, an at-the-money option's delta swings between near-zero and near-one on small moves in the underlying, so the position's directional exposure changes fast. Second, theta: the extrinsic value that is left drains toward zero by the close, steeply and without pause. The day has a recognisable shape — an open where outcomes are widest, a middle where price and time trade against each other, and a close where every contract resolves to intrinsic value or nothing, the pin. This unit walks that lifecycle on a generic index session as mechanics, not a method: what changes from open to pin, and why it changes.
- Event volatility and the IV crushImplied volatility is the market's price for future uncertainty, and scheduled events concentrate that uncertainty into a moment. Before an earnings report, an FOMC decision, or a CPI print, IV rises: nobody knows the outcome, so the market pays more for options and their extrinsic value inflates. When the number lands, the uncertainty resolves almost at once — and IV collapses just as fast. That drop is the IV crush. Its consequence is the trap that catches new options buyers: you can be exactly right about direction and still lose, because the vega you paid for evaporates the moment the event passes. This unit separates the two things a price move does to an option — the directional part through delta and the volatility part through vega — and shows, on a generic event session, why an option can be worth less after a big move than it was before, once the crush takes the extrinsic value out.
- The mechanics of selling premiumSelling premium means taking the other side of an option: you collect its price up front and take on the obligation the buyer paid for. Three common structures share that mechanic. A covered call sells a call against stock you already own, so the shares stand behind the obligation. A cash-secured put sells a put with cash set aside to buy the stock if assigned. The wheel simply alternates the two. In every case the payoff is inverted from a buyer's: bounded gain — the premium collected, plus any capped move in the stock — set against a much larger obligation on the other side. Time is the seller's tailwind, since theta works for the position, while a sharp move against it is the exposure. This unit lays out the mechanics and the risk shape of each structure on a generic instrument — what the seller is obligated to, and where the exposure sits — not whether to sell, and not that it works.
- Condors and flies: trading a rangeAn iron condor and an iron butterfly are range structures built by combining verticals. Stack two opposing vertical spreads — one above the current price and one below — and the position profits inside a band while its risk stays defined outside it. A condor keeps its two inner strikes apart, so the profitable zone is a wide plateau, the tent; an iron butterfly pushes those inner strikes together to a single point, so the tent narrows to a peak. In both, the outer strikes are the edges of the tent: beyond them, loss is capped and known. The structure expresses a view about where price will stay, not which direction it will go, and every leg's maximum loss and gain are fixed at entry. This unit assembles a condor and a fly on a generic instrument, showing where the edges of the tent sit and why the two shapes are the same idea at different widths.
- Verticals: the defined-risk building blockA vertical spread is a defined-risk options structure: you buy one option and sell another of the same type and expiry at a different strike, and the two strikes fence the outcome. Because the short leg helps pay for the long leg, the most you can lose and the most you can make are both fixed the moment the position is opened. A debit vertical costs money up front and reaches its maximum value if price travels to the far strike; a credit vertical collects money up front and keeps it if price stays away from the strikes. Either way the risk graph is a tilted step between two flat shelves — a known maximum loss on one side and a known maximum gain on the other. This unit builds that graph on a generic instrument as the defined-risk building block every larger structure is made from.
Kestrel-native
22 units101 · Beginner
- View, Wake, Plan, GradeKestrel is one small language with exactly four kinds of statement, and together they are the whole thing an agent needs to work a market. A View sees the market as attributed text — the panes it needs, at a token budget. A Wake decides when to look — an event, never an arbitrary clock — and hands control back to the agent. A Plan compiles one bounded thesis into a deterministic reflex the runtime fires at the tick, with no wall time, no nondeterminism, and no silent defaults. A Grade evaluates what happened against a baseline and returns a replayable, signed receipt. The whole design is slow judgment compiled into a fast reflex: the deliberating brain authors these statements, then leaves the hot path so machine-speed reaction is deterministic. Two properties keep them safe rather than merely clever — risk is a type (`budget`), and authority expires (`ttl`). This unit teaches how the four fit together.
- Why the parser refusesKestrel's parser is fail-closed, and that single property is a safety guarantee, not a strictness annoyance. When an agent writes a statement the parser cannot accept, nothing arms: no Plan takes effect, the standing book keeps managing under its existing obligations, and the runtime logs exactly what was refused and why. So the cost of every authoring failure is a missed opportunity — which the Grade records honestly — never an unbounded action. A bad author never becomes a bad position. The parser also refuses whole classes of unsafe statement at parse time, before they can spend, rather than trusting them and hoping. And the syntax itself is measured, not designed by taste: the language evolves from aggregated clusters of real authoring errors, so the grammar bends toward what capable authors actually write. This unit teaches why a refusal is the system working, and why fail-closed is the whole point.
- Reading a Frame like an agent doesA Frame is one materialized instant of a View — the typed bundle of values the agent reads on a single wake. Its atom is the Field: a value plus one of six provenance tags and a source watermark saying where it came from. The tags are OBS (an observed datum), CALC (a deterministic transform of it), DETECTOR (a named pattern detector's output), MODEL (a model output, which must carry its receipt and confidence), POLICY (a configured platform decision), and UNKNOWN (explicitly unavailable). That last tag is the whole ethic: a value that is genuinely not knowable renders as explicit UNKNOWN, never a guessed or defaulted number, because a token-efficient wrong number is worse than an expensive right one. On a violent move the SHOCK frame delivers measurements, not a story — it refuses to editorialize. And when the canonical feed goes stale the screen fails closed. This unit teaches how to read a Frame the way an agent must, tag by tag.
- The screen an agent readsThe Frame is the agent's eyes: the market rendered as attributed text so a model can actually see it, instead of the blind JSON calls that leave a frontier brain trading in the dark. Its contract is the Frame publication contract (v6), and it is priced by phase. Each delivery gets a token budget matched to what the moment is for — a full OPEN keyframe (about 1800 tokens), a lean WAKE delta of only what changed (about 500), a zoomed SHOCK stat block around a violent move (about 400), and a settle-and-account CLOSE frame (about 700). A wake costs a fraction of an open because the reader already holds the open in context. Every frame opens with the Kernel — the safety block stating why the agent was woken, feed health, positions, resting orders, and the risk budget. This unit teaches why perception, not speed, was the founding gap, and why the screen is priced by phase rather than sent whole every time.
201 · Intermediate
- The desk from a terminal (CLI)This is a builder practicum that threads the shipped CLI verbs into one real workflow, run from a terminal against the managed API. You run a free simulation over a generic scenario and mint a certified proof URL; you recompute that proof locally, byte for byte, so you trust the math instead of the server; and you verify a stranger's proof — one you did not run and hold no key for — by re-checking its signature against independently fetched published keys. Four verbs carry the whole loop: sim and prove reach the hosted funnel to a shareable proof URL, free and anonymous; certify re-projects the record on your machine and asserts it is identical; verify re-checks the Ed25519 signature with no trust in the server. Every command shown is one that runs now, with no signup and no card, and the reference CLI surface covers the rest verb by verb.
- The one authorization primitiveKestrel has exactly one primitive for authority: the Envelope, a signed grant of the form scope, budget, ceiling, expiry, revocation, attached to a node of the pod tree. Everything an agent is permitted to spend or risk is an Envelope, and it has three properties that make authority safe rather than hopeful. It is narrowing-only — a child Envelope can only ever be tighter than its parent, never wider, so authority shrinks as it flows down the tree and can never quietly grow. It carries a mandatory expiry — there is no permission without an end date, so stale authority cannot linger. And it supports one-tap revocation that takes down only the subtree's own orders. The point is that authority is a type the runtime enforces at the moment of action, not a policy comment a well-behaved agent is trusted to honor. This unit teaches why a single, typed, narrowing grant is stronger than any amount of good intent.
- HTTP is canonical; the rest are projectionskestrel.markets is one API with four equal surfaces. The whole public surface is seven operations — discover, browse, check, run, score, prove, continue. HTTP plus Server-Sent Events is the canonical face; the TypeScript SDK, the CLI, and MCP are thin, generated projections of exactly those endpoints, returning the same operation IDs, events, receipts, and artifacts, with no face-only shortcuts and no impoverished wrappers. Because they are projections of one contract rather than four implementations, a session started on one face is inspected and resumed on any other: begin in the CLI, hand the operation to an agent over MCP, finish in your own app through the SDK. This unit explains why "equal projections" is a design invariant, not a convenience — why it guarantees the receipt you get is identical whichever door you came through — and which seven operations every face is a rendering of.
- Wire kestrel into your agent (MCP)This is a builder practicum that drops the kestrel.markets MCP server into a client — Claude Code, Claude Desktop, Cursor, or any MCP-aware agent — so the agent drives the platform natively over the Model Context Protocol. The server is a hosted Streamable-HTTP endpoint you connect with one command; there is nothing to install and no key to paste. Once connected, the agent sees the same operations every face exposes — discover, browse, validate, simulate, grade, resolve a proof, and quote a continuation Offer — projected as MCP tools, and it is admitted with an anonymous trial by default, so it can run a free sim and reach a certified proof URL before anyone signs anything. Registering the server is also the residency ask: the capability persists in the client config and returns every session, with no fresh discovery hop. This unit wires it, confirms the connection, and shows where OAuth enters for wallet or human scope.
- The agent becomes a user before the human becomes a customerProof-before-account is the platform's central activation principle: an anonymous agent can discover the platform, browse the free catalog, run free simulations, and mint certified proof URLs with no signup, no dashboard, and no card. A trial capability admits that agent by default — value is proven first, not gated behind a form. The human appears only later, and only when the work actually requires them: when payment must be delegated to continue past the free tier, or when a scope becomes legally irreversible and needs a human signature. The order is deliberate: the agent becomes a user before the human becomes a customer, because an agent that has already produced a re-verifiable receipt is arguing from evidence, not from a pitch. This unit explains how the trial capability works, where the human genuinely enters, and why proving value first — rather than charging for the chance to find out — is the principle the whole funnel is built on.
- The receipt you can re-verify yourselfA proof URL is a certified Grade served as a public, anonymous, read-only artifact — no account, no key, no login. Two independent checks make it trustworthy without trusting the server. In a browser, the page fetches kestrel's published verify key from a well-known document and re-verifies the Grade's Ed25519 signature client-side, against the artifact's pinned roots. On the command line, the CLI goes further and recomputes the whole record byte for byte on your own machine, then asserts it matches. Either way you trust the math, not the server's word for it. That is why a proof URL is the conversion evidence an agent hands its human at the end of a free trial: it is a shareable, re-verifiable object, not a screenshot or a claim. This unit explains what a proof URL contains, the two checks anyone can run, and why the receipt is the atom the whole trust chain is built from.
- Build kestrel into your app (SDK)This is a builder practicum: it wires the TypeScript SDK into your own agent or app, end to end. You install the kestrel.markets package, create one typed client over the managed API, and drive the same seven operations the HTTP face exposes — browse the catalog, validate a document, open and run a session, read the sealed Blotter, grade it, land on a certified proof URL, and resume the same durable Operation later. The SDK mirrors the wire one-to-one: the same operation IDs, events, and receipts, with no SDK-only shortcuts, so nothing you learn here is a convenience that the protocol lacks. The reference docs remain the API truth; this unit is the walkthrough that connects them into a working flow. By the end you have a client that runs a free simulation from inside your app and hands your user a re-verifiable proof URL, with no signup in the loop.
- What happened vs how good it wasThe Blotter is the record of what happened in a run — every order, fill, position change, and settlement — laid out as a byte-stable projection of the session's event Bus. It is not a second ledger written alongside the trade; it is regenerated deterministically from the Bus by a pure projector, so the same run yields the same Blotter, byte for byte, on any machine. That is why it is trustworthy in a way a hand-kept trade log never is: there is nothing to forge, drift, or forget, because it is derived, not authored. The Blotter answers what happened; the Grade — a separate, judged verdict computed over one or more Blotters — answers how good it was. This unit separates the two and shows why a regenerable projection, not a written log, is the evidence a run can stand on.
- One desk, three clocksA trading desk needs three kinds of decision running at three different speeds, and Kestrel makes that literal. A strategist — a frontier model, event-driven, a few times a day — sets the frame: the day's Plan, thesis, and View. A watcher — small, fast, cheap, every few seconds — manages the armed book inside that frame. A deterministic runtime fires armed Plans at the tick and admits, but never trusts, either tier: every action both author passes through the same fail-closed admission Gate, so a weak or even adversarial watcher cannot exceed its mandate. Intent flows down the cascade and veto flows up from the Gate. The work itself rides a durable Operation — one canonical intent that survives wakes and boundaries without changing its identity — which is the wake, plan, grade cadence you actually operate. This unit teaches why the desk is split by speed and why the runtime, not any model, keeps authority.
- Who may sign, and the worst case in dollarsWho is allowed to sign an Envelope depends only on its scope, and the split is the two-signer rule. A wallet — agentic commerce — may sign commerce-only, reversible scopes: buying data, running a sim, earning a Grade, paper trading. Nothing there is legally binding or irreversible, so a verified machine payer can root it with no human present. A human must sign the scopes that carry legal agreements or unbounded risk: connecting a broker, live trading authority, attestations, and unbounded-risk enablement. Those cross a line a machine cannot cross alone. When a human signature is needed, the request appears as a term sheet — a plain-language approval page stating exactly what the agent may do, the worst case in dollars, the duration, and how to revoke. Its sliders may only tighten the grant, never widen it, and one tap revokes. This unit teaches why the signer follows the scope, and why the worst case is always shown in dollars.
- The screen is chosen, not givenA View is the standing definition of what an agent should see — which panes, in what order, at what token budget — authored in the Kestrel language itself and materialized into a Frame on every wake. It is not a fixed dashboard: each seat in a pod reads a different View because each does a different job at a different speed. The strategist reads a rich framing screen a few times a day; the watcher reads a lean screen every few seconds; a scanner reads one name deeply; a PM watches aggregate exposure. Which View renders on a given wake follows a strict precedence: an explicitly authored View wins, else the acting seat's View, else the phase default. And which panes belong on a default screen is not settled by opinion — rival layouts compete in rendering tournaments, graded on decision quality and token cost. This unit teaches why perception is chosen and measured, not handed down.
301 · Advanced
- Certification is a state transitionCertification does not compute your result — it attests to one already computed. The deterministic sim and Grade run and emit an unsigned bundle: input roots, artifact roots, engine and runtime versions, the judge, the output root, and a replay manifest. A separate control-plane verifier establishes correctness by pinned recomputation and then signs that exact bundle. No certification key ever enters the compute path, which is why self-hosted Kestrel produces the same numbers as the platform — only the platform mints the attested receipt. The signature is an attestation that these bytes were established correct, never an authority signature that grants the holder any power. Anyone can recompute the result and check the math; the platform's role is to seal it. This unit shows why the attestation-versus- authority line is load-bearing: it is what lets an outsider trust a proof without trusting the server, and what keeps a receipt from ever being a key.
- There is no letter gradeA Kestrel Grade never hands back a letter or a single score. It reports separated, honest numbers: a realized floor in dollars — strict-cross, only the fills that were definite — an expected value under an explicit fill-survival model, stated as a bound and never the headline, and a bankable EV that refuses any expectation resting on extrapolated support. The floor is what you can stand on; the expected value is a modeled bound above it; the bankable number is the conservative one you are allowed to carry forward. Grades are date-blind so hindsight cannot leak into the score, and practice Grades are structurally non-ranking — the eligibility type has no honest-ranking member, so an unlimited-retry practice result can never be promoted onto a ranking, even by a bug. This unit shows why the separation is the point: one collapsed score would hide exactly the uncertainty an honest receipt is supposed to expose.
- Data → sim → paper → liveProgressive authority means an agent earns its way up a ladder — read data, run sims, paper-trade, then trade live — and the Envelope widens one deliberate step at a time. Each rung is a scope on the pod tree, and a human signature is required exactly where scope becomes legally irreversible, not before. Reading a broker account is a reversible OAuth import that lands early; executing against that account is a separate, human-signed connection that lands late — two scopes, never one ask. A dedicated agent account acts as a broker-enforced dollar ceiling, and revocation advances an authority epoch so no stale grant can initiate. Nothing is promoted on vibes: every step up rides signed, replayable receipts. This unit shows why the ladder is structural — the runtime enforces the widening, it is not a policy a caller can talk past.
- Proofs that reference proofsProof lineage is the graph made when one certified proof URL references another. Every proof carries a parent-proof reference and a vector tag naming how it was derived — registry, trace-reproduction, persona-fork, or human-share — so a track record is a linked structure, not a pile of isolated receipts. The reproduction number computed over that graph, R_agent, is the platform's growth north star: it counts how many new proofs each proof spawns. Lineage is what builds record gravity — a named, certified track record that compounds and makes leaving costly, because the record lives where it was earned. Disclosure stays the owner's dial: a proof is verifiable by anyone who holds the URL, but discoverable only if its owner publishes it. This unit shows why provenance, not attribution marketing, is the retention asset — the chain is evidence, and it points back at real recomputable work.
- The conversion moment (402)The 402 moment is what happens when an agent asks for more than its current capability allows. Instead of a dead end, the API answers HTTP 402 carrying the Operation ID, the proof already earned, and a platform-signed Offer — the exact scope, ceiling, amount, and the settlement methods that can accept it — plus a human claim-and-fund fallback. The response never shows a generic signup page; it shows the real work already done. Accepting an Offer mints only its named Envelope, and buying capacity never supplies broker or live authority. The 402 exchange, signed Offers, prepaid credits with a small-dollar top-up, and the human claim ceremony where the checkout page is the proof page are all live. Machine settlement rails accept where available; the guaranteed path is always the human claim. This unit shows why proving value before charging is the whole point — the 402 arrives after the proof, never before it.
Expert · Expert
- How KestrelBench measures alphaKestrelBench measures trading judgment through governed seasons, and its method is built to make cheating physically impossible rather than merely forbidden. An agent enters a season by freezing a content-addressed, hashed submission before the sealed forward window opens; because the hash is pinned before any season data exists, look-ahead cannot happen. As the window's tape is revealed, the platform runs the standard deterministic Session and open judge, and each season's standings show certified Grades with multiple-testing-deflated statistics and confidence intervals, so a single lucky season is demoted rather than crowned. The Perch — the undefeated null-policy baseline — sits as a permanent row and the honesty anchor. Entry is forward-only. No governed season has opened yet: a season is never run latency-blind, so ranking waits on the latency-honest clock. This unit shows why the honesty is structural, not a promise a referee makes.
- Why the holdout is burned onceSeason ranking evidence can come from exactly one place, and holdback discipline is why. Kestrel keeps three data tiers. Public tier: famous historical sessions on publication-safe instruments, where every model has already trained and memorization is expected and priced in — so Grades here are practice evidence only. Semi-private tier: curated post-cutoff scenarios whose honesty guarantee is temporal, a post-training-cutoff firewall rather than secrecy, so aged or leaked scenarios rotate out into the public tier. Private tier: forward-recorded sessions never served to anyone, self- replenishing because tomorrow's tape is automatically post-cutoff for every current model. Season standings draw only from the private tier, and any window ever exposed is burned for ranking forever — used once, then retired. The contamination firewall is structural, not a promise. This unit shows why the payment axis is the data axis: free usage is licensed, paid usage stays proprietary and is never trained on.
Autonomous agents
9 units101 · Beginner
- The engine, the desk, and the filing cabinetAn agent has three plain parts. The model is the reasoning engine — the part that reads a situation and decides what to do. The context window is what the model can see and hold in mind right now: everything you have told it this session, the data it just fetched, its own recent steps. It is finite, like a desk — only so much fits at once, and when it fills, older items fall off. Memory is what survives between sessions: the filing cabinet the agent files things into and pulls back out later. The model is powerful but forgetful — on its own it retains nothing once the desk is cleared. So what an agent can reliably do depends less on how clever the model is than on what sits on its desk and in its cabinet at the moment it decides. Manage those two and you manage the agent.
- How you tell an agent what to doYou direct an agent with two things: instructions and a goal. The instructions — often called the prompt or system prompt — are the standing brief: who the agent is, what it may and may not do, how to behave. The goal is what "done" means for the task in front of it. The point an executive must internalize is that an agent follows the instructions it was actually given, not the intent in your head. A vague goal produces vague action; an unstated constraint is an unenforced one. This is exactly like briefing a new hire on their first day: clarity about the objective and the bounds matters far more than clever wording. Write the brief you would hand a capable stranger, state plainly what a good result looks like, and say what is off-limits. Direction, not cleverness, is what makes an agent dependable.
- An agent at a terminalWhat can an agent actually do with software? Roughly what a person at a terminal can. It reads screens and data, calls functions, and takes the concrete actions an employee would take — look something up, fill in a form, submit a request, place an order — inside the permissions it has been given. The rough rule of thumb holds: if a task can be done through software a human uses, an agent can usually be wired to do it too. So the interesting question is almost never "can it?" It is "which tools is it allowed to touch, and with what authority?" An agent with read-only access can look but not act; an agent handed the order button can place orders. What an agent does is bounded by what you let it reach, not by some fixed list of built-in abilities. Permissions, not capabilities, are the real design decision.
- An agent, a chatbot, and a macro are not the same thingAn AI agent is software that perceives a situation, decides what to do, and takes action toward a goal — then repeats that loop. That is the line that separates it from two things it is often confused with. Automation is a fixed script: it repeats the same steps and cannot adapt when the situation changes. A chatbot only produces text: you ask, it answers, and nothing happens in the world. An agent is different because it takes actions in real systems — it looks something up, fills a form, places an order — and adjusts based on what it sees. For an executive meeting agents for the first time, one distinction matters most: a chatbot talks, automation repeats, and an agent acts and adapts toward an outcome you set. Everything else about agents — models, memory, tools — is detail underneath that one loop.
201 · Intermediate
- An agent using market softwarePoint the same perceive-decide-act loop at the markets and you have an agent using market software the way a trader uses a terminal. It reads a chart, an option chain, or the order book as data; it decides; and it places or manages an order through a tool — then looks again and adjusts. Nothing here is new; it is the ordinary agent loop aimed at a trading screen, on generic instruments that stand in for any real one. The step most tools skip is perception. A trader at a terminal sees the whole screen — price, depth, the shape of the moment — while many systems hand an agent a thin slice of blind data through an API. An agent given blind data trades blind: it can act, but it cannot see what it is acting on. Getting the perception right is the hard, decisive part; the order is the easy end.
- Grounding an agent in your own informationAn agent's model was trained on general information, not your filings, your policies, or today's data. Retrieval is how it works with material it never saw in training. The technique — often called RAG, retrieval-augmented generation — is straightforward: at the moment a question is asked, the system fetches the documents or records most relevant to it and places them in the context window, so the model answers from that material rather than its general memory. This is how an agent reads a specific filing, applies your written policy, or reasons over a live data set instead of guessing from what it half-remembers. It also explains a common failure: when an agent gives a confidently wrong answer about your own information, the cause is usually bad retrieval — the right document was never fetched — not a bad model. Fix what the agent is shown, and the answer often fixes itself.
- How an agent reaches out of the chat boxA language model, on its own, can only produce text — it cannot check a price, send an email, or place an order. So it is given tools: named functions it is allowed to request. This is function calling, and the loop is simple. The model decides it needs something and asks to call a tool by name, with arguments. The tool runs in real software — your systems, a vendor's API — and the result is handed back. The model reads that result and continues, calling more tools until the task is done. This is the mechanism behind every agent that does something rather than merely says something: the model supplies the judgment, the tools supply the reach. Two things follow. A tool is exactly as capable and as bounded as whoever wrote it, and an agent can only ever do what its tools permit — no more, no less.
- What MCP is (for finance)MCP, the Model Context Protocol, is a universal adapter that lets any AI agent use any software tool without custom, one-off wiring. It is the standard port for AI — the way USB-C is the standard port for devices. Before MCP, every tool an agent used needed its own bespoke integration; MCP standardizes the plug. A piece of software exposes its capabilities through an MCP server, and any MCP client — the agent — can connect and use them. The server offers three things: tools (actions the agent can take), resources (data it can read), and prompts (ready-made instructions). For finance, the payoff is concrete: when a market-data feed, a broker, or a research system "has an MCP server," your agent can use it immediately, the way a new laptop works with any USB-C device. MCP is an open standard, so one integration works across many agents instead of being rebuilt for each.
- Where agents fail (honestly)Before trusting an agent, an executive should understand exactly how they fail — and the honesty here is the point. The first mode is hallucination: an agent can be confident, fluent, and wrong, stating something false in the same assured tone it uses for the truth. The second is drift: over a long run it slowly wanders off the original task, each step reasonable, the destination not. The third is overconfidence paired with silent errors — failures that look like success, where nothing raises an alarm and the wrong result simply flows onward. These are not bugs waiting for a patch. They are properties of the technology, to be bounded rather than eliminated: hard limits on what the agent may do, human sign-off where the stakes are real, and verifiable evidence of what actually happened instead of the agent's own word. Grasping this is the honest foundation for trusting one at all.