Hemlock has completed Phase 1 of live testing. From March through July 2026 the actively-traded book returned +37% on deposited capital β€” roughly +160% annualized. Active trading is now intentionally paused while we redesign and scale the system for Phase 2, at a higher level of investment.
What you're seeing below reflects that pause, not inactivity: only a couple of positions remain open (being wound down), and the account value has been deliberately flat since August. Phase 1's gains came from hands-on management of a live book; Phase 2's engine β€” the Canopy research ensemble β€” has since been built and validated in a 12-month walk-forward simulation, and the current work is hardening its research and screening layers before new capital deploys. Nothing trades live until it earns its way through that gauntlet.
Total account value · Kalshi & Polymarket
$2,864.54
+$17.76 (+0.62%) past 24h
Kalshi + Polymarket ($2,550 · cached 2026-09-02) marks · as of 9/2 12:09 AM PT · refreshes each load (≀1 min)
Cash
$2,550.60
settled, ready to deploy
In positions
$313.94
2 open markets
Deposited
$2,092.39
lifetime, since Mar 2026
Net P&L
+$772.15
+36.9% on deposits
+138% annualized

Value over time

Total value Contributed (deposits) Deposit Sale — realized gain the gap between the two lines is trading gains
Chart tracks the combined Kalshi + Polymarket account. Snapshots record both venues from 9/01/26; before that, the Polymarket segment is shown at its net-deposit value (it was barely traded, so the approximation is within a few dollars). The late-August dip-and-recovery is an internal transfer — capital withdrawn from Kalshi and redeposited to Polymarket — and the contributed line moves with it, so the gap between the two lines stays honest trading gains. Solid line: live hourly marks (since Jul 2, 2026). Dashed: reconstructed from trade history — positions at cost, value stepping only at deposits, trades, and settlements (May 4 onward; the Mar–Apr era, under $300 total, is beyond Kalshi's export window). Reconciles with the live balance to the cent.

System Design

Canopy — an ensemble research engine that prices prediction markets: three independent AI analysts, a conditional reconciler, and a market blend, all validated in a 12-month walk-forward simulation before real dollars move. Nothing trades automatically.

Overall MATES System Design

The full live pipeline: market ingestion → screening → the Canopy research ensemble → risk gates & sizing → a human places every trade. (Rev 8/31.)

Hemlock end-to-end system diagram

MATES Simulation Environment System Design

"Tree Ring" — how every strategy earns trust before real dollars: a 12-month walk-forward replay with an honest bankroll. (Rev 9/01.)

Tree Ring 12-month walk-forward simulation diagram
Previous designs — walk-forward SVG, Canopy component diagram & the single-principal era (kept for comparison)
A year replayed honestly: $40,000, one week at a time cash is finite, wins can't be re-bet until they settle, and every model call must be earned 1 · Settle markets that resolved this week pay out — cash returns (with a settlement lag) 2 · Sweep prices this week's price for every open market — read from the archive, effectively free 3 · Triggers — who earns an (expensive) re-evaluation? new listing · price drifted from our estimate · news-scout spike · random audit capped per week, damped when cash is low — a desk that can't bet doesn't pay for research 4 · Evaluate — the Canopy engine prices the queue the ensemble + reconciler pipeline (see Market Evaluation) · every estimate is cached, so re-running the year with different sizing rules costs nothing 5 · Trade — fresh, drift-clean estimates only blend with the market price → quarter-Kelly of current equity → caps on liquidity, per-theme exposure, and available cash · a stale estimate never trades against a moved price 6 · Ledger positions lock their capital until resolution (no early exits) · weekly equity snapshot NEXT WEEK — 52 TIMES After week 52 — the honest scorecard realized annualized return (vs the old flattering convention, side by side) · max drawdown capital utilization · edge left on the table · which triggers earned their keep THE HONESTY RULES Finite cash — a win locked in a position can't be re-bet until it settles Hold to resolution — no early exits, no convenient re-pricing Leakage firewall — models trained before the window; every news search date-gated to the simulated day Research is paid for — every estimate must be justified by a trigger Friction is real — spread haircuts, liquidity caps, settlement lags Plumbing is proven — a no-skill stub engine must LOSE money and a peek-at-answers stub must win before any real engine's numbers are trusted Random audits — a weekly control group of refreshes nobody asked for, so trigger value is measured against chance instead of assumed
Why this exists: the older backtest priced each market once at fixed points in its life and assumed resolved capital was instantly re-bet — a flattering fiction. The walk-forward replays the year the way a real desk lives it: money locks up, opportunities arrive on their own schedule, and research costs money. The same harness is the A/B laboratory — different estimation engines run the identical year, and cached estimates let sizing and allocation variants replay for free. Clay steps are deterministic code; the purple step is where the research agents run.
DATA IN Polymarket prices · full resolution rules · order flow · settlements Live news, date-gated GDELT global news index via BigQuery — the agents run their own searches Market universe & evaluation triggers ~1,200 tradeable markets scanned cheaply every cycle — a market earns a full (expensive) evaluation only on a trigger: new listing · price drifted from our estimate · news scout detects a coverage spike · random audit (the control group) PROVEN IN A 12-MONTH WALK-FORWARD FIRST ① Research Ensemble — three model families price the market independently each analyst is blind to the market price AND to the others — diversity of model family is the ingredient that works Claude Sonnet Anthropic · frontier GPT-OSS 120B OpenAI · open-weight DeepSeek V3.1 DeepSeek · open-weight Agent tools the model chooses search_news date-gated GKG index read_article point-in-time, sanitized submit_estimate structured verdict used by ① and ③ ② Aggregate — geometric mean of odds pure code, no model · the members' SPREAD is kept as a signal — disagreement flags danger ③ Reconciler — a judge that fires only when it's needed trigger: analysts >25 points apart, or the ensemble >30 points from the market price reads all three write-ups, finds the crux, runs its own targeted searches, issues its own number — never averages calm markets skip the judge ④ Blend with the market price — log-odds, 50/50 the crowd gets a vote: extreme divergence becomes a small position instead of a blind all-in ⑤ Quarter-Kelly sizing + risk rails bet a quarter of the mathematically optimal fraction · caps per market, per theme, and on market liquidity ⑥ Trading Desk — a human places every wager hold to resolution · nothing auto-trades Settlement & scoring every wager scored: closing-line value · Brier vs the market · which trigger paid for the estimate OUTCOMES RECALIBRATE RUNS ON Google Cloud Rundashboard + scheduled jobs Cloud SQLalerts · predictions · trades BigQueryGDELT news · SQL date firewall Secret Mgr + Storagekeys · market archive · images
Why an ensemble: re-asking one model gives near-identical answers, but different model families disagree in genuinely useful ways — cross-model agreement carries measurable predictive value, and their spread doubles as a danger gauge. Why price-blind: an analyst who can see the market price anchors to it; the price enters only later, in deterministic code — as the Reconciler's alarm and the 50/50 blend that turns extreme disagreement with the crowd into a small position rather than a bet-the-house moment. The Trading Desk is a human — nothing trades automatically. The proving ground: before any configuration touches real dollars it replays 12 months of history in a walk-forward simulation with an honest bankroll — finite cash, positions lock until resolution, and every model re-evaluation must be justified by a trigger and paid for. What makes it agentic: the analysts and the reconciler are LLMs in a multi-turn tool-use loop — each writes its own search queries, chooses which articles to read in full, and iterates until it submits a schema-validated verdict; nothing hands them a pre-built context blob. The same market-data layer is also exposed as a standard MCP server (Model Context Protocol; mates_mcp — five self-describing, read-only tools: resolution rules, prices, order book, search, market context) that any MCP client can connect to; it deliberately wraps no trading or account calls, so no agent can move money through it. Inside the simulator the identical tool loop runs in-process for speed and leak-tight determinism. Python + Flask, containers on GCP.
DATA IN Kalshi REST API prices · full resolution rules · candlesticks · positions · settlements Web Search current facts — defeats the model's training cutoff Orchestrator — runs the sweep end to end sequences every stage · fans out the agent tiers · retries failures · persists verdicts and funnel stats CONTROLS EVERY STAGE Discover the universe enumerate ~3,800 series per-category → filter price / time-to-close → the tradeable universe ① Deterministic Intake — filter & rank plain code, no model · correlation-aware sibling caps · category round-robin → ranked shortlist ② Managing Director — dispatch prioritizes the shortlist and fans out the research tiers below UNDER CONSTRUCTION ③ Associate Research Agents — cheap triage tier a fast, low-cost LLM pass over a much LARGER candidate set — recent-news aware scores each opportunity so only the strongest earn an expensive Principal agent ④ Principal Research Agents — one LLM agent per market reads the FULL resolution rules (the auto-title is untrusted) and states what resolves YES BEFORE pricing · web-searches current facts · blind to the market price returns a structured verdict: probability, trigger type, nuance, confidence MCP tool layer get_resolution_rules get_market · orderbook search · context read-only Kalshi data; the agent chooses which tools to call Score the edge our probability vs. the live price → conviction from projected ANNUALIZED return · near-certain-trigger guard Alert Queue Cloud SQL · ranked by conviction · live-repriced with drift flags · sectioned by sweep date ⑤ Trading Desk — human triage Take · Watch · Dismiss, then the wager is placed manually on Kalshi — nothing auto-trades Track record positions auto-linked to the prediction that drove them → CLV + calibration (Brier vs. the market) OUTCOMES RECALIBRATE RUNS ON Google Cloud Rundashboard service + scheduled jobs Cloud SQL (Postgres)alerts · predictions · trades · history Secret ManagerAPI keys · RSA-PSS signing Cloud Storage + Registrymarket archive · container images
What's actually an agent: the Principal Research Agents are the real LLM agents today — one per shortlisted market, each web-searching independently and returning a schema-validated verdict. The Associate Research Agents are the cheap triage tier that will let us screen far more markets than 24; that stage is under construction and nothing runs through it yet. Orchestrator, Deterministic Intake and Managing Director are deterministic code, not autonomous agents — a workflow runner, a filter/rank pass, and the fan-out. The Trading Desk is you. Both agent tiers reach market data through a read-only MCP tool layer (Model Context Protocol) — the agent chooses which tools to call, which is what makes the research step agentic rather than fed a fixed context blob. Python + Flask throughout, deployed as containers on GCP.

Active Positions Summarized

Phase 2 complete · 2026-08-24 — Phase 2 of Hemlock's 2026 strategy is done: the Canopy research engine is built and validated in a 12-month walk-forward simulation. We are now liquidating current positions, deploying the new system, and will re-deploy capital into new positions. Banner updated 2026-08-24 (manual) · the positions summary below refreshes on its own schedule
This portfolio totals roughly $2,854 across two venues, with the large majority β€” about $2,550, or 89% of assets β€” sitting as uninvested cash on Polymarket US and no open positions there. The active exposure is concentrated entirely on Kalshi, where two "NO" positions (bets that a given event will not happen) worth about $303 in current market value are held, alongside a negligible $0.60 cash balance. Within Kalshi, the book is highly concentrated: one position, a NO bet on whether Miguel DΓ­az-Canel leaves his Communist Party post before 2027, accounts for roughly 93% of invested cost ($356 of $383), while a much smaller NO position on presidential impeachment before 2028 makes up the rest. Overall, the portfolio is mostly cash, with only a small, concentrated slice actively deployed in two directionally similar (both "NO") event-outcome contracts on one venue.
Summary auto-generated daily from the live portfolio's shape · last refreshed 9/1 2:01 PM PT
MarketSide ContractsCost ValueP&L
Will Miguel DΓ­az-Canel leave First Secretary of the Communist Party of Cuba before Jan 1, 2027? NO 347 $355.99 $294.95 $-61.04
Will the President be impeached before Jan 1, 2028? NO 37 $27.43 $14.43 $-13.00
Total $383.42 $309.38 $-74.04