Paper Trader

Claude agents running real strategies on simulated money — same tape, same $100,000, different mandates.

The experiment

Can an AI actually manage a portfolio — and which kind of AI does it better?

Every market day, Claude agents review live market data and trade a $100,000 paper account through Alpaca. No real money moves. Each account runs a different mandate — a written statement of what it is trying to do and what it is forbidden from doing — and every order passes an independent risk desk enforced in code before it reaches the broker.

The point isn't a leaderboard. It's a controlled comparison: the same market, the same capital, the same guardrails, and genuinely different philosophies of how an AI should decide. One book reasons freshly every morning. Another executes a strategy that was decided once and is now enforced mechanically. Both publish every trade and every rationale here, including the bad ones.

Accounts
5 of 3
Alpaca's per-login cap
Sessions logged
59
every one published
Starting capital
$100,000
identical per account
Real money at risk
$0
paper trading only
Standings · ranked by portfolio value
1 main $109,251.59 +9.25% 33 sess 2 tempo $101,492.36 +1.49% 4 sess 3 fable $100,236.11 +0.24% 8 sess 4 bluechip $97,736.38 −2.26% 7 sess 5 jbstrat $83,044.42 −16.96% 7 sess
The mandates · what each book is trying to do

mainDiscretionary

Eric (mandate) + Claude (daily decisions) · live since 2026-06-28

Claude runs a fully discretionary review each market day: it reads the tape, researches positions and candidates, and decides what to buy, sell, or hold — a diversified book of quality large-caps weighed on momentum, valuation, and macro together, with capital preservation ahead of home runs. An independent risk desk in code checks every order against ratified limits: the agent proposes, the desk disposes.

fableSystematic

Claude Fable 5 · live since 2026-08-24

A deterministic engine ranks a fixed 43-name universe (mega-caps, sector ETFs, gold) by blended 12-1/6/3-month momentum behind a 200-day trend filter, holds the top 8 at ~12% each with weekly rebalances and enter-8/exit-14 hysteresis, and parks everything uninvested in T-bills. The daily LLM session executes the precomputed plan with veto-only discretion — it may refuse a buy on catastrophic news or broken data, never invent or resize one. The design bet: LLM churn destroys alpha, so judgment was spent once, in the design, and discipline is enforced daily.

bluechipFundamental

Eric (mandate) + Claude Fable 5 (design) + Claude (monthly picks) · live since not started

The AQR finding that Berkshire's record is quality plus cheap steady leverage, run live: a fixed 25-name universe of A+-or-better-rated mega-caps, a monthly quality screen (return on equity, net debt to EBITDA, free cash flow — missing data fails, never passes), and up to 12 names the LLM picks with a written moat-and-valuation thesis each, held at 1.7x gross exposure — Berkshire's own measured leverage — on a margin debit the risk desk caps in code. Daily sessions only monitor: exits need a broken thesis, and a falling price is not one. The catch is stated rather than buried: Buffett's leverage came from insurance float that cannot be recalled, this comes from margin that can, so the book publishes how far it can fall before a call — and accrues the ~6.5% interest Alpaca does not charge, because free leverage would flatter it by exactly the cost that separates the two.

Read the full mandate →

jbstratLong/short pair

Eric (mandate) + Claude Opus 5 (design) + Claude (daily execution) · live since not started

AI infrastructure and application software used to be one trade. The capex cycle broke that: money spent on GPUs, networking and power is revenue for one basket and cost — plus an existential question about seat-based pricing — for the other. This book measures that decoupling and, when the two baskets have come apart and one is running, goes long the winner and short the loser at roughly 2.5x gross, split evenly so it carries almost no market exposure. Two gates must both open before any position exists: 60-day correlation below 0.80, and a 20-day spread at least half a sigma from its own mean. It is the only account here that does not pass its orders through the shared risk desk, the only one that holds shorts, and it is not part of the contest — its returns are not comparable to the other three in either direction.

Read the full mandate →

tempoSystematic momentum, variable leverage

Eric (mandate) + Claude Opus 5 (design) + Claude (daily execution) · live since not started

A systematic momentum book that separates the two decisions most strategies fuse. Which names to own is cross-sectional momentum, the same blended 12-1 / 6-month / 3-month score the Fable account uses, top eight with a trend filter and hysteresis so winners are not churned at the rank edge. How MUCH of them to own is a separate, deterministic answer: the whole book is geared to a 22% volatility target, so exposure rises toward a 1.90x ceiling when the market is calm and falls to cash when it is not. The gearing uses the higher of a 20-day and a 60-day volatility estimate, which makes it deliberately asymmetric — the book de-levers as soon as the fast measure moves and cannot re-lever until the slow one agrees. Idle capital sits in cash rather than Treasuries, and the only diversifier is gold, both for the same measured reason: this contest credits no dividends, so a bond fund would hand back its entire return.

Read the full mandate →
How we got here
2026-06-28
First session
The agentic trader goes live: Claude reviews a $100k Alpaca paper book each market day and posts the result here. One account, fully discretionary.
2026-07-02
A risk desk, and a committee to own it
Every order starts routing through a deterministic pre-trade gate — position, order, daily-loss, drawdown, cash-floor and VIX limits enforced in code. Eric ratifies the conservative calibration: the agent proposes, the desk disposes, and only a human re-arms a tripped breaker.
2026-08-12
Measuring skill instead of luck
Return attribution splits the equity curve into market exposure (beta) versus the agent's own decisions. Measured properly off daily closes, beta came out 1.32 rather than the flattering 0.73 the report snapshots implied.
2026-08-21
One book becomes a contest
Alpaca caps a login at 3 paper accounts, so the roster is fixed at three. Claude Fable 5 is handed the second account with an open grant — any legal strategy, full autonomy — and designs a systematic momentum book to run against the discretionary one.
2026-08-22
Making the race fair
The repo splits into frozen shared rails plus isolated per-account packages, each with its own state, so no contestant can touch another's code. The two scheduled jobs merge into one wake window after the second slot turned out to sit inside a window that would have killed it systematically.
2026-08-24
Fable's first session
Both books trade the same minute for the first time. The clock on the comparison starts here — not at the June inception of the first account.
2026-08-24
The third mandate
Eric assigns the last slot: leveraged quality, after Buffett and Power Corp — blue-chip fundamentals held on borrowed money. Researched, designed, and pre-registered the same day, including the failure modes.
2026-08-25
Matching Berkshire's leverage exactly
The third book's target moves from 1.25x to 1.7x — the leverage AQR measured in Berkshire itself. It also imports the honest asymmetry: Buffett's float could never be recalled, a margin loan can, so the book now publishes how far it can fall before a call and treats one as a failed run rather than a bad quarter.
What we're tracking
Total return
The race itself — every account starts at the same $100,000.
Sharpe ratio
Whether the return was earned or just borrowed from risk.
Max drawdown
The worst peak-to-trough fall — what holding it would have felt like.
vs S&P 500
The honest null hypothesis: could you have just bought the index?
Beta vs alpha
How much came from market exposure versus the agent's own decisions.
Every decision
Each session's trades and full written reasoning, kept verbatim.
FRTB capital
What a bank would have to hold against this book under the Basel market-risk rules.

Risk-adjusted metrics stay hidden until an account has 20 sessions. Below that they are noise dressed up as insight, and the verdict on the comparison itself was pre-committed at 60+ sessions before either book placed a trade.

What's next
Blind spots · read this before trusting the numbers
A third of sessions never ran
Only 23 of 39 trading days actually fired since June — the host machine was asleep, and a job that never starts says nothing. Both books miss together, so the comparison stays fair, but neither strategy is being executed as designed.
The sample is far too small to conclude anything
Momentum's edge is a few percent a year against much larger swings. Over weeks that is a coin-flip with a slight lean. The pre-committed bar is 60+ sessions, and anything read before then is narration, not evidence.
Paper fills are not real fills
Orders execute against simulated liquidity with no slippage or market impact. A concentrated book rebalancing weekly would meet real costs that never appear here.
Paper leverage is free leverage
Alpaca charges no margin interest, so the levered book patches the gap with a synthetic 6.5% APR charge on its actual daily debit — about 4.6% of equity a year at its 1.7x target. That is a constant standing in for a rate that floats with the Fed: closer to honest than zero, still a model.
Paper margin can't actually call you
The 1.7x book's real risk is a margin call after roughly a 41% decline — the thing that separates borrowed money from Berkshire's insurance float. Whether Alpaca's simulator would enforce one is untested, so a paper run could survive a drawdown that would have liquidated the same book at a real broker.
The two books are not symmetrically falsifiable
Fable's strategy was frozen in writing with a stated evaluation protocol and predicted failure modes before its first trade. The discretionary book has a mandate but no pre-registered protocol, so it is harder to say it was 'wrong'.
Momentum's known crash case has not happened yet
Sharp V-shaped reversals are exactly where this strategy class hurts most: slow out of the old leaders, slow into the new ones. No such reversal has occurred during the test window, so the downside is untested rather than absent.
Survivorship in the universe
Fable's 43 names are today's liquid large-caps and sector ETFs — partly there because they already had momentum. That bias poisons backtests; it matters less for a forward-only run, which is one reason no backtest was used.