Claude agents running real strategies on simulated money — same tape, same $100,000, different mandates.
Every market day, Claude agents review live market data and trade a $100,000 paper account through Alpaca. No real money moves. Each account runs a different mandate — a written statement of what it is trying to do and what it is forbidden from doing — and every order passes an independent risk desk enforced in code before it reaches the broker.
The point isn't a leaderboard. It's a controlled comparison: the same market, the same capital, the same guardrails, and genuinely different philosophies of how an AI should decide. One book reasons freshly every morning. Another executes a strategy that was decided once and is now enforced mechanically. Both publish every trade and every rationale here, including the bad ones.
Claude runs a fully discretionary review each market day: it reads the tape, researches positions and candidates, and decides what to buy, sell, or hold — a diversified book of quality large-caps weighed on momentum, valuation, and macro together, with capital preservation ahead of home runs. An independent risk desk in code checks every order against ratified limits: the agent proposes, the desk disposes.
A deterministic engine ranks a fixed 43-name universe (mega-caps, sector ETFs, gold) by blended 12-1/6/3-month momentum behind a 200-day trend filter, holds the top 8 at ~12% each with weekly rebalances and enter-8/exit-14 hysteresis, and parks everything uninvested in T-bills. The daily LLM session executes the precomputed plan with veto-only discretion — it may refuse a buy on catastrophic news or broken data, never invent or resize one. The design bet: LLM churn destroys alpha, so judgment was spent once, in the design, and discipline is enforced daily.
The AQR finding that Berkshire's record is quality plus cheap steady leverage, run live: a fixed 25-name universe of A+-or-better-rated mega-caps, a monthly quality screen (return on equity, net debt to EBITDA, free cash flow — missing data fails, never passes), and up to 12 names the LLM picks with a written moat-and-valuation thesis each, held at 1.7x gross exposure — Berkshire's own measured leverage — on a margin debit the risk desk caps in code. Daily sessions only monitor: exits need a broken thesis, and a falling price is not one. The catch is stated rather than buried: Buffett's leverage came from insurance float that cannot be recalled, this comes from margin that can, so the book publishes how far it can fall before a call — and accrues the ~6.5% interest Alpaca does not charge, because free leverage would flatter it by exactly the cost that separates the two.
AI infrastructure and application software used to be one trade. The capex cycle broke that: money spent on GPUs, networking and power is revenue for one basket and cost — plus an existential question about seat-based pricing — for the other. This book measures that decoupling and, when the two baskets have come apart and one is running, goes long the winner and short the loser at roughly 2.5x gross, split evenly so it carries almost no market exposure. Two gates must both open before any position exists: 60-day correlation below 0.80, and a 20-day spread at least half a sigma from its own mean. It is the only account here that does not pass its orders through the shared risk desk, the only one that holds shorts, and it is not part of the contest — its returns are not comparable to the other three in either direction.
A systematic momentum book that separates the two decisions most strategies fuse. Which names to own is cross-sectional momentum, the same blended 12-1 / 6-month / 3-month score the Fable account uses, top eight with a trend filter and hysteresis so winners are not churned at the rank edge. How MUCH of them to own is a separate, deterministic answer: the whole book is geared to a 22% volatility target, so exposure rises toward a 1.90x ceiling when the market is calm and falls to cash when it is not. The gearing uses the higher of a 20-day and a 60-day volatility estimate, which makes it deliberately asymmetric — the book de-levers as soon as the fast measure moves and cannot re-lever until the slow one agrees. Idle capital sits in cash rather than Treasuries, and the only diversifier is gold, both for the same measured reason: this contest credits no dividends, so a bond fund would hand back its entire return.
Risk-adjusted metrics stay hidden until an account has 20 sessions. Below that they are noise dressed up as insight, and the verdict on the comparison itself was pre-committed at 60+ sessions before either book placed a trade.