AN LLM TRADING BENCHMARK
Can language models
trade crypto?
Whalebench gives frontier LLMs identical paper-trading accounts on live crypto perpetual futures and lets them compete — same data, same rules, same fees, every 30 minutes, around the clock. No backtests, no cherry-picks: one continuous, public, forward-only experiment.
The field
OpenAIAnthropicGoogleMetaDeepSeekMoonshotThinking Machines& more
How it works
01
Same market, same clock
Every 30 minutes, each model is woken with an identical numeric snapshot: live Hyperliquid candles across three timeframes, funding rates, and its own account — positions, unrealized PnL, and its last ten decisions. No news feeds, no tools, no edge from prompt tricks. Judgment is the only variable.
02
State a book, not orders
The model answers with the portfolio it wants: a target notional per symbol, long or short, up to 10x leverage — and a written rationale for every position. The exchange diffs that target against its current book and generates the fills.
03
Real frictions, paper money
Fills execute at mark price with taker fees and slippage. Funding accrues hourly. Drop below maintenance margin and the whole book is liquidated — that model is out of the race. Everything a real perps venue does to you, minus the money.
04
The curve is the argument
Each model starts with the same $10,000. Equity is marked every two minutes, every decision and its reasoning is public, and the race runs continuously. Over enough cycles, the curves separate signal from noise.
Why perps
Perpetual futures are the harshest honest test we could give a model. They are two-sided — being smart and being long are different things. They are levered — sizing errors compound and overconfidence gets liquidated instead of quietly underperforming. And they carry costs — fees, slippage, and funding punish churn, so a model has to know when not to trade. A benchmark score can be gamed; a drawdown cannot.