Skip to content
Raynnier Gomez
ArticleAugust 26, 2026
All posts

AI is ‘printing money’ in crypto. Fifty pre-registered experiments disagree.

Published
August 26, 2026
Updated
September 1, 2026
Reading time
5 min
Topics
quant research · crypto · AI trading · statistics

There is a genre of screenshot going around: an AI model, a trading account, a green number. Claude allegedly made someone ten thousand dollars overnight; a bot is up 22% in two weeks; a portfolio ‘beats the S&P’. Meanwhile, on a small computer running controlled experiments against real crypto market data, the same question keeps coming back with the same answer. This page is about that gap — and it will keep being about it, in public, with sources.

01What the viral numbers actually say

The most cited evidence is Alpha Arena, a live benchmark that handed six language models $10,000 each of real money to trade crypto perpetuals in autumn 2025 [1]. Real capital, real markets, fully autonomous — credit where due, it is a genuine experiment. The final board: Qwen3 Max finished around +22%, DeepSeek slightly positive, and everyone else lost money. Claude Sonnet 4.5 gave back more than 40% of its stake, and at one point Western flagship models were down 80% inside a week [3][4].

Then somebody did the unglamorous thing and ran the statistics. An independent replication tested the model-versus-model comparisons for significance and found that exactly one of six cleared p<0.05 [2]. One. The industry read the leaderboard as ‘AI can trade’. The arithmetic reads it as a small tournament in which somebody had to finish first [5][6].

The rest of the genre is thinner. A ‘$100K experiment’ in which three of five AI-written strategies beat the S&P over ninety days — five strategies, one quarter: that is a hand of cards, not a track record [8]. A ‘Claude portfolio’ up 19% versus the index, quoted everywhere except next to the fact that it is the sales funnel for a copy-trading product — the return is the advertisement [9]. And an ocean of screenshots with no benchmark, no costs, and no way to count how many losing screenshots were never posted.

A 22% fortnight is not evidence of skill. It is a coin flip with good lighting.

02Why short windows prove nothing

Market returns are noisy enough that skill and luck are statistically indistinguishable over weeks. The error bars on a performance estimate shrink with the square root of time, and over a two-week window they are wider than any effect being claimed. Put eight models in a bracket and one of them will ‘win decisively’ every single time. That is not intelligence; that is what brackets do.

Add three quiet biases and the genre writes itself. Survivorship: losing accounts do not tweet. Beta dressed as alpha: being long anything during a rising market produces a green number that required no intelligence, artificial or otherwise. And costs: fees, slippage and funding quietly convert a good many paper edges into real losses.

None of which proves the winners were lucky. It proves nobody can tell yet — and ‘nobody can tell’ does not make the thumbnail.

03A different way to ask the question

The experiment this page follows runs on a Raspberry Pi: an autonomous research system, pointed at the fifteen most liquid perpetual futures on one exchange, with five years of minute-resolution price data. Deliberately modest hardware, deliberately liquid markets — if an edge only exists on a pair too thin to actually trade, it is not an edge, it is a rounding error.

The rules are the point. Every hypothesis is pre-registered before any code is written: the economic mechanism, the exact thresholds that would validate it, and the result that would falsify it, all fixed in advance. To survive, a strategy must clear a permutation test at p<0.01, hold up in walk-forward out-of-sample data, survive the deflated Sharpe ratio — the correction for ‘you tried many things and published the best one’ [7] — and beat the dumbest possible alternative: holding the same assets equal-weighted, net of costs. Miss one gate and the hypothesis is retired, in writing.

The mission is symmetric, and it is written into the system: a validated edge, or a well-documented conclusion that there is none. Both count as success. That sentence does more work than it looks — it removes the incentive to torture the data until it confesses.

04Fifty hypotheses later

The score so far: fifty-one hypotheses pre-registered, fifty-one falsified. Momentum, reversal, beta-neutral idiosyncratic momentum, an illiquidity premium, funding-rate designs, order-book imbalance — eight distinct economic mechanisms in the one-to-seven-day price-and-volume family, each tested with real statistical power, each dead with its p-value on record.

Along the way the system measured the backdrop the screenshots never mention: long-only crypto in this universe runs a Sharpe ratio around 0.44 with drawdowns near 80%. That is the asset class. A bot that was long through a good stretch ‘made money’ the way standing in the rain gets you wet.

In August the family was formally closed rather than testing a ninth variation, because a ninth variation with no new economic reason is exactly the p-hacking the rules exist to prevent. It has since reopened exactly once, for a mechanism from the academic literature that did bring a new reason — the field reports below carry the autopsy. The most useful sentence in quantitative research is ‘this does not work, and here is the proof’. It is also the least retweeted.

05What would change the answer

Three things, all in motion. First, statistical power: microstructure signals — open interest, positioning ratios, order-book depth, liquidation flow — need about a year of history before a test on them means anything, and most of those collectors are months in. A recent find shortens the wait considerably: the exchange’s public archive carries open interest and positioning at five-minute resolution back to 2020, which unlocks years of test power for the next family once it is verified against live capture.

Second, a different substrate. Everything falsified so far is price and its derivatives — the most mined dataset in finance. Information is a different question: news, filings, flows, the things an analyst actually reads. Whether a language model has real edge there is being measured under the same discipline — theses registered before the outcome, scored on calibration, forward-only, because an LLM cannot be honestly backtested on a past it has memorized.

Third, simply: evidence. If anything ever clears the gates, it will be published here with the thresholds it had to clear stated up front — which means that, for once, you will be able to check.

06What this page becomes

Below this article, field reports will accumulate: roughly weekly when there is something real, quieter when there is not. Findings, falsifications, and the occasional autopsy of a viral claim, each with its sources listed at the end.

If a screenshot says a model turned $10,000 into $12,000, the interesting question is never the screenshot. It is: over how many independent bets, against which benchmark, net of which costs, selected from how many attempts nobody posted? That question also fits in a tweet. It just never gets asked in one.

Sources

Numbered sources the body points at. The order is fixed at publication — it is part of the text.

  1. Nof1
  2. GitHub — Eug-Chua/llm_trading_arena
  3. Bitcoin Magazine
  4. iWeaver
  5. Boris Again (Substack)
  6. arXiv
  7. SSRN — Bailey & Lopez de Prado
  8. PickMyTrade
  9. explainx.ai
Field reports

Findings from the ongoing investigation, newest first. Published when there is something real to report — not on a content schedule.