Backtesting ArenaBacktesting Arena
← Back to blog

Why Experts Forecast Worse Than Chimpanzees

A chimpanzee guessing at random beats educated experts on fact questions — and forecasters on predictions. That's not a punchline but the finding of several of the most-cited studies. Why it happens, what Michael Burry has to do with it, and why a real prediction needs two things: a what and a when.

Backtesting Arena·July 6, 2026·4 min read·8 views
Why Experts Forecast Worse Than Chimpanzees

A chimpanzee pressing answer buttons at random beats educated experts. That's not an insult — it's one of the best-documented findings in forecasting research, and it has direct consequences for who you should believe in financial markets.

The chimp and the Nobel laureates

The Swedish physician Hans Rosling spent years asking highly educated groups — medical students, professors, scientists, investment bankers, journalists, decision-makers — simple fact questions about the world. Thirteen questions, three options each. A chimp choosing at random would score about 33%. Across fourteen industrialised countries, the human average was under two correct out of twelve. Some of the worst results came from Nobel laureates and medical researchers.

The key point isn't that people are stupid. It's that they are systematically wrong. They don't miss in all directions — they lean consistently the same way, too pessimistic. The world got better; the mental model didn't. Dice would score higher, because at least dice aren't biased.

"About as accurate as a dart-throwing chimpanzee"

The psychologist Philip Tetlock ran the hardest test. Over twenty years he tracked 284 people who make their living commenting on political and economic trends, gathering more than 82,000 concrete forecasts. His now-famous verdict: the average expert was about as accurate as a dart-throwing chimpanzee. Many would have done better simply guessing at random.

To be fair: Tetlock himself dislikes the chimp soundbite — the point is that experts barely cleared the random baseline, not that they were hopelessly clumsy. But one detail makes it sting: there was an inverse correlation between fame and accuracy. The more sought-after the expert, the worse the forecasts on average. Because the pundits who make it onto the screen are selected for one thing — telling a confident, simple story. Accuracy is optional. They are rarely in doubt and often wrong.

Markets don't make it better

CXO Advisory graded 6,582 public stock-market forecasts from 68 experts. The average accuracy: 47.4%. A coin flip wins. The best guru managed about 68%, the worst barely 22% — and being a household name didn't help.

On recessions the record is thinner still. An IMF analysis put it bluntly: the failure to predict recessions is "virtually unblemished." Of forty-nine economies in recession in 2009, the number forecasters saw coming a year earlier — in 2008 — was exactly zero. And when they finally took fright, they over-corrected, seeing recessions where there were none.

The core of the problem: what AND when

Why is forecasting so hard? Because a real prediction has two parts: what will happen and when. Most pundit predictions deliver only the what — "a crash is coming" — and quietly drop the when. That's not a forecast. It's a horoscope with a chart.

Without a when, every crash call is eventually right — the way a stopped clock is right twice a day. In markets, "early" costs you exactly what "wrong" does. A prediction you can't time is a feeling, not an edge.

The textbook case: Michael Burry

Michael Burry is the perfect example — precisely because you can tell it fairly. In 2008 he placed a specific, mechanism-based bet against subprime mortgages. He was early, and it hurt; but the thesis was concrete, checkable, and it resolved. "The Big Short" immortalised the trade. That prediction was falsifiable. And it was right.

Since then Burry has called crash after crash — a "greatest bubble ever" in 2021, a "sell" in 2023 — while markets kept rising. Those were open-ended warnings: a what with no mechanism and no clock. To his credit, he owns the misses. The point isn't that Burry is unintelligent — he's brilliant. The point is that one correct call doesn't make a permanent oracle. We remember the movie, not the misses. That's survivorship bias and authority bias in one package.

The good news: forecasting is learnable — but differently

Tetlock didn't stop at the chimp. In a later project he found people with real, measurable forecasting talent — "superforecasters" who beat the professionals. Not through higher IQ. Through method: they framed things specifically, attached a probability and a date to each claim, and corrected themselves the moment they were wrong.

And here the circle closes back to what we build. The mathematics that separates a skilled forecaster from a lucky one — penalising vague claims and many guesses — is the same logic an honest backtest applies. Fittingly, it was Marcos López de Prado, who gave the Deflated Sharpe Ratio its name, who re-graded the famous guru forecasts, weighting them by specificity and time horizon. That same Deflated Sharpe Ratio runs inside our validate_strategy.

What this means for you

Don't trust the confident voice on TV. Trust a claim with a what, a when, a sample size, and a benchmark you can check. That's a backtest. It doesn't predict the future — it makes your assumption falsifiable, timed and measurable. That's how you beat the chimp: measure, don't narrate.

Study the Past — Improve your Future. 🥋

Try it yourself

Run the backtest with your own parameters and time ranges.

Run backtest →

More on this topic

Market Analysis

The Bitcoin 4-Year Cycle Isn't a Law. It's a Superstition — and That's Exactly What Makes It Exploitable.

Backtesting Arenatradingstrategies.work

Every time Bitcoin corrects, the same reflex kicks in: "Relax, it's just the cycle." But where is this cycle written down? In which whitepaper, in which law of nature? Spoiler: nowhere. The 4-year cycle is pattern recognition built on three data points — and a self-fulfilling prophecy that makes exactly those traders exploitable who believe in it most. A critical dissection.

BitcoinMarket structureBacktesting
May 9, 20261 min
Market Analysis

Stocks in Wartime: Why the Recovery Statistics Only Contain the Winners

Backtesting Arenatradingstrategies.work

"Stocks have recovered from every war" is true — for one country. After the April 1940 high the Dow needed four years and nine months to get back to even, in 1914 the exchange was closed for four and a half months, and St Petersburg, Vienna, Tokyo and Shanghai are missing from the statistics because their exchanges ceased to exist. Across 39 markets, a diversified investor is down in real terms after 30 years in 12 % of cases.

MethodologyBacktestingDrawdown+1
Sep 21, 20261 min
Market Analysis

Bitcoin doesn't tip when enough managers understand it. It tips when a zero allocationbecomes the risky choice.

Backtesting Arenatradingstrategies.work

The 25 percent tipping point quoted in every adoption argument comes from a naming game with 18 to 30 players per group and no right answer. It counts people; prices respond to dollars. What tips in large portfolios is not understanding but the answer to one question: what happens to me if I am the only one at zero?

Bitcoin
Sep 12, 20261 min
Market Analysis

Popular Trading Myths, Tested Against the Actual Numbers

Backtesting Arenatradingstrategies.work

Max pain, the yen carry trade, rate cuts, the four-year cycle: seven trading myths, each held against the numbers. The answer is almost never "myth" — usually it is something less comfortable.

BacktestingMarket structureMethodology+2
Aug 6, 20261 min
📬

Don't miss new blog posts

One short email per new post — strategies, backtests, market analysis. No spam, unsubscribe with one click anytime.

By subscribing you accept our privacy policy. We use Resend for delivery. Double opt-in confirmation required.

Comments (0)

Join free to post comments.

Sign up →

No comments yet. Be the first!