In brief

Earnings prediction is one of the toughest games in finance, and the analyst consensus is a genuinely strong opponent — thousands of professionals spend their quarters modeling revenue, margin and EPS, and their averaged forecasts are the baseline the market trades against. Samaya ran each model through a point-in-time gate, forcing predictions one week before earnings with nothing published after the cutoff visible, and scored them on how well they forecast the surprise.

Only the three newest models cleared the harder bias-corrected consensus baseline, which adjusts for analysts' habit of lowballing estimates before earnings. Older models beat raw consensus but not the corrected version. The inflection point, in other words, is recent.

Read the setup before the score

The test design is the reason to take this seriously, and also the reason to read carefully. The four target metrics were revenue, gross margin, operating income and adjusted EPS. The companies all reported after July 14, chosen to sit past the knowledge cutoffs of every model tested, so no one could win from training memory. A programmatic authentication gate blocked access to post-cutoff information.

Samaya then ran ablations, and those are the numbers to quote. Giving the models quarter-start data instead of one-week-before data raised performance by 25 percentage points. Taking data away entirely flattened even the best models; parametric memory of finance is nearly worthless at prediction time. And expert-guided instructions pushed models to research 1.6 to 2.7 times more, with errors falling.

Put plainly: this was not a test of seven naked models. It was a test of seven models inside Samaya's production finance harness — time-gated retrieval, structured data tools, analyst-written prompting. GPT-6 Astra nearly matched its own performance even without seeing consensus at all, which is striking. But the harness is doing load-bearing work in every row of these charts.

Who ran the test

Samaya sells the harness. That does not invalidate the result, but it is the disclosure the headline needs. The study is vendor research, and the component the vendor markets is the component that contributed the most. A fair reading is that Samaya has demonstrated a genuine capability milestone and also written its own sales deck in the same document. Both things can be true.

What it means for the people who do this for a living

Here is the practical reading. The models' edge was strongest on below-consensus calls — the misses, the quarters where being right means disagreeing with the crowd. Samaya's trace analysis suggests the wins came from gathering fresh evidence and being willing to move away from the herd. Analysts lowball and herd; the models, properly supplied with data, do not have to.

That does not mean funds will trade purely on model output tomorrow. Predicting a number more accurately is not the same as predicting what the market will do with it; separate research put AI's share of explaining earnings-day stock moves at 17%, up from 5% — an advance that still leaves most of the move unexplained. But for the research process itself, the implication is hard to dodge. The scarce input is no longer the model. It is the data pipeline and the instructions that keep the model honest. The analyst's job does not disappear. The spreadsheet part of it just got cheaper.

Sources

  1. [1] Samaya — “Frontier AI models outperform human experts on earnings prediction for the first time” (Oct 2026)Read source
  2. [2] Phys.org — “AI models explain 17% of earnings-day stock moves, up from 5%” (July 2026)Read source