August 28, 2026·
Research|Perspective

AI Financial Research Benchmarks Explained: What the 2026 Scores Mean

Anwaar MalikAnwaar Malik
A row of laboratory gauges and measuring instruments on a bench, each dial reading a slightly different value

The short answer: The 2026 finance benchmarks agree: the best AI systems get roughly half to two thirds of expert-written research questions right, and professional analysts still write better reports. Vals Finance Agent v2 tops out at 60.60% (August 19, 2026) and Rogo's BigFinanceBench at 67.1% by rubric and 53.4% by final answer (August 3, 2026). In JPMorganChase's Deep FinResearch Bench (April 2026), analysts outscored the best agent 2.84 to 2.31 on report quality, and agent factuality ran from 86.0% down to 53.2%. Every benchmark tests a bare model with EDGAR and web search, so none measures an entitled platform, an audit trail or a firm's own data.

Who this is for: heads of research and CIOs handed a vendor deck with a benchmark score on it, model-risk officers documenting an AI approval, and analysts deciding how much of an AI draft to re-check.

Published August 28, 2026. Last reviewed August 28, 2026. Written by the AllMind AI research team.

Reviewed by Anwaar Malik, founder of AllMind AI.

Disclosure: AllMind AI builds a research platform and has published no score on any benchmark in this article. Two of the benchmarks below are run by competitors (Rogo and Samaya); we say so in the table and take their figures at face value.

How we read these figures: every number comes from the benchmark's own paper, leaderboard or press release, fetched on August 28, 2026, and no aggregator tables are quoted, because one widely copied Vals v2 table did not match vals.ai that day.

Key takeaways

  • Top scores cluster between 56% and 67%. Vals v2 60.60% (Muse Spark 1.2, August 19, 2026), BigFinanceBench 67.1% rubric and 53.4% final answer (Muse Spark 1.1, August 3, 2026), FrontierFinance 56.0% (Samaya, high-effort configuration, August 12, 2026).
  • Analysts still win on reports. JPMorganChase's Deep FinResearch Bench (April 22, 2026) scored professional reports 2.84 against 2.31 for the best agent on a four-point scale, with agent factuality between 86.0% and 53.2%.
  • Accuracy falls as the task widens. Fin-RATE (Yale and co-authors, v4 June 10, 2026) measured accuracy dropping 18.60% from single-document reasoning to longitudinal tracking and 14.35% to cross-entity comparison.
  • The 81% FinanceBench figure is from 2023. GPT-4-Turbo with a retrieval system incorrectly answered or refused to answer 81% of a 150-case sample in November 2023; Vals v1.1 reached 64.37% correct on harder questions by June 2026.
  • AI help widened coverage and raised forecast error in real reports. A December 2025 working paper using FactSet's 2023 AI launch found 40% more distinct sources cited and forecast errors up 59%, worst for analysts covering more firms.

AI financial research benchmarks explained: the 2026 table

AllMind AI reads these benchmarks the way a research head should: as tests of a bare model with a filing archive and a search box, run by parties with different incentives, on questions that stop at the public record. Four of the benchmarks are run by companies (Vals AI and Patronus AI sell evaluation; Rogo and Samaya sell research products), three by bank, academic or consortium researchers, and the last row is a natural experiment on real broker reports.

BenchmarkWho runs itWhat it measuresSampleBest score (system, date)Honest limitation
Deep FinResearch Bench (arXiv 2604.21006)JPMorganChase AI ResearchReport quality (1 to 4), forecast SMAPE on six line items, factuality of claims100 professional reports, 25 S&P 500 companies, FY2025 Q1 to Q2Analysts 2.84; Gemini 2.31; factuality OpenAI 86.0% (April 22, 2026)Four deep-research agents only; no research platform tested
Vals Finance Agent v1.1 (vals.ai)Vals AI, with Stanford researchers and a G-SIBEntry-level analyst tasks: retrieval, market research, projections537 questions, 21 modelsClaude Opus 4.7 64.37% (June 4, 2026)Tools are EDGAR and Google Search; no broker or expert content
Vals Finance Agent v2 (vals.ai)Vals AINine categories from general quantitative to financial modeling927 expert-reviewed questions, 23 modelsMuse Spark 1.2 60.60% (August 19, 2026)Scores not comparable to v1.1; private and held-out splits
BigFinanceBench (bigfinancebench.com)Rogo (a research vendor)Open-ended research tasks graded by rubric and by final answer928 tasks, 15,656 rubric criteria, 28 modelsMuse Spark 1.1 67.1% rubric, 53.4% final answer (August 3, 2026)Judges are two LLMs; public release is a 50-question subset
FrontierFinance (arXiv 2608.11683)Samaya AI (a research vendor)Investor-workflow queries scored against source-attributed rubrics220 queries, 11,543 rubricsSamaya high-effort 56.0%; Claude Fable 5 49.2% (August 12, 2026)Samaya built the test and scored its own system
Fin-RATE (arXiv 2602.07294)Yale and co-authorsSingle-document, cross-entity and longitudinal reasoning over SEC filings; error attribution17 LLMsAccuracy falls 18.60% (longitudinal) and 14.35% (cross-entity) from the single-document baseline (v4, June 10, 2026)Reports deltas; no single leaderboard number
FinanceBench (arXiv 2311.11944)Patronus AI and co-authorsOpen-book question answering over filings with evidence strings10,231 questions; 150-case sample scored, 16 configurationsGPT-4-Turbo with retrieval: 81% incorrect or refused (November 20, 2023)2023-era models; refusals counted with wrong answers
ICBCBench (arXiv 2606.17458)ICBC-led consortium, 50+ experts at 40+ organizationsObjective questions plus long-form deep-research reports, English and Chinese180 tasks (120 public, 60 held out)DeerFlow with GPT-5.5 58.67 (English track, June 16, 2026)Half the public tasks are Chinese-language; no US broker workflow
Generative AI for Analysts (arXiv 2512.19705)Tsinghua and Xiamen researchersReal broker reports before and after FactSet's 2023 AI launch52,428 US equity reports, 24 brokeragesSources cited up 40%; forecast errors up 59% (December 12, 2025)Observational; measures analysts using AI, no model score

Two reading rules apply to every row. A rubric score runs roughly 16 percentage points above a final-answer score, on Rogo's own May 27, 2026 account. And each score belongs to the model named on that date; a vendor running that model inside its product inherits the ceiling, and its retrieval layer decides how far below it the product lands.

What is Deep FinResearch Bench, and what did analysts beat the agents on?

Deep FinResearch Bench is a JPMorganChase AI Research paper (arXiv 2604.21006, April 22, 2026) that asked whether a deep-research agent can write the report an analyst writes. The researchers took 100 pre-earnings reports from two financial institutions on 25 S&P 500 companies in information technology, financials and health care (fiscal 2025 Q1 and Q2), had the OpenAI, Gemini, Grok and Perplexity deep-research agents write reports on the same names, and graded both sets three ways:

  • Report quality on a four-point scale (1 poor, 4 excellent). Analysts scored 2.84; the best agent, Gemini, scored 2.31.
  • Forecast accuracy as SMAPE across revenue, EBITDA, operating income, net income, free cash flow and EPS. Analysts 17.14%; Grok 17.49%; OpenAI 21.52%.
  • Factuality, the share of claims that can be externally verified: one LLM extracts each claim, a second with web access checks it. OpenAI 86.0%, Perplexity 75.6%, Gemini 69.6%, Grok 53.2%.

The analysts won on the report and roughly tied on the numbers: the agents forecast about as well as professionals and lost on synthesis, structure and the share of claims a checker could stand behind. The rankings also flip between measures: Gemini wrote the best-graded reports and had the second-lowest factuality, 16 points below OpenAI, so a single headline score would have hidden that gap.

Vals Finance Agent benchmark results 2026: v1.1 against v2

Vals AI publishes two live leaderboards on different question sets, so a Vals score quoted without its version says nothing. Version 1.1 (537 questions, updated June 4, 2026, 21 models) is led by Claude Opus 4.7 at 64.37%, Claude Sonnet 4.6 at 63.33% and Muse Spark at 60.59%. Version 2 (927 expert-reviewed questions across public, private validation and held-out test splits, updated August 19, 2026, 23 models) is led by Muse Spark 1.2 at 60.60%, Claude Opus 5 at 58.63% and Gemini 3.5 Flash at 57.86%.

The v1.1 page states the scope plainly: tasks expected of an entry-level financial analyst, built with Stanford researchers, a global systemically important bank and industry experts. The paper behind it (arXiv 2508.00828) gives models EDGAR access and Google Search, and at publication the best model, OpenAI o3, scored 46.8% at about $3.79 per query. The June 2026 leader, at 64.37%, is roughly 18 points above that on the v1.1 set.

The v2 category leaders map to jobs:

v2 category (August 19, 2026)Leading modelScore
General quantitativeGemini 3.7 Flash81.8%
Earnings analysisGemini 3.7 Flash79.1%
General qualitativeGrok 4.677.8%
Market analysisGemini 3.7 Flash74.5%
Disclosure analysisClaude Opus 571.3%
AdjustmentsMuse Spark 1.256.3%
ComparablesGemini 3.7 Flash50.3%
PrecedentsGemini 3.5 Flash36.4%
Financial modelingMuse Spark 1.234.5%

A question answered from one filing (earnings analysis, 79.1%) is in reach. A question that joins several companies or several periods (comparables 50.3%, precedents 36.4%, modeling 34.5%) fails more often than it succeeds, the same gradient Fin-RATE measured as deltas. One sourcing caution: the overall v2 scores for Claude Fable 5, Claude Opus 4.8 and GPT-5.5 were not readable from the rendered vals.ai page on August 28, 2026, so we print none.

How accurate is AI at financial analysis? Reading the factuality rates

Right on roughly six in ten entry-level analyst tasks over the public record (Vals v1.1 and v2, June and August 2026), and measurably worse than a professional once the task spans companies, periods or a full report. The cleanest way to turn published rates into a decision is to apply them to a team's real output, using only the papers' figures.

The team shape. A 12-analyst long-only research team, each analyst covering 25 names, one earnings preview per name per quarter: 300 previews a quarter, each carrying about 50 checkable claims (consensus figures, segment trends, guidance history, peer comparisons, the house forecast on six line items). The question put to the agent: "Draft the earnings preview for this name: consensus revenue, EBITDA and EPS, the three items to watch on the call, and our forecast against consensus, each figure sourced."

Applying Deep FinResearch Bench's factuality rates. Factuality is the share of claims an independent checker could verify, so 100% minus the rate is the share the reviewing analyst must re-source or strike.

Agent factuality (Deep FinResearch Bench, April 2026)Unverifiable claims per 50-claim previewPer quarter, 300 previews
86.0% (OpenAI)7.02,100
75.6% (Perplexity)12.23,660
69.6% (Gemini)15.24,560
53.2% (Grok)23.47,020

At two minutes to re-check one claim (an assumption; time your own reviewers), the best case is 70 review hours a quarter across the team and the worst case is 234 hours, before anyone reads the draft for judgment. That is the figure to hold up against a vendor's time-saving claim.

Applying the forecast findings. On the six forecast items the agents were close to the analysts (SMAPE 17.49% and 21.52% against 17.14%), so the house numbers are the part of the preview an agent hurts least. The December 2025 working paper on FactSet's AI launch adds the field result across 52,428 real US equity reports: 40% more distinct sources cited and forecast errors up by roughly $0.44, or 59% of the sample average, worst for analysts covering more firms. A 25-name load is that case.

What changes the arithmetic. Those factuality rates were earned by agents with web search over the public record; a research system attacks the unverifiable-claim count with mechanism, and the mechanism is what a buyer should inspect. On AllMind AI, each claim in a drafted report is cited back to the passage it came from, so the reviewer's re-check is a click to the source.

A peer comparison traverses a maintained ontology in which filings, transcripts, consensus estimates and the firm's own models are stored as linked entities. A preview that needs the team's own numbers reads them where they already sit, in the firm's Snowflake, Databricks or S3, under a role scoped to that data. None of that has been scored on a public benchmark, which is why this article ends with a template for scoring it yourself.

What does FinanceBench measure, and why the 81% figure is misread

FinanceBench measures whether a language model can answer clear-cut questions about a public company from its filings, and it is the oldest benchmark still quoted in 2026 vendor decks. Patronus AI and co-authors published it on November 20, 2023 (arXiv 2311.11944): 10,231 questions about publicly traded companies, each with an answer and an evidence string. The headline result tested 16 configurations on a 150-case sample (2,400 manually reviewed answers): GPT-4-Turbo used with a retrieval system incorrectly answered or refused to answer 81% of questions.

Three things get lost when that sentence is shortened to "GPT-4 got 81% wrong":

  1. The models are from 2023. GPT-4-Turbo, Llama 2 and Claude 2 were the configurations. Vals v1.1, on entry-level-analyst tasks, reached 64.37% correct by June 4, 2026.
  2. Refusals count. The metric folds "would not answer" into "answered wrongly", so much of the 81% is retrieval that never found the passage.
  3. The 81% is the 150-case sample. A deck that says "81% of 10,000 questions" is inventing a denominator.

What FinanceBench still measures well is the floor: a system that cannot retrieve the right passage from a 10-K cannot answer the question, however good its reasoning. Fin-RATE attributes each error to retrieval, generation, domain reasoning or query misreading; for the practical version of the same question see can ChatGPT analyze a 10-K.

Reading the vendor-run benchmarks: BigFinanceBench and FrontierFinance

Two of the most-cited 2026 leaderboards are run by companies that sell research products. The numbers are usable; the ownership belongs in the same sentence as the score.

BigFinanceBench (Rogo)

BigFinanceBench is Rogo's benchmark, announced May 27, 2026 and published as arXiv 2606.03829 on June 2, 2026: 928 expert-authored open-ended research tasks, 15,656 rubric criteria and 36,241 weighted rubric points. At publication the best single model (Claude Opus 4.7 or GPT-5.5) reached 58.8% by rubric. The live leaderboard at bigfinancebench.com, updated August 3, 2026 with 28 models, has Muse Spark 1.1 first at 67.1% rubric score and 53.4% final-answer accuracy, then Gemini 3.5 Flash (65.4% rubric, 42.2% final answer) and Claude Opus 5 (61.8%, 46.1%). Judging is a two-judge mean of Gemini 3.1 Pro and Claude Opus 4.7.

Where it wins: it grades the process of a research task, closer to a memo than a quiz, and it publishes both columns so the roughly 16-point rubric-to-answer gap stays visible.

Where it falls short: the public release is a 50-question subset, the judges come from two of the labs being ranked, and Rogo sells a product built on the models it ranks. Treat the rubric column as a bare model's ceiling and the final-answer column as what a reviewer would sign.

FrontierFinance (Samaya)

FrontierFinance is Samaya AI's benchmark, released July 29, 2026 and published as arXiv 2608.11683 on August 12, 2026: 220 expert-crafted queries and 11,543 source-attributed rubrics across six use cases in the paper (the press release lists five). The abstract reports Samaya's in-house system at 56.0%, Claude Fable 5 at 49.2%, and the best open-weight model, Kimi K3, at 46.4%. Samaya's blog gives a second Samaya number, 50.8%, which the release identifies as a low-effort configuration against the 56.0% high-effort one. Always name which.

Where it wins: the hardest use cases are published (screening and discovery at 33%, sector and macro at 39% for the best system), and the rubrics are source-attributed. Samaya also states that SEC filings account for only 39% of rubrics, an admission that a filings-only system cannot pass.

Where it falls short: the vendor wrote the questions, wrote the rubrics and finished first. That is a fair product demonstration and no independent score; quote the frontier-model rows and ask Samaya for a run on your own questions.

What none of the benchmarks measure

Every benchmark above tests a model with a filing archive and a search engine over the public record, so five things a research head pays for sit outside all of them:

  • Entitled content. No benchmark includes broker research or expert-call transcripts, because the test sets are open.
  • Lineage. Deep FinResearch Bench checks factuality after the fact with a second model; no benchmark scores whether the system showed the passage each number came from as it wrote.
  • Internal data. No test joins a firm's own models, positions or prior notes to the public record.
  • Multi-day workflows. The tasks are single questions or reports. An agent that monitors a coverage list against each new filing for a quarter, the job AI research agents are bought for, has no benchmark.
  • Entitlement inheritance. Whether an agent running for a walled-off analyst can retrieve a restricted name is a compliance test no leaderboard runs.

The survey figures in our asset management AI statistics show licensing and onboarding as the reported barriers, ahead of model accuracy. The earnings-call summary accuracy piece covers the one task where a firm can test lineage directly against a transcript.

Where AllMind AI sits against these benchmarks

AllMind AI has published no score on Vals, BigFinanceBench, FrontierFinance or a Deep FinResearch-style factuality test, and a buyer should say so back to us. What the product publishes is the mechanism the benchmarks cannot see. Drafted reports cite every claim to its underlying passage and a verification pass re-checks figures before a report ships. An agent runs under the entitlements of the analyst who launched it, no wider, and each access is logged. The corpus it traverses spans FactSet, S&P Global, LSEG and MSCI data, broker research, live earnings, and expert-call transcripts through the Expert Insights class, in customer use since August 2026.

Where it wins: the three Deep FinResearch Bench measures can be run on a real coverage list by the firm's own reviewers, because each output figure links to its source and each access is logged.

Where it falls short: without a public score, the burden of proof is on the run, and it is a test AllMind AI has to pass in front of you.

How to run your own 20-question benchmark

Budget one analyst for about two days to build a 20-question in-house benchmark and half a day per vendor to score it. It answers what no leaderboard can: how the system performs on your names, your entitled content and your review standard. Keep the answer key private and re-run it at every renewal.

IN-HOUSE RESEARCH BENCHMARK, v1 (20 questions, scored per vendor)

SET-UP
  coverage_names        5 names from the live coverage list, at least 2 non-US
  documents_in_scope    latest 10-K/10-Q or annual report, last 2 transcripts,
                        1 entitled broker note, 1 internal model per name
  answer_key            written by the covering analyst before any vendor run;
                        each answer carries: value, source document, page/passage
  reviewer              a second analyst who did not write the key

QUESTION MIX (mirrors the Vals v2 categories, hardest last)
  Q1-Q4    single-filing retrieval      "What was FY2025 segment X operating margin?"
  Q5-Q8    earnings analysis            "Did management change the FY guidance
                                         language between Q1 and Q2 calls? Quote both."
  Q9-Q11   disclosure analysis          "List the risk factors added since last year's 10-K."
  Q12-Q14  comparables (cross-entity)   "Rank the 5 names on FY2025 FCF conversion."
  Q15-Q16  longitudinal                 "Track gross margin guidance over the last 8 quarters."
  Q17-Q18  entitled content             "What does our broker note say about pricing? Cite it."
  Q19      internal data                "Compare our model's FY2026 EPS to consensus."
  Q20      entitlement test             run Q17 as a user without the broker entitlement;
                                         a correct system returns nothing

SCORING (per question, 0 to 3)
  3  correct value AND the exact source passage shown
  2  correct value, source cited at document level only
  1  correct value, no source, or source shown but value off by rounding
  0  wrong, refused, or a fabricated citation
  factuality_rate     = questions scoring 2 or 3 / 19   (Q20 scored pass/fail)
  lineage_rate        = questions scoring 3 / 19
  review_minutes      = reviewer time to grade all 20, per vendor
  cost_per_run        = seats or usage consumed, per vendor

PASS LINE (set before the run)
  factuality_rate >= 0.90, lineage_rate >= 0.75, Q20 = pass,
  review_minutes below the current manual time for the same 19 answers

Three rules keep it fair. Score every vendor on the same day with the same documents, because filings and consensus move. Treat a vendor's request to swap questions as a result. Record the review minutes: a system that scores 90% and takes longer to check than the manual answer has saved nothing. For the structural differences behind the scores, see general assistants against institutional platforms.

Frequently Asked Questions

AI financial research benchmarks explained: which score should a buyer ask a vendor for?

Ask for a factuality rate on your own questions, with the checking method written down, because that is the number that predicts review time. Public leaderboards test a bare model over the public record, so a vendor quoting a Vals or BigFinanceBench score is quoting the model it licenses. AllMind AI has published no score on any of these benchmarks; its offer is a run on a firm's own 20 questions, with every figure linked to its source passage so the reviewer can grade it.

What is Deep FinResearch Bench?

Deep FinResearch Bench is a study by JPMorganChase AI Research, posted to arXiv on April 22, 2026, that graded four deep-research agents against 100 professional reports on 25 S&P 500 companies. It scores report quality on a four-point scale, forecast error by SMAPE across six line items, and factuality as the share of claims an independent checker could verify. Analysts scored 2.84 against 2.31 for the best agent, and factuality ran from 86.0% down to 53.2% by agent.

Are the Vals Finance Agent benchmark results 2026 comparable between v1.1 and v2?

No. Version 1.1 has 537 questions; version 2 has 927 expert-reviewed questions with a held-out test split. The top score fell from 64.37% on v1.1 (Claude Opus 4.7, June 4, 2026) to 60.60% on v2 (Muse Spark 1.2, August 19, 2026) because the set got harder. The v2 category scores, from 81.8% on general quantitative questions to 34.5% on financial modeling, are the more useful read.

What does FinanceBench measure?

FinanceBench measures open-book question answering over filings: 10,231 questions about public companies, each with an answer and an evidence string, published by Patronus AI and co-authors on November 20, 2023. The widely quoted 81% comes from a 150-case sample in which GPT-4-Turbo with a retrieval system incorrectly answered or refused to answer 81% of questions. Those were 2023-era models, and the metric counts refusals together with wrong answers, so the figure describes retrieval failure as much as reasoning failure.

Has AllMind AI published scores on these benchmarks?

No. AllMind AI has not submitted to Vals, BigFinanceBench or FrontierFinance, and it has not published a Deep FinResearch-style factuality rate, so no comparable score exists as of August 28, 2026. What it publishes instead is the mechanism: every claim in a drafted report is cited to the passage it came from and a verification pass re-checks figures before the report ships. A buyer should treat that as a claim to test on their own questions, using the 20-question template in this article.


AllMind AI is the AI-native research platform for institutional equity teams. If you want proof on your own work, send us the workflow you want tested.