LLM evaluation in finance: measurement, leakage, and verification
Validating LLM outputs for financial and economic research requires strict safeguards like data contamination checks, bias correction, and multi-model cross-verification to keep flawed calculations from corrupting market and risk models.

Mark Vandergon
Lead Data Scientist
Federal Reserve Bank of New York
What you'll take away from this session
Bias correction for AI-generated predictors is solvable
Researchers developed correction methods with software libraries to address the bias LLM outputs introduce when used as regressors in downstream models. Naive two-step approaches produce systematic error; applying these corrections should be standard practice for any quantitative team using AI-derived predictors.
Training leakage is a risk in financial analysis
Holdout tests, memorization probes, and robustness checks using models with earlier knowledge cutoffs can surface training leakage before it contaminates a finding. These tools can detect risks in LLMs that recall rather than reason with trained data from the period being analyzed.
Different models disagree substantially on financial sentiment
Pairwise agreement scores across frontier and open-weight models on Federal Reserve transcript data are notably lower than on standard financial filings. Selecting one model without benchmarking against others introduces a verification gap.
Supervised learning and human annotation remain essential
Fine-tuned smaller models still outperform larger general models on specialized financial classification tasks. Human-annotated validation sets remain necessary to confirm that what the model measures corresponds to what researchers intend.
Reliability lags behind accuracy across agentic financial tasks
Accuracy benchmarks for financial AI tasks are improving rapidly. Reliability measures covering consistency, calibration, and cross-model agreement are not moving at the same pace, a gap that becomes more consequential as agentic workflows extend task horizons from minutes to hours.
LLM finance research has an accuracy and verification problem, yet only one of them is being solved. At the Federal Reserve Bank of New York, Lead Data Scientist Mark Vandergon has spent years building the methodological infrastructure to use language models as scientific instruments in economic research. His conclusion is that LLM evaluation benchmarks for financial tasks are improving rapidly, but the reliability infrastructure needed to trust those results has not kept pace.
Three failure modes structure the session. Using LLM outputs as regression inputs without bias correction produces systematically wrong coefficients. Correction methods exist in published software libraries but are not yet standard practice. Models trained on data covering the period being analyzed may be recalling rather than reasoning, a leakage risk detectable through holdout tests and memorization probes. And when the team ran sentiment classification across frontier and open-weight models on Federal Reserve transcript data, pairwise agreement was substantially lower than on standard financial filings — a finding that applies to model risk management workflows that rely on a single model's output without cross-model validation.
The session's transparency research sharpens the stakes further: when the team modeled market access to Fed deliberation transcripts, reduced forecast error was driven by sentiment in the text, not quantitative forecasts or hawk-dove signals. Institutions building AI-powered analytics around communications should reorder where verification infrastructure matters most.
FAQ
How can LLMs be used as measurement tools in financial and economic research without systematic error?
Why do frontier LLMs disagree on financial sentiment classification, and what does that mean for model risk management?
What did Federal Reserve Bank of New York research reveal about the information markets most need from central bank communications?
Transform the work that matters most
See how Domino helps the world’s most regulated enterprises build, scale, and govern AI-powered applications.