Selected work on making AI output reliable enough to act on in regulated financial workflows.
Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable
P. Ahuja, 2026 · arXiv & SSRN
A seven-dimension reliability metric for LLM output in capital-markets workflows. Demonstrated across five workflows and four models, scored by four independent LLM judges from three model families, with a deterministic verification script (95 checks, all passing).
A notable result: two frontier models were statistically indistinguishable on reliability despite a material cost difference, which argues for choosing models on workflow fit and cost rather than headline quality.
Read on SSRN
Read on arXiv
Code & data
CC BY 4.0
The Checking Problem: What Must Be True Before AI Ships in a Regulated Firm
P. Ahuja, 2026 · arXiv & SSRN
A controlled study of why enterprise AI programmes stall. Six document-heavy workflows of the kind performed daily in regulated financial services, run across four model families and three tool configurations, three times each: 5,093 scored output elements across 72 configurations. 57 cleared a demonstration bar; 32 cleared a production bar requiring sustained accuracy, reproducibility across repeats, verifiable attribution and a confidence signal that carries information. A survival rate of 56.1%.
A tool that states no confidence requires review of 100% of its output, because it offers a reviewer no basis for triage. Requiring it to cite sources and state a confidence reduces that to 49%. The value of an AI workflow is set less by how often it is right than by how much of it a human must still check.
Read on SSRN
Read on arXiv
PDF
CC BY 4.0