Learn the art of finance engineering →
← All posts

Finance ]

10 Questions to Ask AI Vendors About Accuracy

Use a 10-question answer decoder to test an AI vendor's accuracy claims, evidence, change controls, human authority, and failure handling.

Vanessa Galarneau

6 min read
Share
A focused woman at a bright review desk touches violet-lit clear testing modules arranged as cell banks, gated rails, and a circular monitoring loop.
On this page

The slide says 95% accurate. That number is useless until the vendor names the workflow, denominator, ground truth, system version, and failed cases, then proves the claim on buyer data. Without that chain, the percentage is sales copy, not decision evidence.

The missing denominator matters. In a February 2026 evaluation paper, six NIST researchers studied 22 API-accessible frontier language models across three benchmarks. They separated benchmark accuracy on a fixed question set from generalized accuracy across similar questions. One number can answer the first question while leaving the second wide open.

Finance buyers need a harder standard. The vendor must say what was tested, which mistakes count, how severe failures are handled, and whether the evidence survives a model update. Then the buyer should make the system perform.

What does accuracy mean in a finance AI workflow?

In finance, accuracy belongs to the workflow, not the model. A cash forecast, a board-pack narrative, and an invoice classifier face different populations, error costs, and human gates. The metric has to follow the work.

A four-part definition of finance AI accuracy
DimensionQuestion the buyer must answerEvidence
Numeric correctnessWas every material figure calculated correctly from approved inputs?Known-answer tests, expected-versus-actual results, control totals, and recalculation
Contextual fitnessDid the result use the right entity, period, currency, metric definition, and source population?Scope record, definition version, filters, mappings, and exceptions
Evidence completenessCan another person reconstruct the inputs, logic, system version, changes, and review?Trace ID, source lineage, configuration history, run log, and approval record
Decision fitnessIs the residual error acceptable for the intended action, audience, and materiality?Error tolerance, severity classes, human release rule, monitoring threshold, and fallback

The four dimensions are not a composite score. A perfect citation trail cannot rescue the wrong legal entity. A repeated calculation cannot rescue an unapproved formula. One material break can invalidate the answer even when the average looks excellent.

NIST says it is revising the AI RMF 1.0 Core, but the current Core still calls for documented test sets, metrics, deployment-similar conditions, regular evaluation, production monitoring, and independent assessment. It also asks whether a system can fail safely outside its knowledge limits. That is closer to a finance definition of accuracy than a leaderboard screenshot.

Which 10 questions should finance ask AI vendors?

Ask the ten questions against one real workflow, not a tour of prepared features. The answers below are patterns, not quotations from any named vendor. An answer counts only when the promised artifact exists and the buyer can inspect or replay it.

The Accuracy Answer Decoder, questions 1 to 5
QuestionAn incomplete answerAn answer worth hearingProof to requestStop if
1. What exactly does accuracy mean for our use case?"Our model is 95% accurate."A named metric tied to the workflow, decision, population, error classes, tolerance, and measurement uncertainty.Metric definition with denominator, tolerance, uncertainty, and failed casesThe vendor cannot name the denominator or error classes
2. Was the system tested under conditions like ours?"We lead the public benchmarks."Public benchmarks are preliminary. We also tested representative tasks on unseen, production-like data and support a buyer-supplied blind set.Dated report with sample, environment, edge cases, and raw failuresNo blind buyer test or raw failures are available
3. How are facts, citations, and calculations verified?"Retrieval prevents hallucinations."Sources are versioned and checked against the cited passage. Material calculations are reproduced against known ground truth outside free-form text generation.Source trail and calculation harness with expected-versus-actual resultsMaterial numbers or citations cannot be opened and checked
4. Can we reproduce a result later?"Everything is in the audit log."A trace ID resolves to the input snapshot, retrieval sources, model and workflow versions, configuration, timestamp, exceptions, and overrides.Trace ID, retained run, system version, and successful replayA prior material result cannot be reproduced
5. What happens when the model, data, prompt, or workflow changes?"Upgrades are automatic."Material changes trigger notice, regression testing, a before-and-after comparison, approval, and a rollback rule.Changelog, regression result, approval, and rollback demonstrationMaterial updates bypass regression testing or rollback
The Accuracy Answer Decoder, questions 6 to 10
QuestionAn incomplete answerAn answer worth hearingProof to requestStop if
6. When does the system abstain or escalate uncertainty?"It always gives users the best available answer."Workflow-specific thresholds block, abstain, or route uncertain and out-of-scope cases to a named owner.Abstention rules, threshold tests, exception route, and override recordThe system never refuses or the threshold is hidden
7. What did the system get wrong, and how severe were the failures?"The remaining error rate is only 5%."Failures are shown by error class and severity, including material cases, false positives and negatives where relevant, and remediation.Representative failed cases, severity taxonomy, and remediation recordFailed cases are withheld or all errors are treated alike
8. How was ground truth created and protected from test-set leakage?"Our experts labeled the benchmark."The vendor names the labelers, adjudication method, source provenance, holdout design, version, and controls that kept the test set unseen.Ground-truth protocol, sign-off, holdout lineage, and leakage controlsLabels have no owner or the holdout may have entered development
9. How does accuracy vary by workflow, entity, period, and edge case?"Performance is consistent across customers."Results are disaggregated by task and relevant operating slice, with the worst-performing cases and limits shown beside the average.Slice-level results, worst-case failures, limits, and edge-case testsOnly an aggregate average is disclosed
10. What happens after a material answer is wrong?"A human reviews important outputs."The workflow stops distribution, notifies an owner, corrects and invalidates affected results, reruns the work, and adds the failure to regression testing.Correction trail, notification rule, incident rehearsal, and new regression caseThe vendor cannot invalidate, correct, notify, and rerun

How should buyers decode an accuracy answer?

Move every important answer through three rungs: claim, artifact, replay. The claim says what the system does. The artifact records how the claim was tested. The replay lets the buyer reproduce or challenge the result on a relevant case. The trouble starts when the ladder ends there.

Audit practice offers a useful, narrow analogy. PCAOB AS 1105 says inquiry of company personnel, by itself, is insufficient for the relevant assertion or control conclusion. The standard separately defines recalculation as checking mathematical accuracy and reperformance as independently executing a procedure or control. A software purchase is not an audit. But inquiry, recalculation, and reperformance still describe three different levels of assurance.

Consider the answer, “Every figure has lineage.” The artifact should identify the source population, filters, joins, transformations, definition, and calculation. The replay should open one material figure, change one approved input, return the expected result, and retain the change. If the trail stops at a chart or citation, finance still cannot reproduce the number.

The technical boundary behind numeric claims is explained in Why LLMs Hallucinate Financial Numbers. LLMs generate text token by token; by themselves, they are not dependable calculation engines. A tool-enabled language model can call deterministic code or a validated quantitative engine for the calculation, but finance still needs the inputs, logic, result, and version record. The test is simple: rerun the same inputs and logic and confirm the same number.

What happens after the model or workflow changes?

Accuracy evidence has a half-life. A model swap, retrieval refresh, prompt edit, connector release, mapping change, new data population, or revised finance definition can make yesterday's test irrelevant. The contract and operating process should say which changes trigger notice, regression testing, reapproval, or rollback.

The NIST Generative AI Profile recommends empirically validated capability testing, known-ground-truth comparison, provenance records, version control, incident processes, and fallback planning. Every material change should trigger another test.

Ask for the last material change, not the vendor's change-management policy. What changed? Which regression set ran? Which metrics moved? Who approved the release? What would have caused a rollback? A real example forces the operating system behind the slide into view.

Monitor outcomes after launch, including abstentions, overrides, exceptions, near misses, corrections, and drift. A system can hold its benchmark score while failing on the new entity, revised chart of accounts, or unusual quarter the buyer actually cares about.

Do certifications and security reports prove AI accuracy?

No. ISO/IEC 42001 is an AI management-system standard. Certification can support the case that a company operates a governed process within the stated scope. It does not establish that one product returns the right answer for one finance workflow.

The same boundary applies to security assurance. Encryption, penetration testing, access control, and incident response matter. They answer whether the system and data are protected. They do not answer whether gross margin used the approved revenue policy, whether a forecast generalized beyond its test period, or whether the cited ledger rows reconcile.

The U.S. Federal Trade Commission's January 2025 IntelliVision order concerned a specific facial-recognition company, not finance software. Its narrow lesson travels: objective claims about AI accuracy and efficacy need competent, reliable testing that supports the scope of the claim.

Security evidence answers whether the data were protected. Workflow evidence answers whether gross margin used the approved revenue policy. Finance needs both.

What must stay human in an AI vendor decision?

Finance leaders must set the intended purpose, ground truth, materiality, error tolerance, and release authority. The vendor can propose metrics and controls. The buyer owns which mistakes are acceptable and which stop the workflow.

A reviewer needs the source records and the authority to stop distribution. Without both, the review is theater.

Humans also own the exit. Finance should know how to export the evidence, preserve required records, return a critical workflow to manual operation, and invalidate prior approvals after a material change.

How should finance run the next vendor demo?

Send one closed-period workflow in advance. Bring an approved source population, a known answer, one restricted user, one missing definition or mapping, and one material change. Ask the vendor to complete the workflow, expose a failure, and leave the evidence behind. The Accuracy Answer Decoder is a Pluvo editorial framework, not a standard or certification.

The seven fields in the buyer's leave-behind
FieldWhat finance records
Vendor answerThe claim in the vendor's own words
Artifact suppliedTest report, trace, log, contract term, certificate, or incident record
Evidence ownerThe person accountable for producing and explaining it
Last-tested dateWhen the evidence was generated or independently reviewed
System versionModel, workflow, retrieval, connector, and configuration scope
Buyer validationWhat finance reproduced, challenged, denied, or reconciled
Red flagMissing denominator, hidden failures, stale version, unsupported generalization, or no rollback

The buyer's guide chooses the category. The governance checklist tests the control environment. The workpaper guide defines audit evidence. This decoder does one job: it turns a vendor's accuracy claim into evidence finance can inspect.

Where does Pluvo fit the accuracy test?

Pluvo says its product computes figures deterministically from source data and preserves the lineage behind each number. Those claims are testable. Buyers should challenge them exactly as they would any vendor's.

Ask Pluvo to open a material figure through financial lineage, show the control and review boundary, and replay the calculation on your approved workflow. If you are evaluating a recurring finance process, bring the ten questions to a demo.

Write down 95% if it appears on the slide. Then ask the question that makes the number useful: 95% of what, on which version, under whose conditions, with which failures?

Frequently asked questions

What questions should I ask an AI vendor about accuracy?

Ask how accuracy is defined, which population and version were tested, what the failed cases were, how facts and calculations are verified, whether a result can be reproduced, how changes trigger retesting, and who can override, stop, or exit the system.

Is 95% AI accuracy good enough for finance?

Not by itself. The answer depends on the workflow, denominator, error severity, ground truth, test conditions, uncertainty, and human release rule. One material error can invalidate a finance output even when the aggregate rate is high.

Should an AI vendor show buyers its failed cases?

Yes. Aggregate accuracy can hide material failure classes. Buyers should inspect representative failed cases, error severity, false positives and negatives where applicable, remediation, and whether those failures entered the regression set.

How can finance test an AI vendor without sharing live sensitive data?

Use an approved closed-period population, a de-identified or minimally scoped blind set where appropriate, a known answer, a restricted user, and deliberate missing context. Preserve enough realism to test the workflow without exceeding approved data access.

What is the biggest red flag in an AI vendor demo?

The biggest red flag is an accuracy claim that cannot resolve into raw failures, a named version, representative test conditions, a source and calculation trail, and a buyer-run replay. A polished answer without inspectable evidence is still only a claim.

About the author

Vanessa Galarneau

CFO & COO

Get new articles in your inbox

FP&A, close automation, and finance-ops writing — no spam, unsubscribe anytime.

Turn your data into a system for real decisions

Book a demo