[ Finance ]
AI Audit Trails in Finance: Six Questions an AI-Touched Workpaper Should Answer
A six-question evidence packet helps finance preserve the source, method, configuration, review, change history, and independent testing path behind AI-touched analysis.

On this page
An AI audit trail in finance links a material figure to its approved source population, calculation, system configuration, human review, and subsequent changes. Saving the PDF, prompt, and chat history is not enough. The evidence required still depends on the engagement, materiality, reliance, risk, and intended use.
The independent-replay test is blunt: can a qualified person who did not perform the original work retrieve the source data, rerun the calculation, inspect the exceptions, and see why the reviewer approved it? A transcript alone cannot prove reconstruction, recalculation, or review.
This article offers a six-question preparation packet for finance teams. It draws on PCAOB audit-evidence standards and AI-control guidance from COSO and NIST, but it is not an auditing standard. The evidence needed will depend on the engagement, materiality, reliance, risk, and intended use.
What is an AI audit trail in finance?
An AI audit trail links a material output to the inputs used, the calculation performed, the configuration that ran, the human review, and later changes. Activity logs, lineage records, and workpapers each supply part of that chain. None is sufficient alone.
For public-company audits under PCAOB AS 1215, the auditor's own documentation must show procedures performed, evidence obtained, conclusions reached, who performed the work, who reviewed it, and the relevant dates. The file must be understandable to an experienced auditor who had no previous connection with the engagement. AS 1215 governs the auditor's documentation. It does not impose a blanket retention rule on a company's AI logs.
PCAOB AS 1105 focuses on whether audit evidence is sufficient, relevant, and reliable. Saving the same unsupported answer as a screenshot, PDF, and chat history does not strengthen it.
Which six questions should an AI-touched workpaper answer?
A finance team should prepare six answers: where the inputs came from, how the result was computed, which configuration ran, who reviewed the work, what changed, and whether another qualified person can reconstruct the run, recalculate material figures, and, where applicable, reperform the underlying procedure or control.
For each question, the packet identifies evidence to retain, a likely owner, an internal test, and the change that should trigger reassessment.
| Question | Evidence to consider retaining | Likely owner | Internal test | Trigger for reassessment |
|---|---|---|---|---|
| 1. Where did the inputs come from? | Source-system and report IDs; entity, book, period, currency, filters, extraction time, row count, control total, source snapshot or immutable record references | Finance process owner and source-system owner | Retrieve the same approved population and tie it to the recorded control total | Late journal, reopened period, remapping, source correction, or changed extraction criteria |
| 2. How was each figure computed? | Formula, query, rule, transformation path, assumptions, thresholds, reconciliation, and the boundary between calculated figures and generated narrative | Metric or calculation owner | Recalculate the material figure outside free-form model output | Formula, query, mapping, assumption, threshold, or upstream-data change |
| 3. Which configuration produced the output? | Run ID and time; code, query, prompt, system prompt, model, parameter, retrieval index, connector, plugin, and policy versions where material | Workflow or system owner | Resolve the recorded run manifest to the exact approved configuration | Model or vendor update, prompt edit, parameter change, connector change, retrieval refresh, or policy revision |
| 4. Who performed and reviewed the work? | Named preparer and reviewer, timestamps, evidence inspected, exceptions challenged, overrides, dissent, and approval decision | Finance control owner | Confirm the reviewer had authority, inspected the evidence, and resolved the recorded exceptions | Reviewer change, late evidence, new exception, unsupported override, or altered intended use |
| 5. What changed since the approved run? | Source, logic, configuration, and output diffs; reason; impact assessment; test result; approver; effective time; rollback record | Change owner and finance approver | Compare the current run with the last approved state and trace every material difference | Any unapproved change or a change whose impact was not tested |
| 6. Can an independent person reconstruct the run and test the result? | Runbook, permissions, retained source references, run manifest, calculations, output, exceptions, review record, and expected control totals | Process owner, tested by an independent operator | A qualified operator reconstructs the run, recalculates material figures, and, where applicable, independently executes the underlying procedure or control without asking the original preparer what happened | Missing artifact, failed control total, inaccessible source, irrecoverable version, or unexplained variance |
Keep the artifacts that let a reviewer challenge the result. A prompt belongs in the file when it changes the procedure; a token count rarely does. More telemetry can make a weak workpaper harder to inspect.
When does AI-generated analysis enter audit scope?
AI does not bring every forecast, board memo, or internal analysis into external-audit scope. The relevant question is how the output is used. If management relies on it for a financial-statement assertion, accounting estimate, material journal, management review control, or company-produced information supplied as audit evidence, its sources, calculations, and review may become relevant to the auditor's procedures.
For PCAOB engagements, AS 1105 requires the auditor to test the accuracy and completeness of company-produced information used as audit evidence, or test the relevant controls, and to evaluate whether the information is sufficiently precise and detailed. For audits of financial statements for fiscal years beginning on or after December 15, 2025, AS 1105 also requires procedures over electronic information the company received from external sources and provided to the auditor: understand its source and the company's process for receiving, maintaining, and, where applicable, processing or modifying it, then test the information or the relevant controls.
AICPA Statements on Auditing Standards apply to audits of nonissuers whose audits are not required to be conducted under PCAOB standards; PCAOB standards govern audits of issuers and SEC-registered brokers and dealers. The AICPA's explanation of SAS No. 142 identifies accuracy, completeness, authenticity, and susceptibility to management bias as factors in evaluating evidence. Different standards apply, but finance still needs to show which information mattered, where it came from, how it was used, and who reviewed it.
The right first question is not, "Did AI touch this?" It is, "What decision, assertion, control, or audit evidence depends on this output?" That answer helps determine the documentation depth.
How should finance preserve logic, model, and prompt changes?
Finance should preserve material AI changes the way it preserves changes to a calculation or controlled workflow: name the previous state, the new state, the reason, the test, the approver, the effective time, and the rollback path. A vendor's model update can change an output even when finance never edits the prompt.
COSO's February 2026 GenAI guidance applies the Internal Control-Integrated Framework to generative-AI capabilities. For higher-impact uses, the guidance discusses inventories and version histories for prompts, system messages, model versions, retrieval connectors, outputs, exceptions, approvals, and rollback records. COSO is guidance, not an auditing standard or law.
NIST's voluntary, cross-sector Generative AI Profile calls for documenting data origin, human oversight, limitations, version history, and testing. It does not prescribe a finance workpaper, but it supports retaining provenance and change history alongside performance results.
Do not make audit readiness depend on regenerating identical prose from a probabilistic model. Preserve the controlled inputs, method, configuration, and evidence well enough to reconstruct the run, recalculate material figures, and, where applicable, reperform the underlying procedure or control. If wording changes but the source-backed figure, exception set, and reviewer conclusion still meet the documented criteria, the work can still be tested.
Can a new reviewer test the work without calling the preparer?
Run an independent replay: give the workpaper to a qualified colleague who did not prepare it, without a walkthrough. Ask that colleague to retrieve the source population, recalculate one material figure, inspect the configuration and exceptions, and decide whether the recorded approval still applies.
| Step | Operator action | Pass condition | Failure diagnosis |
|---|---|---|---|
| 1. Retrieve | Open the recorded source population for the stated entity, book, period, currency, and filters | Population and control total match the workpaper | Source lineage or extraction control is missing |
| 2. Recalculate | Run the recorded formula, query, or rule without using generated prose | Material figures tie within the approved tolerance | Calculation logic, mapping, or assumptions are missing |
| 3. Resolve | Open the run manifest and identify the material logic, prompt, model, connector, and policy versions | Every version resolves to a retained approved state | Configuration history is incomplete |
| 4. Challenge | Inspect exceptions, contradictory evidence, overrides, reviewer comments, and dissent | Disposition and authority are visible for every material exception | Review was ceremonial or exceptions were suppressed |
| 5. Compare | Check changes since approval and rerun the invalidation rules | Approval remains valid or the item has been reopened | Change control cannot show whether the old sign-off still applies |
Apply the five-step independent replay to an automated variance-analysis workpaper or one task in an AI-assisted monthly close. Variance and close workflows already have source populations, control totals, exceptions, and named review points. A general chatbot answer is harder to test because interaction history is not financial lineage.
Which decisions must remain human?
AI cannot bear responsibility for deciding that audit evidence is sufficient, for management's financial-statement assertions, for approving an accounting judgment, for waiving a broken control total, for resolving contradictory evidence, or for signing a review. Those decisions remain with accountable people who have the authority and context to challenge the work.
The PCAOB's July 2024 GenAI staff outreach spotlight reported that participating firms expected GenAI to augment, not replace, people. The engagement-team member using a GenAI tool remained responsible for the results and documentation, and reviewers were expected to apply the same diligence as they would to work without GenAI. The outreach covered mainly larger audit firms and a limited set of preparers. It was a snapshot, not a rule or a market census.
A reviewer name and timestamp show that someone clicked. A useful review record shows what the reviewer inspected, which exceptions they challenged, what changed, and why they approved the result. The AI governance checklist for finance tests that broader control environment. This article's packet tests whether one result can survive independent challenge.
How does Pluvo preserve the evidence chain?
Finance Engineering is the discipline of building AI-native finance systems that are accurate, governed, auditable, model-agnostic, and directly tied to how the business actually operates. Pluvo's public lineage and controls pages describe an evidence chain that links approved sources to definitions, transformations, controlled calculation logic, output versions, and approval records.
Language models can help investigate exceptions and draft explanations, while material figures remain traceable to governed calculation logic and human review. LLMs predict text; they don't compute. In this article, that line is a finance-system boundary: language-model output is not, by itself, a controlled calculation. A number that's 95% right is 100% useless. The 95% line is a quality bar, not a measured error rate.
For practical field notes on governed finance workflows, subscribe to the Finance Engineering Newsletter.
A preserved sentence can be evidence when it formed part of the control or conclusion. It cannot substitute for the population, calculation, run manifest, exception record, and review that produced it.
Frequently asked questions
What is an AI audit trail in finance?
An AI audit trail is the linked evidence behind a material output: approved inputs, calculation method, system configuration, human review, changes, exceptions, and records used to reconstruct the run, recalculate figures, or reperform a procedure or control.
Is a ChatGPT or AI conversation history an audit trail?
Not by itself. Conversation history can show prompts and responses, but it does not prove source completeness, controlled calculation, configuration integrity, meaningful review, change history, or successful independent testing.
Does every AI-generated finance analysis fall within external audit scope?
No. Relevance depends on the engagement, materiality, risk, reliance, and whether the analysis affects financial reporting, ICFR, a material assertion, or information that management or the auditor uses as evidence.
Does an AI system need to reproduce the same prose for an audit?
Not necessarily. Preserve controlled inputs, methods, versions, calculations, exceptions, and review evidence so a qualified person can reconstruct the run, recalculate material figures, and, where applicable, reperform the underlying procedure or control.
Which changes should trigger reassessment of an AI-touched workpaper?
A late journal, source correction, mapping change, model or prompt update, retrieval refresh, changed threshold, new exception, altered intended use, or other untested change can require the workpaper to be reopened and reviewed.



