Why “AI in accounting” is really three different questions

Asking whether AI can be trusted in accounting is really three separate questions with three separate answers, and treating them as one leads either to over-automating judgment calls or under-using tools that are genuinely reliable.

Layer 1: data capture and transaction processing (automate this)

Reading structured or semi-structured data off an invoice, matching a payment to an open item, categorizing a routine transaction against a known chart of accounts. These tasks share one property: the correct output is verifiable against the source document itself. If the model misreads a number or misfiles a category, a downstream check against the source will catch it. This layer tolerates automation without a human decision on every single instance.

Layer 2: compliance and consistency checks (assist, then review)

Flagging entries that look inconsistent with VAT treatment, the chart of accounts, or prior-period patterns. AI is useful here as a first pass, surfacing candidates a person should look at, but the check itself needs a reviewer, because “looks inconsistent” is not the same as “is wrong,” and the tool can both miss real problems and flag non-problems depending on how well its rules match your specific chart of accounts and business.

Layer 3: judgment calls dependent on context outside the document (do not automate unsupervised)

Whether a specific expense is tax-deductible given its actual business purpose, whether a supply belongs to this accounting period or the next, how to assess the substance of a related-party transaction. These questions turn on facts and context that are often not present in the document being processed, and a model can produce an answer that sounds confident and is wrong, without the error being visible to anyone who could not already judge the question correctly. That is a materially different risk than layer one, where a wrong output is checkable against the source and gets caught.

What this means practically

Automate layer one without hesitation; the risk of a silent, uncaught error is low because the output is verifiable. Use AI as a first pass at layer two, but keep a reviewer in the loop. Introduce a documented control step at layer three specifically, someone with the authority and the context to catch a wrong judgment call before it reaches a filing or a client-facing number, rather than treating the model’s answer as final because it sounded right.

Where the AI Act’s Article 4 fits

Most AI-assisted accounting features do not automatically fall into the AI Act’s high-risk categories; that classification depends on the specific system and use case. What applies more broadly is Article 4, replaced by Regulation (EU) 2026/1744 with effect from 27 July 2026, which requires providers and deployers of AI systems to take measures supporting the AI literacy of the people who operate or use those systems, taking into account their role and the context in which the system is used. For a firm using these tools, the practical version of that obligation is straightforward: the people relying on the tool’s output should understand which of the three layers they are looking at, and that layer three specifically still needs a human judgment call. The current wording does not require guaranteeing a specific literacy level for each person or issuing a certificate.

For how a similar reliability question plays out in legal work, see AI cannot replace lawyers under Czech law, but it reshapes legal work. More answers are in the answer hub; to talk through where your accounting tools sit across these three layers, use the inquiry form.

Sources and limitations

The description of Article 4 reflects Regulation (EU) 2026/1744 and the European Commission’s current published guidance, both verified 6 August 2026. The three-layer framework is CIAD’s own analytical framing, not a cited external study, and is offered as a practical way to think about where automation is and is not safe, not as a formal risk classification under the AI Act. This page does not provide tax or accounting advice for a specific transaction.

Frequently asked questions

Which accounting tasks are actually safe to fully automate?

Tasks where the correct output is verifiable against the source document itself: reading structured data off an invoice, matching a payment to an open item, categorizing a transaction against a known, stable chart of accounts. These are checkable mechanically, which is why they tolerate automation without a reviewer in the loop for every instance.

Why can't AI handle tax-deductibility and period-allocation questions the same way?

Because the correct answer depends on context the document does not contain: the business purpose of an expense, which financial period a supply actually belongs to, or the substance of a related-party arrangement. A model can produce an answer that reads as confident and turns out to be wrong, and the error will not be visible to someone who could not have judged the question correctly themselves. That is a different risk profile from a task where the output can be checked against the source.

Does accounting software count as a high-risk AI system under the AI Act?

Not automatically. Most accounting AI features do not fall into the AI Act's high-risk categories by default; that classification depends on the specific system and use case, not on the fact that it touches financial data. What does apply broadly is Article 4, the AI-literacy obligation, which is a separate and lighter-touch requirement from the high-risk rules.

What does Article 4 actually require for a firm using AI accounting tools?

Article 4, as replaced by Regulation (EU) 2026/1744 with effect from 27 July 2026, requires providers and deployers of AI systems to take measures supporting the AI literacy of the people who operate or use those systems, proportionate to their role and the context of use. For a firm using an AI-assisted accounting tool, that means people relying on the tool's output should understand where it is reliable (layer one) and where it is not (layer three), not that the firm must guarantee a specific literacy level for each employee or produce a certificate.

Should a human always review layer-three output before it is used?

Yes. A documented control step, someone with the authority and context to catch a wrong judgment call, before the AI's output on a context-dependent question is relied on for a filing or a client-facing figure, is the practical safeguard for the layer where automation is least reliable. This is a process recommendation, not a description of a specific legal audit requirement for every firm.

SOURCES AND VERIFICATION