What “AI system audit” actually means here

The phrase gets used loosely. This page describes what CIAD, and the AI Act itself for high-risk systems, actually checks: not a single test of whether the model gives good answers, but whether the organisation around the model can manage what happens when it does not.

The five categories an audit covers

Risk management. Is there a documented, repeatable process for identifying what can go wrong with this system, in this use, for these people, and for reducing that risk to an acceptable level, reviewed as the system or its use changes rather than assessed once at launch.

Data governance. Where the training and input data came from, whether it is relevant and representative for the intended use, whether known gaps or biases were examined, and whether that examination is documented rather than assumed. This is a distinct check from accuracy testing.

Technical documentation. Whether the system’s purpose, capabilities, limitations, and design choices are written down in enough detail that someone other than the original builder could understand what it does and why, and that a regulator or auditor could review it after the fact.

Logging. Whether the system’s operation is recorded in enough detail, and retained for long enough, to reconstruct what happened if something goes wrong, who is affected, and when.

Human oversight. Whether a specific person or role has the authority, the information, and the real practical capacity, not just the formal title, to review, override, or stop the system’s output before it causes harm.

For a high-risk AI system under the EU AI Act, these five categories are legal obligations (Articles 9 to 15 of Regulation (EU) 2024/1689), not audit best practice. For a system that does not fall into the high-risk category, an audit still checks the same five categories, scaled to what is actually at stake for that particular use.

What the audit does not test

It does not grade whether the model is “smart,” compare it against a competitor’s model, or certify that it will never produce a wrong or harmful output. No audit or test can guarantee that. What it can show is whether the organisation would catch a wrong or harmful output before it does damage, and whether it can demonstrate that to a regulator, a client, or its own leadership after the fact.

Where AI literacy training fits, and where it does not

Article 4 of the AI Act, replaced by Regulation (EU) 2026/1744 with effect from 27 July 2026, requires providers and deployers of AI systems to take measures supporting the development of AI literacy among their staff and other people who operate or use the system on their behalf, taking into account those people’s knowledge, experience, education, training, the context of use, and the people the system is used on. The current wording does not require guaranteeing a specific literacy level for any individual, a defined training format, or a certificate; the European Commission’s AI Literacy Q&A confirms no certificate is required and no knowledge measurement is mandated.

An audit can reasonably ask whether such measures exist and whether there is some record of them, because a record is useful evidence that a measure was actually taken, not a bureaucratic add-on. But the record is not itself the legal requirement, and an audit that treats “a training record exists” as the finish line is checking the wrong thing. The actual question is whether the people operating and using the system understand its limits well enough for the other four categories, oversight in particular, to function.

For the difference between this kind of audit and an adversarial test of the system itself, see penetration test or security audit: what is the difference. For how often a security audit needs repeating once you have one, see security audit frequency. More answers are in the answer hub; to scope an audit for a specific system, use the inquiry form.

Sources and limitations

This page describes audit scope in general terms and is not a compliance checklist for a specific system or a substitute for a legal risk classification of your AI system under the AI Act. The Article 4 description reflects the regulation as amended and the European Commission’s current published guidance, both verified 6 August 2026; consult a lawyer for how the obligation applies to your specific organisation and systems.

Frequently asked questions

Does an AI audit test model accuracy?

It can measure accuracy against a defined benchmark or test set, but that is only one input. The larger question an audit asks is whether the organisation knows the accuracy, has a process for re-checking it as data or usage drifts, and has defined what happens when the model is wrong, not whether the model is right in a single test run.

What counts as 'human oversight' in an audit?

A named person or role with the authority, the information, and the actual practical ability to stop or override the system's output before it causes harm, not a person who technically reviews outputs they have no time or context to meaningfully evaluate. Auditors typically ask for evidence: who is that person, what do they see, and can they show a case where they actually intervened.

Does the audit check data quality?

Yes, for high-risk systems this is a distinct legal requirement, not a subset of accuracy testing. It covers where training and input data came from, whether known biases were assessed, whether the data is representative of the population the system is used on, and whether that assessment is documented, not assumed.

Does the audit look at AI literacy training records?

It can, as one of several supporting practices. Article 4 of the AI Act (as replaced by Regulation (EU) 2026/1744) requires providers and deployers to take measures supporting the AI literacy of staff and other people who operate or use the system on their behalf, taking into account their knowledge, role, and the context of use. It does not require a guaranteed level for each individual, a specific training format, or a certificate. Keeping a record of what training or other measures were taken is good practice for showing the measures were real, but the record itself is not what the law requires, the measures are.

Is an AI audit the same as a penetration test?

No. A pentest asks whether someone can break into or manipulate a system within a defined scope. An AI audit asks a broader question about risk management, data governance, documentation, and oversight. See [penetration test or security audit](/en/answers/penetration-test-or-security-audit-what-is-the-difference/) for the general distinction, which applies here too: an AI audit may include adversarial testing of the model (prompt injection, jailbreak attempts, data extraction) as one input, but that testing alone is not the audit.

SOURCES AND VERIFICATION