Dossier · AI systems audit

Evidence and sampling in an AI systems audit

Complete documentation does not prove that a control works. Useful evidence connects an assertion, a population, a reproducible procedure, a result and the exact limits of the conclusion.

On this page
Assertion

Start with what the evidence must be able to contradict

A claim such as “the system is supervised” must be decomposed. Who intervenes, on which outputs, at what point, with what information, authority and trace? The assertion becomes testable: for the defined population and period, every output subject to validation must retain the validator’s identity, decision and the version examined.

The chain remains explicit: use, risk, criterion, population, procedure, result, conclusion, decision. If an item supports no precise link, it documents context but does not establish the assertion.

Population

Define what could have been selected

A sample only has meaning after its population has been defined: period, versions, case types, users, environments and events. The auditor then retains the selection method, its justification, exclusions and the limits of extrapolation.

AI systems often require several selection angles: nominal cases, edge cases, degraded conditions, business exceptions, human interventions and exposed subgroups. A global average may remain stable while a critical subset deteriorates; stratification therefore follows risk, not an arbitrary sample size.

Admissibility

Distinguish an available item from admissible evidence

Availability is not enough. Evidence is examined across four complementary properties: relevance to the criterion, reliability of origin and integrity, sufficiency of the set, and reproducibility of the procedure. Required strength depends on risk and the assurance sought.

PropertyControl questionTypical failure
RelevanceDoes the item answer the assertion for the relevant period?A general policy used to conclude on one precise case
ReliabilityAre origin, version and integrity established?A screenshot without source or timestamp
SufficiencyDoes the set cover relevant risks and variants?One nominal case extended to the whole population
ReproducibilityCan another team replay the procedure?Missing parameters, data or environment
Replay

Retain enough context to reproduce the result

The work file records the evidence identifier, origin, date, version and custodian. For a test, it adds the data or its fingerprint, parameters, environment, protocol, useful raw result and uncertainty. An isolated output, however impressive, cannot distinguish stable behaviour from a demonstration accident.

Limit

Never conclude beyond what was observed

The conclusion identifies the population, period, versions, procedures and access covered. It distinguishes what is established, what remains interpreted and what could not be verified. Restricted access, a version discrepancy or an absent subgroup becomes a visible reservation; it does not disappear into an average score or cautious wording.

Primary sources

Sources this page relies on

Last documentary review: 5 September 2026.

Before you write to us

Frequently asked questions

Is there one sample size that works for every AI system?

No. Selection depends on the population, decision, risks, subgroups, frequency of exceptions and assurance sought. The method and its limits must be justified.

Is supplier documentation sufficient evidence?

It may establish what the supplier declares or designs. On its own, it does not prove configuration, use or control effectiveness in the actual operating environment.

Audit dossier

Can your evidence reproduce the conclusion?

Describe the assertion, population and available traces. We identify the procedures required and the foreseeable limits.

Tell us about your situation