·Peter PanPractice

The Truth About AI Governance: How Internal Audit Opens the Black Box

The Truth About AI Governance: How Internal Audit Opens the Black Box — From five ideas to a life cycle audit route you can actually walk
Written by: ARKO AI Editor | Editor in charge: Peter Pan

■ 1. Introduction: the transparency crisis inside the AI wave

Organisations today face a formidable challenge: we are adopting artificial intelligence faster than we have adopted any technology before, and at the same time our sense of control over these systems is draining away. Once AI reaches automated decisions, credit scoring and even hiring, management tends to retreat behind a wall marked "too technical to explain". But AI should never be a black box that cannot be questioned. The current transparency crisis is not only a technical challenge; it is a failure of corporate governance. Auditing AI is no longer the preserve of engineers — it is central to how an organisation secures its survival and its credibility. As regulatory pressure builds worldwide, auditors have to step out from behind the curtain.

From several years of teaching ISO/IEC 42001:2023, I have found that the sticking point is almost never "we don't know AI carries risk". It is "we know there is risk, but we don't know when to look, what to look at, or what evidence to ask for". Most AI governance writing stops at the level of principles: be fair, be explainable, keep a human in the loop. None of that is wrong, but auditors need checkpoints and working papers, not adjectives.

This article therefore comes in two halves. The first sets out five ideas that reframe AI governance and help readers see how an uncontrollable technology becomes a controlled business asset. The second supplies the missing axis — using the AI system life cycle processes of ISO/IEC 5338:2023 to pin those five ideas to specific points in time and specific evidence. Along the way it clears up a misconception that circulates widely in AI circles: does the AI life cycle have seven stages or eight?

fig01_article_overview
fig01_article_overview

■ 2. Idea: you do not need to be an AI engineer to audit AI

Many audit professionals find AI intimidating and assume they must master machine learning before they can engage. This is a serious misreading. In the spirit of the Global Internal Audit Standards (GIAS 2024, effective 9 January 2025), auditing begins with a risk-based approach. The auditor's job is not to read code line by line, but to assess whether the governance framework built by the first line (operations and development) and the second line (risk, compliance, security) is actually working.

I have used this comparison many times: an auditor does not need to be a paper engineer to audit financial statements, and by the same token an auditor does not need to master neural networks to audit AI. What matters is verifying control points and accountability. Working through the Three Lines Model, internal audit should focus on the effectiveness of data governance, the procedural transparency of model validation, and the accountability structure around decisions.

"Internal audit does not need AI engineering expertise to perform testing; but where the subject matter is highly technical — adversarial testing, model robustness — the standards call for external expert support."

Both halves of that statement have to hold at once. Audit cannot absent itself because of a technical threshold; equally, audit should not pretend it can run a red team. The healthier arrangement in practice is this: audit designs the questions and the evidence requirements, the technical testing is outsourced or performed by the second line, and audit retains judgement over whether the test scope was sufficient and whether findings were tracked to closure. That is what independence actually looks like.

■ 3. Idea: shadow AI is the largest risk you cannot see

In AI governance the biggest exposure is rarely the system you know about; it is the unauthorised tool. Employees trying to work faster may feed protected customer data or company confidential material into a public model without ever thinking of it as a disclosure. Procurement records and financial statements will not surface this, because most shadow AI is free or already bundled inside software you own. Auditors have to move from a purely financial mindset (read the invoices) to an operational technology mindset (read the traffic).

Key techniques for detecting shadow AI:
▸ Cloud and SaaS subscription reconciliation: cross-check IT spend against cloud service access logs.
▸ Network egress monitoring: use data loss prevention (DLP) to watch traffic heading to known AI endpoints.
▸ Browser extension review: inspect AI assistant plug-ins installed on employee devices.
▸ Vendor release notes: check whether your ERP or office suite has quietly switched on AI features nobody approved.
▸ Identity provider logs: review SSO and OAuth grants to find third-party AI services signed in with corporate accounts.
▸ API key inventory: search code repositories and CI/CD configuration for model vendor keys.

The last two are often more productive than DLP. DLP sees known endpoints, but the shadow AI that causes real trouble is frequently an API a development team wired up themselves — it never appears in browser traffic, only in the environment variables of some repository. Looking ahead, as SaaS vendors embed AI directly into subscriptions you already pay for, the definition of shadow AI is itself changing — the question stops being "what did an employee sneak in" and becomes "when did the software we bought quietly turn into an AI system". NIST's AI 800-4 report, published in March 2026, points in the same direction: browser-based agent tools and automation wired up with personal API keys are becoming the next generation of shadow AI.

■ 4. Idea: your role determines your audit depth — deployer vs. provider

Effective AI governance refuses a one-size-fits-all approach. A scoping gateway is what allows audit resource to be allocated sensibly. Under the EU AI Act, organisations fall broadly into two roles:
Deployer: if you are simply using a third-party tool (Microsoft Copilot, ChatGPT Enterprise), audit should concentrate on vendor contract management, input data controls and output review.
Provider / Developer: if you build the system yourself or substantially modify a model (fine-tuning, for instance), you carry the highest level of governance obligation, including full life cycle management and technical robustness testing.

This distinction is what lets audit aim at real risk. For a deployer, you can dial down technical testing of the model internals and dial up scrutiny of vendor assurance — SOC 2 reports, model cards, DPA clauses — rather than burning resource on tests you are not positioned to run. There is a practical trap here worth flagging. Article 25 of the EU AI Act provides that a deployer becomes a provider in defined circumstances — putting a high-risk AI system on the market under its own name or trademark, substantially modifying a high-risk system, or changing the intended purpose of a system so that it becomes high-risk. In other words, "we only bought it" is not a defence with indefinite shelf life. When scoping, audit should ask two questions at once: what role are we in today? And is any product or project route quietly pushing us towards being a provider? Fine-tuning, bolting on a RAG knowledge base, and repackaging model output as your own branded service are the three most common escalation triggers.

■ 5. Idea: board-level AI literacy is now a legal obligation

Accountability has shifted: senior management no longer enjoys an exemption for technical ignorance.

AI governance succeeds or fails on oversight from the top. This is not merely an ethical recommendation but a legal duty. Article 4 of the EU AI Act expressly requires AI literacy among relevant personnel. More pointedly, in jurisdictions such as Switzerland, Article 716a of the Swiss Code of Obligations sets out non-delegable board duties, which means AI oversight cannot simply be handed to the IT department.

Why does a literacy gap turn into a governance failure?
▸ Automation bias: management becomes over-reliant on model output and skips the human oversight that was supposed to be there.
▸ Misapplied controls: a board without AI knowledge cannot recognise bias or compliance gaps in AI-generated content.
▸ Risk appetite mismatch: a board that does not understand model uncertainty cannot set a meaningful tolerance threshold for AI-driven decisions.

Internal audit should examine the board skills matrix and confirm that senior decision-makers have enough AI literacy to assess what the technology does to the organisation's risk appetite. AI literacy is also very easy to turn into a compliance performance — run a ninety-minute webinar, collect the attendance sheet, and the evidence exists. If audit only checks whether training took place, it has effectively endorsed an empty shell. A more meaningful checkpoint is this: over the past four quarters, has the board changed a single decision because of AI risk information? If the answer is never, that is not because there was no risk. It is because the risk information never reached that level in usable form.

■ 6. Idea: what is distinctive about generative AI — hallucination and leakage

Generative AI (GenAI) breaks a core assumption of traditional IT control: that system output should be predictable. Hallucination means the internal control framework has to be extended, which is what work such as the COSO generative AI guidance addresses.
The warning is straightforward: if the control framework does not evolve, the organisation ends up with a control vacuum across all five COSO components.

  1. Control environment: no ethical guidance for AI use.
  2. Risk assessment: hallucination and data leakage risks never enter the register.
  3. Control activities: conventional access controls do not cope with a model that changes.
  4. Information and communication: AI-generated content is not marked as such.
  5. Monitoring activities: no mechanism watches for model drift.

Auditors should pay particular attention to disclosure and marking. Article 50 of the EU AI Act requires organisations to ensure AI-generated content can be identified, limiting over-reliance and reducing reputational exposure. Article 50(2) goes further and requires the output of generative systems to be marked in a machine-readable format as artificially generated or manipulated. Watermarking and provenance marking are therefore no longer a technical option — they are an auditable compliance requirement.

The deeper audit challenge generative AI introduces is non-reproducibility. Traditional system audit can re-run a transaction and compare the result; run the same prompt against the same model twice and the answers may differ. That changes the shape of audit evidence itself — it moves from "was the output correct" to "is the process traceable": prompt records, model version, parameter settings, retrieval sources and human review records. You need all five.

■ 7. Idea: without a life cycle map, an audit is just sampling — the axis ISO/IEC 5338:2023 supplies

The five ideas above all hold. But laid out side by side, a structural problem appears: they are five topics, not a process. An auditor who walks into the field with those five points easily ends up sampling at random — shadow AI this year, board training next year. It looks like coverage, but coverage can never actually be stated.

This is precisely where ISO/IEC 5338:2023 earns its place.

◆ 7.1 What ISO/IEC 5338:2023 is, and what it is not

ISO/IEC 5338:2023, "Information technology — Artificial intelligence — AI system life cycle processes", was published in December 2023 by ISO/IEC JTC 1/SC 42. It defines a set of processes and related concepts describing the AI system life cycle, covering both machine learning systems and heuristic systems. Its technical foundation is two existing systems engineering standards — ISO/IEC/IEEE 15288:2023 (system life cycle processes) and ISO/IEC/IEEE 12207:2017 (software life cycle processes) — extended and modified with AI-specific processes drawn from ISO/IEC 22989:2022 (AI concepts and terminology) and ISO/IEC 23053:2022 (framework for ML systems). The introduction to ISO/IEC 5338:2023 is more candid than most standards: its starting position is that rather than inventing a dedicated AI life cycle from scratch, it is better to extend the life cycle processes of conventional software systems and add what is specific to AI.

That positioning has two direct consequences.
First, where a component of an AI system is in fact conventional software or a conventional system, implement it using the processes of ISO/IEC/IEEE 12207:2017 and ISO/IEC/IEEE 15288:2023 as they stand. Audit does not need a whole new programme for AI projects; it needs AI-specific checkpoints bolted onto the SDLC audit programme it already has.
Second — and this matters for what follows — ISO/IEC 5338:2023 states explicitly that it does not prescribe a specific life cycle model. Its concern is which AI-specific processes exist, not what order they should be arranged in. That sentence decides how we should read the argument about the number of stages.

As for how it divides labour with neighbouring standards, my working understanding is this: ISO/IEC 42001:2023 tells you who is accountable, ISO/IEC 5338:2023 tells you what has to be done, and ISO/IEC 23894:2023 tells you how to judge whether it should be done at all. ISO/IEC 5338:2023 itself states that it provides further detail on the life cycle processes discussed in ISO/IEC 42001:2023.

For audit purposes the split is practical: ISO/IEC 42001:2023 gives you the organisation-level checklist, ISO/IEC 5338:2023 gives you the system-level one — and the second is exactly the half most AI audit programmes are missing. A fuller comparison across six standards is set out below.

fig02_six_standards_division
fig02_six_standards_division

◆ 7.2 Three process types: where to start rewriting the audit programme

Most introductions go straight to the "four process groups" of ISO/IEC 5338:2023, but that structure is inherited from ISO/IEC/IEEE 15288:2023. What the standard itself emphasises more is a different classification.

Processes are grouped by how far they depart from the base standards:
Generic: identical to ISO/IEC/IEEE 15288:2023 and 12207:2017. AI changes nothing.
Modified: an existing process with elements added or adjusted. For these, ISO/IEC 5338:2023 adds a dedicated subclause covering AI-specific considerations.
AI-specific: new processes with no counterpart anywhere in ISO/IEC/IEEE 15288:2023 or 12207:2017.

This three-way split is worth more to an auditor than the four groups. It answers the question that actually matters — how much of my existing SDLC audit programme has to change? Generic processes carry over untouched; modified processes need AI-specific questions added; AI-specific processes have to be designed from scratch.

The structural grouping still follows ISO/IEC/IEEE 15288:2023:
Agreement processes: acquisition and supply. In an AI context this is vendor assessment, model procurement terms and data usage rights.
Organizational project-enabling processes: life cycle model management, infrastructure management, portfolio management, human resource management, quality management, knowledge management. Compute resources and data platform governance sit here.
Technical management processes: project planning, project assessment and control, decision management, risk management, configuration management, information management, measurement, quality assurance. Model version control and experiment tracking belong to this group.
Technical processes: business or mission analysis, stakeholder needs and requirements definition, system requirements definition, architecture definition, design definition, system analysis, knowledge acquisition, AI data engineering, implementation, integration, verification, transition, validation, continuous validation, operation, maintenance and disposal. The AI-specific processes are added at this layer.

fig03_5338_process_groups
fig03_5338_process_groups

◆ 7.3 The three AI-specific processes: where audit meets genuinely new ground
The most important additions ISO/IEC 5338:2023 makes to ISO/IEC/IEEE 12207:2017 and 15288:2023 sit in three AI-specific technical processes. None of them has a counterpart in a traditional systems audit programme, and they are where audit most often has a blind spot.

【One】Knowledge acquisition (ISO/IEC 5338:2023 § 6.4.7)
ISO/IEC 5338:2023 draws on the definition in ISO/IEC 2382:2015: knowledge acquisition is the process of locating, collecting and refining knowledge, and converting it into a form a knowledge-based system can process further. It typically involves a knowledge engineer, but it is also an integral part of machine learning. Put simply, the behaviour of an AI system does not come only from code — it comes from knowledge, whether statistical patterns in training data, expert rules, annotation guidelines or domain ontologies. ISO/IEC 5338:2023 notes that knowledge acquisition matters even more for heuristic models, because there the knowledge is explicitly encoded into the model and directly determines its correctness.

Audit focus: whether annotation guidelines exist and are versioned; training and inter-annotator agreement for annotators; records of domain expert involvement; licensing and lawfulness of knowledge sources.

【Two】AI data engineering (ISO/IEC 5338:2023 § 6.4.8)
Acquisition, annotation, cleaning, splitting, augmentation and version control of data. ISO/IEC 5338:2023 separates this out precisely because, in a conventional software life cycle, data is merely input — whereas in a machine learning system the model's behaviour is not written but learned from data. The data is part of the implementation.

Audit focus: whether train, validation and test splits prevent data leakage; whether data lineage is traceable; whether personal and sensitive attributes are handled lawfully; whether datasets carry version identifiers that map to specific model versions. ISO/IEC 8183:2023, the AI data life cycle framework, breaks this segment down further.

【Three】Continuous validation (ISO/IEC 5338:2023 § 6.4.14)
This is the one with the greatest impact on audit, and the subject of the whole of the next section. Conventional systems enter a "maintenance" state after go-live, on the assumption that function does not change by itself; AI systems keep changing after go-live because of data drift, concept drift and upstream model updates. ISO/IEC 5338:2023 puts it bluntly: one of the differentiating characteristics of AI systems is measurable potential decay — the target behaviour the model is trying to capture shifts over time, and that shift can be measured through validation inputs and outputs. Behaviour changing is not unique to AI. Being measurable is.

Audit focus: whether performance decay thresholds are defined; whether drift detection has alerting and a named owner; whether retraining decisions carry approval records; whether revalidation results feed back into risk assessment.

◆ 7.4 The counterintuitive point: seven stages or eight? A widely misreported difference, and what it actually means for governance
This is the question I am asked most often, and probably the most confused area in AI circles. You have probably seen two claims. One says ISO/IEC 22989:2022 defines seven AI life cycle stages — this version circulates very widely, and even the AWS blog on ISO/IEC 42001:2023 governance uses a seven-stage diagram. The other says ISO/IEC 5338:2023 has eight. Many readers conclude that this is a difference between the two standards. That conclusion is wrong, and the real situation is more interesting.

【Fact one】ISO/IEC 5338:2023 has no stage list of its own
As noted in 7.1, clause 5.3 of ISO/IEC 5338:2023 states plainly that it does not prescribe a specific life cycle model. Its life cycle diagram is drawn on the basis of the figure in ISO/IEC 22989:2022, with the source acknowledged in the figure note. So "the eight stages of ISO/IEC 5338:2023" is, strictly speaking, not a thing — the stage list originates in ISO/IEC 22989:2022.

【Fact two】The text of ISO/IEC 22989:2022 does list eight stages

Open clause 6.2 of ISO/IEC 22989:2022 and the subclauses run in this order:
6.2.2 Inception
6.2.3 Design and development
6.2.4 Verification and validation
6.2.5 Deployment
6.2.6 Operation and monitoring
6.2.7 Continuous validation
6.2.8 Re-evaluation
6.2.9 Retirement

Eight, unambiguously. So "ISO/IEC 22989:2022 has seven stages" does not come from the text either.

【Fact three】The difference originates in a figure, not in a clause
So where do seven stages come from? From the example life cycle model figure in ISO/IEC 22989:2022. In that figure, continuous validation carries a conditional annotation to the effect of "in the case of continuous learning". It is drawn as a conditional stage rather than one every AI system passes through. When the figure is described second-hand, that stage is naturally dropped. Seven normal stages plus one conditional stage becomes, over time, simply "seven stages".

【Fact four】What ISO/IEC 5338:2023 did was remove that condition
Here is the crux. When ISO/IEC 5338:2023 drew its own life cycle figure, it kept the stages from ISO/IEC 22989:2022 but deliberately removed the "in the case of continuous learning" qualifier, and explained why in the figure note: continuous validation applies equally in situations without continuous learning — for example, to detect data drift, concept drift, or technical malfunctions. The accurate formulation is therefore this: the stage list comes from ISO/IEC 22989:2022 and has eight entries; the example figure in ISO/IEC 22989:2022 marks continuous validation as conditional, limited to continuous learning; ISO/IEC 5338:2023 removes that condition and treats continuous validation as a normal stage. The difference is not how many stages there are, but whether continuous validation is necessary.

【Why this difference is substantive for audit】
ISO/IEC 22989:2022 supplies a conceptual-layer AI system life cycle — its job is to give everyone a shared definition of the word "stage". If a life cycle already exists, why is ISO/IEC 5338:2023 needed at all?

The answer lies in that conditional annotation. The better you understand how AI systems actually behave in production, the more you find that assumptions which looked reasonable at the time lead to opposite audit conclusions in practice. Read as seven stages, the inference runs: our model does not learn continuously, the weights are frozen, therefore continuous validation does not apply to us — once acceptance testing is passed at go-live, we are in maintenance. That inference is persuasive, and my reading is that ISO/IEC 5338:2023 removed the conditional annotation precisely in order to close it off.

A frozen model does not mean a frozen world. Input distributions shift (data drift). The real relationship between features and target changes (concept drift). The upstream foundation model may be updated silently by the vendor. The data pipeline may degrade quietly because one field's format changed. None of this has anything to do with whether the model learns continuously. A credit scoring model that has not had a line of code changed in two years may already be operating on a completely different population.

For audit, this difference determines one very concrete judgement:
▸ On the seven-stage reading, a static model needs no periodic revalidation after go-live, and the absence of such evidence is not a finding.
▸ On the ISO/IEC 5338:2023 position, every AI system in production — learning or not — should have a continuous validation mechanism, and the absence of drift monitoring is a control deficiency.

This position is consistent with the wider regulatory direction. NIST AI 800-4, "Challenges to the Monitoring of Deployed AI Systems", published in March 2026, is the first US federal-level effort to map post-deployment monitoring gaps systematically, and its central argument is exactly this: pre-deployment evaluation is not sufficient to support risk decisions, and monitoring has to be a continuous practice rather than a one-time checkpoint. The small change ISO/IEC 5338:2023 made, in other words, is not a committee preference — it is a microcosm of where supervision is heading.

【One sentence for auditors】
When citing the stage model, attribute it to ISO/IEC 22989:2022. When citing the position that continuous validation is a normal stage, attribute it to ISO/IEC 5338:2023. Conflating the two will be challenged in a formal audit report.

fig04_22989_vs_5338_lifecycle
fig04_22989_vs_5338_lifecycle

◆ 7.5 The eight stages: putting processes on a timeline
With the argument settled, the stages can be used with confidence as the audit timeline:
① Inception
② Design and Development
③ Verification and Validation
④ Deployment
⑤ Operation and Monitoring
⑥ Continuous Validation
⑦ Re-evaluation
⑧ Retirement

ISO/IEC 5338:2023 adds a reminder that auditors easily overlook: stages exist only to group activities that have a temporal relationship, in order to show dependency. They do not imply that the activities are separated in time or across the organisation. In agile development, for example, development and operation are distinct stages that run concurrently. A feature must still be implemented before it can be verified and verified before it can be deployed — the dependency holds, but the stage boundary is blurred. The audit implication is that "did the stage gates pass in order" cannot be the only test. In a DevOps or MLOps environment, the absence of a signed document marked "stage three complete" does not by itself mean control failure; what should be tested is whether a dependency was skipped — whether any model version was deployed without verification and validation. ISO/IEC 5338:2023 also notes that different stages of the life cycle can be owned and managed by different organisations (data acquisition and provision, the machine learning model, or the code for other components may sit with different entities), and that an organisation may rely on others to build infrastructure or supply life cycle capability. This is the entry point for supply chain audit: your audit scope may cover only three of these eight stages.

The most underrated exposure for a deployer sits at stage ⑥. You changed nothing, but the vendor did — and the compliance obligation is still yours. This is also the most dangerous application of the "static models need no continuous validation" inference from 7.4: you believe the model is static, but it is not; the change is simply happening where you cannot see it. So the audit programme needs a line for it: is there a mechanism to detect vendor model version changes and assess their impact?

◆ 7.6 What to check at each stage, and what evidence to ask for
Expanding the eight stages into an audit programme produces the table below. The way to use it: first apply the scoping gateway from section 4 to establish whether you are a deployer or a provider, then decide how deep to go on each row. The table has three columns. Checkpoints and evidence are the parts most AI audit guidance already covers; the column worth studying is red flags — patterns that recur in practice but are hard to spot during document review. What they have in common is that the form is complete while the substance has failed, rather than nothing being there at all. An organisation with nothing in place is easy to audit; this category is the difficult one.

fig05_lifecycle_audit_programme
fig05_lifecycle_audit_programme

◆ 7.7 Role determines audit depth: two role vocabularies that must be kept apart
People often ask why a management system standard such as ISO/IEC 42001:2023 makes so much of roles. The reason is that the same life cycle table produces very different audits depending on who you are — a provider has to evidence technical assurance across the whole life cycle, while a deployer mainly has to show it chose the right vendor and is using the system in the right way. There is a trap here, though: AI governance actually runs on two role vocabularies, and they are not compatible.

The first comes from the EU AI Act and is binary: provider and deployer. Its purpose is to allocate legal liability, so the question it answers is "who is responsible when something goes wrong". The second comes from ISO/IEC 22989:2022 and has six roles: AI producer, AI provider, AI customer, AI partner, AI subject and relevant authorities. Its purpose is to describe an ecosystem, so the question it answers is "who is involved in this system".

A fair number of audit reports mix the two — EU AI Act provider obligations in the first half, the ISO/IEC 22989:2022 AI Provider role in the second. It reads smoothly, but the argument has already broken, because those two "providers" are not the same thing. The EU AI Act provider is the party that develops a system and places it on the market, which is closer to the AI Producer of ISO/IEC 22989:2022; the AI Provider of ISO/IEC 22989:2022 is the party supplying a platform or service. And the EU AI Act deployer maps to the AI Customer of ISO/IEC 22989:2022. Same names, different referents — and that will be challenged in a formal audit report. It is worth being clear that there is no official mapping table between these two vocabularies; the correspondences above are a practitioner's reading, not text from either document. Cite the source document, and do not write it as an equals sign.

【How to use them together】
My approach has two steps.
First, use the EU AI Act to establish your own legal role, which sets how deep the audit goes. A deployer can reduce technical testing of the model internals and strengthen scrutiny of vendor assurance; a provider carries technical assurance responsibility across the whole life cycle. This step determines depth.

fig06_deployer_vs_provider
fig06_deployer_vs_provider

Second, use ISO/IEC 22989:2022 to map who else sits on the supply chain, which determines who you ask for which evidence. Auditors in practice rarely face their own organisation's role — they face the roles upstream and downstream: which partner did the data come from, who trained the model, whose outcomes are affected. This step determines breadth.

【What to check for each of the six roles】

AI Producer: the party that designs, develops and trains the model. Audit focuses on the licensing chain for data sources, annotation guidelines and inter-annotator agreement, whether train and test splits prevent leakage, and whether the model card and design decision records were produced during development. The most common red flag is a model card written after handover — what the downstream party received is marketing material, not technical documentation.

AI Provider: the party supplying a platform or service. Audit focuses on whether the terms of service and SLA cover the obligation to notify model version changes, whether customer audit rights are granted, and the channel and deadlines for incident reporting. The red flag is an underlying model updated silently with no version mapping table available.

AI Customer: the party that bought it, which is the deployer under the EU AI Act. Audit focuses on whether actual use matches the vendor's stated intended purpose, input data controls, whether human overseers hold genuine veto authority, and the shadow AI inventory. The red flag is treating "we only bought it" as a defence while the escalation conditions of Article 25 have already been met.

AI Partner: the parties supplying data, systems integration, evaluation or infrastructure. This role is the one audit most often skips, because these parties are usually not the direct contracting counterparty. But after data has changed hands two or three times, the scope of the original consent often cannot be evidenced — and the liability lands on whoever used the data. Audit focuses on whether the licensing chain is traceable to the origin, the integrator's change management, and the independence of third-party evaluators.

AI Subject: the individuals or groups whose interests are affected by a decision. This is the only one of the six roles with no contractual relationship to the organisation, which is exactly why it gets missed. Audit focuses on whether people are told they are interacting with AI, whether appeal and human review channels are reachable and owned by someone who closes them, and whether the decision record retention period covers the appeal window. There is a classic cross-failure here: the appeal window is set by regulation, the destruction schedule is set by the retention policy, each is individually compliant, and together they leave the organisation unable to answer an appeal.

Relevant Authorities: the bodies that make the rules and police the boundaries. This role is not the subject of the audit, so what gets tested is how the organisation responds to it: who owns regulatory tracking, how often it runs, whether risk classification is reassessed when law changes, and whether reporting deadlines and contact points are clearly assigned. The red flag is a regulatory change six months old with the risk classification table still on the previous version and nobody assigned to track it.

fig07_22989_stakeholder_roles
fig07_22989_stakeholder_roles

As the AI supply chain lengthens — foundation model vendors, fine-tuning services, systems integrators, outsourced annotation — the portion any single organisation can audit directly keeps shrinking. That is the value of the six-role model in ISO/IEC 22989:2022: it reminds auditors that their scope may cover only a short segment of the chain, and that the rest has to be covered by contract terms and vendor assurance. If audit rights and change notification obligations are not written into the contract, there will be nothing to find later.

■ 8. Stacking the standards: where ISO/IEC 5338:2023 sits on the map

Adopting ISO/IEC 5338:2023 on its own achieves little; it needs the others around it. This is the stack I currently work with:
▸ ISO/IEC 42001:2023 — AI management system (certifiable, organisation layer)
▸ ISO/IEC 5338:2023 — AI system life cycle processes (system layer, the focus of this article)
▸ ISO/IEC 23894:2023 — guidance on AI risk management (aligned to ISO 31000:2018)
▸ ISO/IEC 22989:2022 — AI concepts and terminology (the actual source of the life cycle stages)
▸ ISO/IEC 23053:2022 — framework for AI systems using machine learning
▸ ISO/IEC 8183:2023 — AI data life cycle framework (breaks down stages ②③ in more detail)
▸ ISO/IEC 42005:2025 — AI system impact assessment (echoes the FRIA logic of the EU AI Act)
▸ ISO/IEC TR 5469:2024 — functional safety and AI systems (pointed to explicitly by ISO/IEC 5338:2023)
▸ NIST AI RMF 1.0 (NIST AI 100-1) — Govern / Map / Measure / Manage

The relationship between the NIST AI RMF and ISO/IEC 5338:2023 is often misread as competition. They operate at different levels of abstraction: the AI RMF organises risk management functions, ISO/IEC 5338:2023 organises engineering processes. Audit can use the four RMF functions as the skeleton of the risk narrative and the eight stages from ISO/IEC 22989:2022 and ISO/IEC 5338:2023 as the timeline — crossed together, they make a reasonably complete audit matrix. It is also worth noting that where your sector has its own life cycle standard (IEC 62304:2006 for medical device software, for instance), ISO/IEC 5338:2023 recommends applying AI-specific considerations alongside it rather than choosing between them. For heavily regulated industries this is an important sentence — it means adopting ISO/IEC 5338:2023 does not require dismantling the compliance architecture already in place.

fig08_standards_stack
fig08_standards_stack

■ 9. Conclusion: the route to responsible AI

AI governance is not an obstacle to innovation; it is the foundation of durable trust. As global supervision tightens (the EU AI Act, FINMA Guidance 08/2024 in Switzerland), organisations need to recognise that only controlled AI is sustainable AI. Internal audit should act as the organisation's compass, steering it steadily through a shifting legal framework and a fast-moving technology. Doing that takes more than principles — it takes a life cycle axis that can be worked through stage by stage, which is exactly what ISO/IEC 5338:2023 contributes.

Looking ahead, AI audit will move in two directions over the next three years.
The first is evidence automation — monitoring logs, model versions and validation reports will flow directly from MLOps platforms into GRC systems, and audit will shift from requesting working papers to querying them.
The second is continuous auditing — if the model changes continuously, the annual audit rhythm has already failed, and audit frequency has to keep pace with retraining frequency.

That is the same logic ISO/IEC 5338:2023 applied when it promoted continuous validation from a conditional stage to a normal one, just moved from the system layer up to the audit layer. A closing question: is your organisation ready, before the next supervisory review, to demonstrate that its AI decisions are traceable, explainable and legally accountable? Start with the AI inventory and the accountability matrix, and do not let a governance vacuum become the thing that kills the innovation.

■ References
▸ISO/IEC 5338:2023, Information technology — Artificial intelligence — AI system life cycle processes
▸ISO/IEC 22989:2022, Information technology — Artificial intelligence — Artificial intelligence concepts and terminology
▸ISO/IEC 23053:2022, Framework for Artificial Intelligence (AI) Systems Using Machine Learning (ML)
▸ISO/IEC 23894:2023, Information technology — Artificial intelligence — Guidance on risk management
▸ISO/IEC 8183:2023, Information technology — Artificial intelligence — Data life cycle framework
▸ISO/IEC 42001:2023, Information technology — Artificial intelligence — Management system
▸ISO/IEC 42005:2025, Information technology — Artificial intelligence — AI system impact assessment
▸ISO/IEC TR 5469:2024, Artificial intelligence — Functional safety and AI systems
▸ISO/IEC 2382:2015, Information technology — Vocabulary
▸ISO/IEC/IEEE 15288:2023, Systems and software engineering — System life cycle processes
▸ISO/IEC/IEEE 12207:2017, Systems and software engineering — Software life cycle processes
▸ISO 31000:2018, Risk management — Guidelines
▸IEC 62304:2006, Medical device software — Software life cycle processes
▸Regulation (EU) 2024/1689 (EU AI Act), Articles 4, 25, 50
▸NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0)
▸NIST AI 800-4, Challenges to the Monitoring of Deployed AI Systems (2026)
▸FINMA Guidance 08/2024, Governance and risk management in the use of artificial intelligence
▸Global Internal Audit Standards (GIAS 2024), The IIA
▸COSO, Internal Control — Integrated Framework: Generative AI considerations

← Back to Blog