Meridian

The Operating-Effectiveness Gap

Why Adopting an AI Governance Framework Is Not the Same as Governing AI, and What Boards Should Ask Instead

AI Q2 2026
Executive Summary

EXECUTIVE SUMMARY

The thesis of this paper is that AI governance has solved adoption and left operation unproven, and that the distance between the two is widening rather than closing. Organisations are adopting AI governance frameworks at speed. The NIST AI Risk Management Framework, ISO/IEC 42001, and the risk-management obligation in Article 9 of the European Union’s AI Act now anchor governance programmes across the United States, Europe, and the Gulf, and by late 2025 the International Association of Privacy Professionals reported that 77% of surveyed organisations were building AI governance programmes, a figure rising toward 90% among those already deploying AI systems. Adoption, in other words, is no longer the question.

Almost no organisation can produce evidence that its governance programme actually governs AI systems once those systems are live. The artefacts of adoption, a chartered committee, a policy mapped to a framework, an attestation, are real and necessary, and they are also precisely what surveys count and what most attestations report; they are not evidence that drift is detected, that incidents are caught before harm, that override authority is exercised, or that a board sees any of it. McKinsey’s 2026 State of AI research found that only about one in four organisations have fully operational AI governance, and that enterprises typically adopt AI capabilities two to three years ahead of the risk-management structures meant to oversee them.

This paper names the distance between those two states the Operating-Effectiveness Gap. The term is imported deliberately from the internal-control discipline, where auditing authorities have long separated design effectiveness, whether a control is well-designed and in place, from operating effectiveness, whether the control actually runs as intended over a period of time. AI governance today is rich in design evidence and poor in operating evidence, and the gap is widening because adoption scales faster than operation: agent deployment roughly doubled across 2024 and 2025, while only about one in five organisations reported a mature model for governing those agents.

Against that gap the paper sets two instruments. The Control-Maturity Curve locates where a programme actually sits across five stages, from Adoption to Adaptation, by the kind of evidence each stage can produce. The Evidence-of-Operation Test, a board diagnostic, shifts the director’s question from whether a programme exists to what evidence demonstrates that it operates. The paper works both tools against current evidence from four jurisdictions and bounds them with stated limits. Its conclusion is that the regulatory, standards, and enforcement environment is converging on operating effectiveness as the thing that matters, that most programmes are not yet built to evidence it, and that the European Union’s pending deferral of its high-risk deadline extends the runway without closing the gap.

INTRODUCTION

The governance of artificial intelligence has reached an unusual moment: the frameworks exist, they are broadly respected, and they are being adopted in large numbers. The NIST AI Risk Management Framework, published in January 2023, set out a voluntary architecture organised around four functions, Govern, Map, Measure, and Manage. ISO/IEC 42001, published in 2023 as the world’s first certifiable AI management system standard, built a Plan-Do-Check-Act improvement cycle against which an organisation can be audited. The European Union’s AI Act, Regulation (EU) 2024/1689, made risk management a legal duty for high-risk systems under Article 9 and required post-market monitoring under Article 72. The scaffolding is in place across the voluntary, the certifiable, and the binding.

What has not kept pace is the evidence that any of this scaffolding bears weight once a system is deployed. A policy can be written and a committee chartered in a quarter; a model inventory that is complete and current, monitoring that detects drift, an override that someone has both the authority and the willingness to pull, and a board that can see whether these things are working, take far longer and are far harder to demonstrate. The result is a governance landscape in which adoption is visible, operation is largely invisible, and the two are routinely conflated.

That conflation is not a matter of bad faith. It is a structural feature of how governance maturity is measured. Surveys ask whether an organisation has a policy, a framework, a committee, an officer; attestations describe the design of a control environment. These are the natural units of self-report, and they all sit at the adoption end of the maturity range. The harder evidence, that a control produced a detection, that an incident was caught, that a decision was reversed, is rarely asked for and rarely volunteered.

This paper treats that distance as a measurable object rather than a rhetorical one. It imports a distinction the audit profession settled decades ago, applies it to AI governance, and tests whether the imported distinction holds against the current evidence. The distinction is between design effectiveness and operating effectiveness; the object it names is the Operating-Effectiveness Gap; and the aim is to give boards, risk officers, and regulators a precise way to ask whether a governance programme does what its adoption implies it does.

I. THE EVIDENCE: ADOPTION IS HIGH, OPERATION IS UNPROVEN

The adoption numbers are unambiguous. By the IAPP’s 2025 AI Governance Profession Report, drawing on more than 670 professionals across 45 countries, 77% of organisations surveyed were working on AI governance, a figure that rose toward 90% among organisations already using AI. Governance has become a near-universal activity among AI adopters, and the framework landscape supports it: NIST’s RMF is the voluntary reference point in the United States, ISO/IEC 42001 offers a certifiable management system internationally, and the EU AI Act makes parts of the discipline mandatory.

The operating numbers tell a different story. McKinsey’s 2026 State of AI research found that only about one in four organisations have fully operational AI governance, and that organisations typically run two to three years ahead of their own risk-management structures. That lag is not incidental; it is the normal condition of an environment in which the capability is acquired before the controls that should accompany it.

The agent data sharpens the point. Deloitte’s State of AI in the Enterprise research found that worker access to enterprise AI tools rose by roughly 50% in a single year, and that only about one in five organisations reported a mature model for governing autonomous AI agents. KPMG’s Global AI Pulse for the first quarter of 2026 found 54% of organisations at the deployment stage for AI agents, and recorded that the share requiring human validation of agent outputs had risen to 63%, up from 22% a year earlier. That last movement is one of the few clean operating-effectiveness signals in the public data, a control, human validation, being wired into a deployed system at scale and rising sharply year over year, and it is the exception that illustrates the rule; most of what is reported is access, deployment, and intent, the vocabulary of adoption.

PwC’s 2026 Digital Trends in Operations survey, drawing on 767 operations and supply-chain leaders at United States companies, captured the self-assessment problem directly. While 85% of leaders judged themselves ahead of most competitors in digital transformation, 89% reported that their technology investments had not fully delivered the expected results, and 87% identified poor data quality as a constraint on value; only 30% reported significant improvement in data quality over a multi-year horizon, and only 27% had fully embedded an AI strategy across business units. The distance between the 85% who believe they lead and the operating reality the same survey describes is, in miniature, the Operating-Effectiveness Gap.

None of these figures, taken alone, proves a governance failure. Taken together, they describe a consistent pattern: high and rising adoption, fast and rising deployment, and operating evidence that is thin, self-assessed, and lagging, and that pattern is the empirical ground on which the rest of this paper stands.

II. THE FRAMEWORK: THE OPERATING-EFFECTIVENESS GAP

The distinction this paper imports is not new; it is foundational to the audit of internal control. Under the Public Company Accounting Oversight Board’s Auditing Standard 2201, an audit of internal control over financial reporting requires testing and evaluating both the design effectiveness and the operating effectiveness of a control. Design effectiveness asks whether a control, operating as prescribed, would prevent or detect the error it targets; operating effectiveness asks whether the control actually operated that way, by whom, and with what competence, over the period under examination. A control can be well-designed and never run, and the audit profession learned long ago that the second question is the one that protects against harm and that it cannot be answered by inspecting the first.

The Operating-Effectiveness Gap is the application of that distinction to AI governance. Nearly every artefact an AI governance programme produces is design evidence: a policy describes how a control should work, a framework mapping shows that a control has been contemplated, a committee charter establishes who is accountable in principle, an attestation describes a control environment as designed. These are the elements that surveys count and that maturity self-assessments reward, and none of them is operating evidence. Operating evidence is a detection that fired, an incident that was caught, an override that was exercised, an inventory that was reconciled against live systems and found complete. The gap is the distance between the design evidence a programme can readily produce and the operating evidence the duty actually requires.

The distinction is now being drawn inside the AI governance literature itself, not merely imported into it. In February 2026 the Committee of Sponsoring Organizations of the Treadway Commission, the body whose internal-control framework underpins much of the audit profession’s practice, released guidance titled Achieving Effective Internal Control Over Generative AI. Rather than build a new model, COSO adapted its five existing internal-control components to generative AI and emphasised ongoing inventories of AI use cases, clear ownership and escalation paths, and continuous monitoring of control performance, with illustrative metrics meant to support audit evidence collection; the guidance treats operating evidence as the object of governance, not its by-product. The Institute of Internal Auditors, through its AI Auditing Framework and its application of the Three Lines Model, carries the same logic into the assurance function.

A useful way to see the gap is to hold four core AI controls against the two standards of evidence. A model inventory can be designed, a register exists, and still fail to operate if it does not reconcile against the systems actually running in production. Model-drift and incident detection can be designed, monitoring is specified, and still fail to operate if no alert has ever fired or been actioned. Override authority can be designed, a policy names who may halt a system, and still fail to operate if no one has ever exercised it or would be supported in doing so. Adversarial robustness can be designed, testing is required, and still fail to operate if testing is one-time rather than continuous against evolving threats. In each case the design artefact is easy to produce and the operating evidence is what the duty actually demands, and the framework’s contribution is to insist on the second column.

The framework must be bounded. It does not claim that design evidence is worthless; a control that is not designed cannot operate, so adoption is necessary, merely insufficient. It does not claim that every organisation reporting adoption has failed to operate; some have, and the public data cannot distinguish them well, which is itself part of the problem. And it imports a distinction from financial-control auditing into a domain, AI behaviour, where controls are probabilistic and the period under examination never truly closes, because the model and its environment keep changing. Those are real limits, and they shape the instrument that follows.

III. THE INSTRUMENT: THE CONTROL-MATURITY CURVE

If the gap is the object, the Control-Maturity Curve is the instrument for locating a programme along it. The curve describes five stages, each defined by the kind of evidence it can produce.

1. Adoption.

The artefacts exist: a policy is written, a framework is mapped, a committee is chartered, an officer is named. This is the stage that surveys and most attestations measure, and it is where the majority of the reported activity sits; the IAPP’s 77% lives here. Adoption is genuine progress, and it is the floor rather than the achievement.

2. Instrumentation.

Controls are wired to live systems: the model inventory is complete and reconciled, monitoring and logging are connected to systems in production, and named owners exist for specific controls rather than for the programme in the abstract. Instrumentation is the bridge between having a control on paper and having it produce evidence, and PwC’s finding that poor data quality constrains 87% of organisations is, in large part, a failure to reach this stage, because controls cannot be wired reliably to systems whose data cannot be trusted.

3. Operation.

The controls run: drift and incident detection produce and action alerts, override authority is exercised when conditions require it, and issues are caught before they cause harm. This is the stage at which the governance duty is actually met, and it is the stage the public evidence most rarely demonstrates; KPMG’s finding that 63% of organisations now require human validation of agent outputs, up sharply from 22%, is one of the few visible signals of programmes reaching operation at scale.

4. Oversight.

Operating effectiveness becomes visible upward: the board and the chief risk officer see evidence of operation, detections, incidents, overrides, reconciliations, rather than evidence of adoption. This is a distinct stage because a programme can operate without its operation being visible to those accountable for it, and a board that receives only adoption metrics cannot discharge its duty regardless of how well the programme runs beneath it.

5. Adaptation.

The programme updates as models drift and rules move, and the post-market monitoring loop closes: findings from deployed systems feed back into design, controls are revised, and the cycle repeats. This is the stage toward which the EU AI Act’s Article 72 post-market monitoring obligation, ISO/IEC 42001’s improvement cycle, NIST’s Manage function, and COSO’s monitoring component all point, and it is where governance becomes a living system rather than a static control set.

The Operating-Effectiveness Gap is the distance between stage one, where adoption metrics and most attestations sit, and stages three and four, where the duty is met and seen to be met. The curve is not a maturity-model marketing device promising that every organisation should reach stage five; it is a diagnostic for honesty, forcing the question of which stage a programme can actually evidence and denying the programme credit for stages it has only adopted on paper. Its limit is that the stages are not strictly sequential in practice, since an organisation may instrument some controls while leaving others at adoption, and the curve should therefore be read control by control rather than as a single programme-wide grade.

IV. THE BOARD DIAGNOSTIC: THE EVIDENCE-OF-OPERATION TEST

For a board, the framework reduces to a single shift in the question asked. The familiar question is whether the organisation has an AI governance programme; by 2026 the answer is almost always yes, and almost always uninformative, because it reports adoption. The Evidence-of-Operation Test replaces it with a question that adoption cannot satisfy: what evidence does the organisation have that the programme operates?

The test is a set of demands for operating evidence rather than design evidence. For the model inventory, the question is not whether one exists but when it was last reconciled against systems in production and what that reconciliation found. For drift and incident detection, the question is not whether monitoring is specified but how many alerts fired in the last period, how many were actioned, and what the response time was. For override authority, the question is not whether a policy names who may halt a system but whether anyone has exercised it, under what circumstances, and whether they were supported. For post-market monitoring, the question is not whether a plan exists but what the deployed systems have reported back and what changed as a result.

Each of these questions is answerable only with operating evidence, and each exposes the gap when only design evidence is available. A board that receives a confident yes to every adoption question and cannot obtain a specific answer to a single operation question has located its own programme on the Control-Maturity Curve, at stage one, regardless of what the policy binder suggests. The test is deliberately uncomfortable, because the discomfort is diagnostic; the value of the question is precisely that it cannot be answered by the artefacts that satisfy a survey

The test also clarifies the board’s own stage. A board operating at stage four receives operating evidence as a matter of routine reporting; a board that receives only programme existence, framework adoption, and committee activity is being shown stage one and should recognise as much. The director’s duty of oversight, in this framing, is not discharged by confirming that a programme exists; it is discharged by confirming that the programme operates and by being shown how that is known.

V. THE LONGITUDINAL VIEW: THE GAP IS WIDENING, 2023 TO 2026

The distinctive claim of this paper is not that a gap exists at a moment but that it has widened across four years, because adoption and deployment have accelerated faster than operation. The longitudinal evidence, assembled from successive survey waves and the standards timeline, supports that trajectory while carrying the limits that cross-survey comparison always carries.

The 2023 baseline is one of frameworks arriving faster than the capacity to operate them. NIST published its AI RMF in January 2023; ISO/IEC 42001 was published as the first certifiable AI management system standard later that year. These are adoption instruments, giving organisations something to adopt, map to, and attest against, and their arrival in 2023 marks the start of the adoption surge, precisely because adoption is the stage that scales quickly, its artefacts, policies and mappings and charters, being producible quickly.

Across 2024 and 2025, deployment accelerated. Deloitte’s research recorded worker access to enterprise AI tools rising by roughly 50% in a single year; KPMG recorded agent deployment reaching 54% of organisations by early 2026 and projected average AI spending of roughly USD 207 million over the following year, close to double the level of the prior year; and the EU AI Act entered into force in 2024, converting parts of the discipline from voluntary to binding and accelerating formal adoption further. Every one of these movements is an adoption-and-deployment movement: the capability and its formal governance scaffolding both scaled.

Operation did not scale at the same rate. By early 2026 Deloitte found only about one in five organisations with a mature model for governing the agents they were deploying at 54% penetration; McKinsey found only about one in four with fully operational governance, and a two-to-three-year lag between capability and control; PwC found 89% of technology investments underdelivering and 87% constrained by data quality, the substrate on which instrumentation depends. The single clearest counter-signal, KPMG’s human-validation requirement rising from 22% to 63%, shows operation beginning to catch up on one specific control, which is encouraging precisely because it remains, so far, the exception.

The trajectory, then, is divergence. Adoption rose from a 2023 baseline toward near-universality among AI users by 2025; deployment, especially of agents, accelerated sharply, with a majority of organisations deploying agents and AI spending close to doubling year over year; operation, measured by mature agent governance and fully operational governance, remained a minority position into 2026. The gap between the curves widened. This longitudinal reading carries an explicit limit: the surveys are run by different organisations, with different samples, populations, and definitions, so the comparison is directional rather than precise, and it establishes a trajectory rather than a measured rate. But the direction is consistent across every independent source, and consistency across independent instruments is the strongest claim survey evidence can support.

VI. THE REGULATORY CONVERGENCE ON OPERATION

The supervisory environment is moving, unevenly but unmistakably, toward operating effectiveness as the thing that is checked. This convergence is the external pressure that will eventually make the Operating-Effectiveness Gap a compliance exposure rather than a governance abstraction.

In the European Union, the AI Act builds operation into the legal text. Article 9 requires a risk-management system for high-risk AI that runs across the lifecycle rather than a one-time assessment, and Article 72 requires providers to operate a post-market monitoring system and to collect and review data on deployed performance; these are operating obligations by construction, since a static risk assessment and an unmonitored deployment do not satisfy them. The timeline, however, is in flux. On 7 May 2026 the Council and Parliament reached provisional political agreement on the Digital Omnibus on AI, which would defer the high-risk obligations for standalone Annex III systems from 2 August 2026 to 2 December 2027, and the European Parliament approved the agreement in plenary on 16 June 2026. As of this writing the measure has not been formally adopted by the Council or published in the Official Journal, and until it is, the deferral is not in force: 2 August 2026 remains the legally operative high-risk deadline. The proposed deferral, if adopted, extends the runway; it does not change the destination, and a board that treats a provisional deferral as settled relief is governing to a date that has not legally moved, which is the precise error this paper warns against.

In the United States, the architecture is voluntary at the framework level and enforced at the conduct level. NIST’s RMF and its July 2024 Generative AI Profile, which sets out twelve generative-AI risk areas and more than 200 suggested actions across the RMF functions, supply the operating vocabulary without mandating it; enforcement supplies the consequence. In March 2024 the Securities and Exchange Commission settled its first AI-washing cases against two investment advisers, Delphia and Global Predictions, imposing civil penalties of USD 225,000 and USD 175,000 respectively for false and misleading statements about their use of AI. These are adoption-versus-reality cases in a different register: firms claimed AI capabilities they could not evidence. The Federal Trade Commission’s market-study orders, including its January 2024 inquiry into generative-AI partnerships and its 2024 surveillance-pricing study, extend the same scrutiny to how AI is actually used rather than how it is described.

In the United Kingdom, the approach is regulator-led and standards-driven rather than statutory. The Department for Science, Innovation and Technology set out a pro-innovation, principles-based framework in 2023 and issued implementing guidance and a government response in 2024, followed by a government AI Playbook in 2025; the Information Commissioner’s Office, in its 2024 audit framework and its AI and data protection risk toolkit, supplies the operating instruments, step-by-step assessments of whether AI processing actually meets data-protection obligations in practice. The United Kingdom route reaches operating effectiveness through audit and guidance rather than through a high-risk regime.

The standards and the international guidance complete the picture. ISO/IEC 42006, the 2025 standard for bodies that certify AI management systems, is significant precisely because it concerns who may verify operation and how, moving operating effectiveness toward third-party assurance; the OECD’s February 2026 Due Diligence Guidance for Responsible AI applies the six-step responsible-business-conduct due-diligence framework to AI, embedding ongoing identification, mitigation, and remediation of adverse impacts, which is an operating discipline rather than an adoption one. Across the European, American, British, and international tracks, the direction is the same: the thing increasingly checked is whether the programme runs.

VII. THE GULF AS GOVERNMENT-SCALE DEPLOYER

The Gulf states present the adoption-versus-operation question in an unusual form, because here the government is frequently the deployer at scale rather than the regulator of private deployers. That inverts the usual relationship and makes operating effectiveness a question about the state’s own systems.

Saudi Arabia’s Saudi Data and Artificial Intelligence Authority adopted its Principles and Controls of AI Ethics on 14 September 2023, a principle-based national framework covering the AI lifecycle, with a risk classification and governance expectations including senior leadership engagement and accountability up to the head of the entity. SDAIA subsequently moved to institutionalise adoption across government, announcing in September 2024 the activation of AI offices across 23 government entities and publishing maturity-oriented adoption frameworks. This is adoption pursued deliberately and at national scale, and the open question, the same one this paper raises everywhere, is what operating evidence those offices and frameworks produce once the deployed government systems are live, evidence that is not yet public.

The United Arab Emirates approaches the question through its financial free zones, where the regulatory craft is most developed. The Dubai International Financial Centre’s Data Protection Regulations, in force from September 2023, include Regulation 10 on the processing of personal data through autonomous and semi-autonomous systems, and Regulation 10 is notable for reaching toward operation: it requires controllers and processors to assess, document, and mitigate privacy risks before deploying such systems, to maintain registers, and, for high-risk processing, to designate accountable roles for autonomous systems. The DIFC’s Consultation Paper No. 3 of 2026, published in June 2026, proposes to strengthen this regime further, clarifying certification obligations and the role of the Autonomous Systems Officer and proposing a new Regulation 11 empowering the Commissioner to recognise accreditation and certification schemes. The Abu Dhabi Global Market’s Financial Services Regulatory Authority has pursued a different instrument, its OpenReg initiative providing machine-readable regulation and supervisory AI models, an attempt to build operating-effectiveness tooling into the supervisory relationship itself.

The Gulf comparison should be weighted carefully and not overclaimed. The frameworks are recent, the deployments are large and often governmental, and the public operating evidence is limited, as it is everywhere. What the Gulf adds to the analysis is the government-as-deployer dimension: where the state deploys AI at scale, the Operating-Effectiveness Gap becomes a question about public systems, and the absence of public operating evidence becomes a matter of public, not merely corporate, accountability. The frameworks adopted are serious; whether they operate is, as everywhere in this paper, the unproven part.

VIII. LEADERSHIP AND THE ACCOUNTABILITY FOR OPERATION

The gap is ultimately an accountability problem, and accountability for operation sits with senior leadership and the board. The frameworks already say so: the EU AI Act assigns provider and deployer obligations that cannot be delegated away, NIST’s Govern function places governance at the centre of the RMF rather than at its periphery, SDAIA’s principles run accountability up to the head of the entity, and COSO’s control environment, the first of its five components, is a leadership construct. The instruments converge on the same point, that operation is a leadership duty rather than a technical one.

The reason leadership matters specifically for operation, as distinct from adoption, is that adoption can be delegated and operation cannot. A policy can be commissioned, a framework can be mapped, a committee can be stood up, all without sustained senior attention; operation requires something leadership alone can supply, the authority to halt a deployed system, the willingness to support the person who exercises that authority, and the insistence on seeing operating evidence rather than accepting adoption evidence. An override authority that no one will back is not an operating control. The COSO guidance’s emphasis on clear ownership and escalation paths, and the IAPP’s finding that fragmented ownership and unclear accountability are among the impediments organisations report, both point at the same failure mode: controls that exist on paper and have no operational owner with real authority.

This is why the Evidence-of-Operation Test is addressed to the board. The board cannot write the monitoring code or reconcile the inventory, but it can refuse to accept adoption evidence in place of operating evidence, and that refusal is the single most consequential governance act available to it. When a board asks what evidence demonstrates that the programme operates, and declines to be satisfied by confirmation that the programme exists, it changes what the organisation beneath it must produce; demand for operating evidence is what moves a programme up the Control-Maturity Curve. The board’s leverage is not technical. It is the leverage of the question it is willing to insist on.

CONCLUSION

AI governance has solved the adoption problem and not yet confronted the operation problem. Frameworks are in place, broadly respected, and widely adopted, and adoption is now near-universal among organisations that deploy AI; the evidence that those programmes operate once systems are live, that drift is detected, incidents are caught, overrides are exercised, and boards can see all of it, remains thin, self-assessed, and lagging. The distance between the two, the Operating-Effectiveness Gap, has widened across four years because adoption and deployment scaled faster than operation, and the European Union’s pending deferral of its high-risk deadline, if adopted, would extend the interval in which that divergence can continue without external consequence.

The instruments offered here, the Control-Maturity Curve and the Evidence-of-Operation Test, are not predictions of failure. They are tools for honesty: they locate a programme by the evidence it can actually produce, deny it credit for stages it has only adopted on paper, and give a board a question that adoption cannot answer. The regulatory and standards environment is converging on operating effectiveness as the thing that will be checked; organisations that can already evidence operation will meet that environment from a position of strength, and those that have evidenced only adoption will discover, when the question is finally asked of them in earnest, how far apart the two have grown.

RECOMMENDATIONS

The following are prioritised and traceable to the analysis above. The highest-priority actions are those that convert design evidence into operating evidence, because that conversion is where the gap closes.

For boards and audit committees:

-  Adopt the Evidence-of-Operation Test as a standing item. Require, for each material AI system, operating evidence rather than programme-existence evidence: the date of the last inventory reconciliation and its findings, alert counts and response times, instances of override exercised, and post-market monitoring results.

- Refuse to accept adoption metrics as oversight evidence. Where reporting describes the framework adopted and the committee chartered without describing what the controls have detected or done, treat the programme as located at stage one regardless of its documentation.

- Confirm the board’s own stage. Establish whether the board receives operating evidence routinely, at stage four, or only programme-existence reporting, at stage one, and close that gap first.

For chief risk officers and heads of governance:

- Locate the programme on the Control-Maturity Curve control by control, not as a single grade. Identify which controls are merely adopted, which are instrumented, and which actually operate, and prioritise instrumenting the controls that protect against the highest-consequence harms.

- Treat data quality as a governance prerequisite rather than a separate workstream. Where controls cannot be reliably wired to systems because the underlying data cannot be trusted, instrumentation fails and operation is impossible.

- Establish operational ownership with real authority, including override authority that leadership will demonstrably support, since an override no one will back is not an operating control.

For policy and standards engagement:

- Align internal evidence to the post-market monitoring direction of the EU AI Act, ISO/IEC 42001’s improvement cycle, NIST’s Manage function, and COSO’s monitoring component, so that operating evidence accumulates in a form that will satisfy external assurance when it is required.

- Treat the European Union’s proposed high-risk deferral as runway to reach operation rather than as relief, and continue to plan to the operative 2 August 2026 deadline until any deferral is actually adopted and published; the destination is unchanged.

Benchmarks that would change these recommendations:

- If future survey waves show mature operational governance becoming a majority position rather than a minority one, the emphasis should shift from reaching operation to sustaining and adapting it, at stage five.

- If a regulator begins routinely demanding and penalising on operating evidence rather than adoption evidence, the Evidence-of-Operation Test moves from prudent governance to compliance necessity, and the timeline for instrumentation compresses accordingly.

- If standardised, audited operating metrics emerge, for example through ISO/IEC 42006 certification practice, self-assessed maturity data should be discounted in favour of third-party-verified operating evidence.

CAVEATS

1. The public evidence is dominated by self-assessment. Most maturity and adoption data is self-reported, and self-report systematically favours adoption evidence over operating evidence, so the gap this paper describes may be wider than the data shows rather than narrower.

2. Cross-survey comparison is directional, not precise. The longitudinal reading draws on surveys run by different organisations with different samples, populations, and definitions, and it establishes a trajectory rather than a measured rate.

3. The imported distinction has limits. Design versus operating effectiveness comes from financial-control auditing, where the period under examination closes; in AI governance the model and its environment change continuously, so operating effectiveness is a moving target rather than a fixed one.

4. Absence of public operating evidence is not proof of operating failure. Some organisations operate their controls well and simply do not disclose the evidence, and the paper claims that operation is unproven in the public record rather than that it is universally absent.

5. The regulatory timeline is unsettled. The European Union’s Digital Omnibus deferral had been approved by the European Parliament on 16 June 2026 but, at the time of writing, had not been formally adopted by the Council or published in the Official Journal, so it was not yet in force and the 2 August 2026 deadline still formally applied; the dates cited may move on adoption.

6. The Gulf analysis is weighted as comparative and limited by disclosure. Government-scale deployment frameworks are recent and their operating evidence is largely not public, so the section identifies the structural question rather than resolving it.

7. The framework is an analytical instrument, not an empirical measurement. The Control-Maturity Curve locates programmes by the evidence they can produce; it does not assign a validated, calibrated score, and its stages are not strictly sequential in practice.

8. The survey figures carry their own definitional ambiguity. Terms such as “mature governance,” “operational,” and “deploying agents” are defined differently across sources, and the paper uses them as the sources do rather than reconciling them to a single definition.

CITATION LIST

1. NIST, AI Risk Management Framework 1.0, 26 January 2023. https://www.nist.gov/itl/ai-risk-management-framework

2. NIST, AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, 26 July 2024. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf

3. ISO/IEC 42001:2023, Information technology — Artificial intelligence — Management system, 2023. https://www.iso.org/standard/42001

4. ISO/IEC 42006:2025, Requirements for bodies providing audit and certification of AI management systems, 2025. https://www.iso.org/

5. Regulation (EU) 2024/1689 (EU AI Act), 13 June 2024, OJ 12 July 2024. https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng

6. EU Digital Omnibus on AI, provisional political agreement 7 May 2026; approved by the European Parliament in plenary 16 June 2026; not yet formally adopted by the Council or published in the Official Journal as of late June 2026 (as reported by the Council of the EU, the European Parliament, Gibson Dunn, and White & Case). https://www.consilium.europa.eu/en/press/press-releases/2026/05/07/artificial-intelligence-council-and-parliament-agree-to-simplify-and-streamline-rules/

7. IAPP, AI Governance Profession Report 2025, November 2025. https://iapp.org/resources/article/ai-governance-profession-report

8. McKinsey, The State of AI (2026 wave), 2026. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-how-organizations-are-rewiring-to-capture-value

9. Deloitte, State of AI in the Enterprise (2025/2026 wave). https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/content/state-of-ai-in-the-enterprise.html

10. KPMG, Global AI Pulse Survey, Q1 2026, March 2026. https://kpmg.com/xx/en/our-insights/ai-and-technology/ai-pulse.html

11. PwC, 2026 Digital Trends in Operations Survey, January 2026, n=767 US operations and supply-chain leaders. https://www.pwc.com/us/en/services/consulting/business-transformation/library/digital-trends-operations-survey.html

12. PCAOB, Auditing Standard 2201, An Audit of Internal Control Over Financial Reporting That Is Integrated with an Audit of Financial Statements. https://pcaobus.org/oversight/standards/auditing-standards/details/AS2201

13. COSO, Achieving Effective Internal Control Over Generative AI, February 2026. https://www.coso.org/

14. The IIA, AI Auditing Framework, September 2024, and COSO GenAI coverage, March 2026. https://internalauditor.theiia.org/en/articles/2026/march/coso-issues-genai-guidance/

15. SEC, Press Release 2024-36, SEC Charges Two Investment Advisers Over AI Statements, 18 March 2024. https://www.sec.gov/newsroom/press-releases/2024-36

16. SEC, Administrative Order, In re Global Predictions, Inc., 18 March 2024. https://www.sec.gov/newsroom/press-releases/2024-36

17. Chair Gary Gensler, SEC public statements on AI washing, 18 March 2024 and 4 September 2024. https://www.sec.gov/newsroom/speeches-statements/sec-chair-gary-gensler-ai-washing

18. FTC, 6(b) study order on generative-AI partnerships and investments, 25 January 2024. https://www.ftc.gov/news-events/news/press-releases/2024/01/ftc-launches-inquiry-generative-ai-investments-partnerships

19. FTC, 6(b) surveillance-pricing orders, 23 July 2024. https://www.ftc.gov/news-events/news/press-releases/2024/07/ftc-issues-orders-eight-companies-seeking-information-surveillance-pricing

20. FTC, staff report, A Look Behind the Screens: Examining the Data Practices of Social Media and Video Streaming Services, 19 September 2024. https://www.ftc.gov/reports/look-behind-screens-examining-data-practices-social-media-video-streaming-services

21. UK DSIT, A Pro-Innovation Approach to AI Regulation (white paper), 29 March 2023. https://www.gov.uk/government/publications/ai-regulation-a-pro-innovation-approach

22. UK DSIT, A Pro-Innovation Approach to AI Regulation: government response, 6 February 2024. https://www.gov.uk/government/consultations/ai-regulation-a-pro-innovation-approach-policy-proposals/outcome/a-pro-innovation-approach-to-ai-regulation-government-response

23. UK Government (Government Digital Service), AI Playbook for the UK Government, February 2025. https://www.gov.uk/government/publications/ai-playbook-for-the-uk-government

24. ICO, Data Protection Audit Framework (including the Artificial Intelligence toolkit), October 2024. https://ico.org.uk/about-the-ico/media-centre/news-and-blogs/2024/10/new-data-protection-audit-framework-launched/

25. ICO, AI and data protection risk toolkit / data analytics toolkit. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/

26. OECD, Due Diligence Guidance for Responsible AI, 19 February 2026. https://www.oecd.org/en/publications/2026/02/oecd-due-diligence-guidance-for-responsible-ai_7831bb49.html

27. SDAIA, Principles and Controls of AI Ethics, 14 September 2023. https://sdaia.gov.sa/en/

28. SDAIA, activation of AI offices across 23 government entities and AI Adoption Framework, announced September 2024. https://www.spa.gov.sa/en/N2217019

29. SDAIA, State of AI and national AI adoption materials, 2024. https://sdaia.gov.sa/en/

30. DIFC, Data Protection Regulations, Regulation 10, in force 1 September 2023. https://www.difc.com/business/registrars-and-commissioners/commissioner-of-data-protection/regulation-10

31. DIFC, Consultation Paper No. 3 of 2026, June 2026. https://www.difc.com/whats-on/news/difc-consultation-amended-data-protection-regulations

32. ADGM FSRA, OpenReg (Open Regulation) AI initiative. https://www.adgm.com/media/announcements/adgms-financial-services-regulatory-authority-launches-its-ai-initiative-on-open-regulation

33. ADGM, AI Applications in Web3 SupTech and RegTech: A Regulatory Perspective, 7 February 2025. https://www.adgmacademy.com/publications/AI-Applications-in-Web3-SupTech-and-RegTech-A-Regulatory-Perspective

34. Latham & Watkins, AI in the UAE: Understanding the Regulatory Landscape and Key Authorities, August 2024. https://www.lw.com/en/insights/ai-in-the-uae-understanding-the-regulatory-landscape-and-key-authorities

35. Margot E. Kaminski & Andrew D. Selbst, An American's Guide to the EU AI Act, Berkeley Technology Law Journal, Vol. 40, No. 4, April 2026. https://btlj.org/wp-content/uploads/2026/04/40.4_Kaminski.pdf

36. EU AI Office / European Commission, General-Purpose AI Code of Practice, 10 July 2025, and the Code of Practice on Transparency of AI-Generated Content, 2026. https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai