Playbooks & Templates

AI Consulting Toolkit

37 implementation playbooks and 16 governance artefact templates, with 252+ fill-in checklists across assessment, pilot, scale, and governance. Free, vendor-neutral, and structured to be completed rather than read.

Start from the job in front of you

Enterprise compliance and regulatory alignment

Regulatory alignment is mostly evidence: a written policy, a classified inventory, assessed risks, named owners, monitoring, and an audit trail. These playbooks produce those artefacts rather than describing them.

Long-term scalability past the pilot

Most AI programmes stall between a working pilot and reliable production. Scaling is an operating-model problem before it is a technical one, so these cover both the platform and the organisation around it.

Ethical and lawful data usage

Ethical data handling turns on two questions: are we permitted to use this data this way, and does the result treat people fairly. These playbooks work through provenance and consent, then through measured fairness.

Risk management templates

Risk work needs written criteria applied consistently, not case-by-case judgement. These supply the register, the classification criteria, the autonomy tiers for agents, and the runbook for when something goes wrong.

Starting a first AI project

Start by establishing where you actually stand and which use case is worth funding, before committing to a build. Note that these playbooks assume a professional audience: if you are new to AI itself, the learning resources are the better first stop.

Small teams and limited capacity

A small team cannot run a four-phase programme, and should not try. Take the readiness assessment, prioritise ruthlessly to one use case, charter it properly so the go/no-go is decidable, and pick the sourcing path with the lowest ongoing burden.

Corporate digital transformation

Transformation fails when it funds models instead of workflow change. These start from where value is actually released: roughly a tenth on algorithms, a fifth on technology and data, and the rest on people and process.

Scale

14 playbooks

Expand from pilot to production with governance and change management.

What the templates cover

AI assessment templates
Readiness assessments, enterprise maturity scorecards, data-estate audits, use-case prioritisation matrices, and capability gap analyses to establish where your organisation actually stands before committing budget.
AI pilot templates
Pilot project charters, success-criteria definitions, vendor evaluation scorecards, build-vs-buy decision frameworks, and AI ethics checklists to run a first deployment that produces a defensible go/no-go decision.
Scaling & operations templates
MLOps maturity roadmaps, change management plans, ROI measurement frameworks, workforce adoption playbooks, and observability and incident-response runbooks for moving from pilot to production.
AI governance templates
AI governance frameworks, EU AI Act compliance checklists, model risk management policies, management-system implementation guides, and agentic-AI autonomy and oversight controls.

Governance artefacts

16 templates

Regulators do not ask whether you have a governance programme. They ask to see artefacts — a register, a classification, a test record, a log specification. These 16 templates turn each obligation into a document someone maintains, with the fields it has to carry and the provision that prescribes them. Free, and structured to be filled in rather than read.

Three of these, no governance platform produces

Reviewing every vendor in Gartner's AI Governance Platforms Magic Quadrant turned up three artefacts none of them was found to generate: Post-market monitoring plan and incident log, Conformity file and Statement of Applicability, Decommissioning record. The first is the central document of any management-system certification, the second is an explicit duty with a deadline attached, and the third closes a lifecycle every platform leaves open.

AI system register

What AI are we actually running? · required by 10 frameworks

Every other obligation is downstream of this one. An organisation that cannot list its AI systems cannot classify, assess, monitor or document them, and cannot answer the first question any regulator or auditor asks. It is also the artefact every commercial platform builds first, which tells you something.

Typically owned by
AI governance lead, with a named owner per entry
Cadence
Continuous intake; full reconciliation quarterly

The template — 10 fields

  • System name and internal identifier
  • Business owner and technical owner — the role, and the person holding it
  • What it does, in one sentence a non-specialist understands
  • Built in-house, bought, or embedded in a purchased product
  • Your role per regime: provider, deployer, importer, distributor
  • Component objects governed in their own right: models, prompts, datasets, agents
  • Third-party models and services it depends on
  • Lifecycle stage: proposed, in development, live, suspended, retired
  • Date entered, date last reviewed, date retired
  • Link to the classification, assessment and documentation for this entry

What to retain

The register and its change history, showing when each system was added and by whom. Retain retired entries: a system switched off last year can still be the subject of a complaint, and several regimes set multi-year record-retention periods that outlive the system.

What usually goes wrong

Counting only the models the data team built. Shadow AI, AI embedded in SaaS products, and AI features switched on by a vendor update are the entries most often missing, and increasingly the ones carrying obligations. Registering prompts, datasets and agents as objects in their own right — as the stronger platforms do — catches what a system-level list misses.

Where these fields come from

  • Presupposed across regimes; explicit in the NAIC AI model bulletin's AIS Program and in US interagency model risk guidance (SR 26-2), both of which expect a model inventory covering models in use, recently retired, and under development

Risk classification record

Which rules apply to this system, and why? · required by 14 frameworks

The tier decides the duties. Classification has to be a written, reasoned decision rather than an assumption, because the reasoning is what gets challenged — and because a wrong answer with recorded reasoning fares far better than an undocumented right one.

Typically owned by
Risk or compliance function, countersigned by the business owner
Cadence
At intake, and again on any substantial modification

The template — 8 fields

  • System identifier, linked to the register entry
  • Regimes considered, and why each does or does not apply
  • Tier assigned under each applicable regime
  • The specific criterion, annex entry or use-case category relied on
  • Reasoning, including why adjacent tiers were rejected
  • Whether anything makes you a provider rather than a deployer
  • Residual uncertainty, and where legal advice was taken
  • Classifier, reviewer, and date

What to retain

The dated classification with its reasoning, and the trail of any reclassification.

What usually goes wrong

Classifying once at launch and never again. A new purpose, a swapped foundation model or a new affected population can move a system between tiers, and nothing prompts you to notice. Putting your own name or trade mark on a third-party system, substantially modifying it, or changing its intended purpose to a high-risk one can each turn a deployer into a provider, with the full provider obligation set attached.

Where these fields come from

  • EU AI Act Article 6 and Annex III for high-risk classification; Article 25 for when a deployer becomes a provider

AI impact assessment

Who could this harm, and what have we done about it? · required by 9 frameworks

Turns a rights-based obligation into an examinable document: who is affected, how, how severely, and what changed as a result. The EU AI Act's fundamental rights impact assessment, a GDPR data protection impact assessment and an algorithmic impact assessment overlap heavily, and the Act now permits a FRIA to incorporate or cross-refer to the relevant parts of a DPIA rather than duplicating it.

Typically owned by
Deployer's business owner, facilitated by privacy or compliance
Cadence
Before first use; updated whenever an element is no longer current

The template — 8 fields

  • The deployer's processes in which the system will be used, in line with its intended purpose
  • The period over which, and frequency with which, the system will be used
  • Categories of natural persons and groups likely to be affected
  • Specific risks of harm to those categories, drawing on the provider's Article 13 information
  • How human oversight measures will be implemented, per the instructions for use
  • Measures if those risks materialise, including internal governance and complaint mechanisms
  • Where a DPIA exists: which parts are incorporated rather than repeated
  • Consultation undertaken, decision, decision-maker, and date

What to retain

The completed assessment, records of consultation, and the decision to proceed or not. Where the FRIA duty applies, the results are notified to the market surveillance authority. Where a mitigation was promised, keep evidence it was implemented.

What usually goes wrong

Assuming it applies to you, or assuming it does not. The FRIA duty is narrower than commonly believed — public bodies, private entities providing public services, and deployers of creditworthiness and life/health insurance pricing systems. Colorado's annual impact assessment, widely planned for, was eliminated when SB 24-205 was repealed and replaced by SB 26-189. Writing it after go-live is the other failure: an assessment that never changed anything invites the question of what it was for.

Where these fields come from

  • EU AI Act Article 27(1)(a)–(f) sets the six required elements; scope is public bodies, private entities providing public services, and Annex III 5(b) creditworthiness and 5(c) life and health insurance pricing
  • GDPR Article 35 for the DPIA; Regulation (EU) 2026/1744 permits cross-reference between the two

Technical documentation pack

Can we explain how this system was built and what it does? · required by 18 frameworks

The reference record handed to a regulator, an auditor or a downstream deployer. Assembled once and maintained, rather than reconstructed under deadline. The EU AI Act enumerates its contents in nine points, and the open model card schema covers much of the same ground in a form engineers will actually complete.

Typically owned by
Technical owner
Cadence
At release; updated with every version

The template — 14 fields

  • General description: intended purpose, provider, version and its relation to previous versions
  • How it interacts with hardware, software and other AI systems not part of it
  • The forms in which it is placed on the market — embedded, download, API — and the hardware it runs on
  • Development process: methods and steps, including any pre-trained third-party models and how they were used, integrated or modified
  • Design specifications: general logic, key design choices and their rationale, what it optimises for, expected output and output quality, trade-offs accepted
  • System architecture, and the computational resources used to develop, train, test and validate
  • Training methodologies and datasets: provenance, scope, main characteristics, labelling and cleaning
  • Human oversight measures, and the technical measures helping deployers interpret outputs
  • Validation and testing: procedures, data, metrics for accuracy and robustness, potentially discriminatory impacts, and dated signed test logs
  • Cybersecurity measures
  • Capabilities and limitations, including accuracy for specific groups, and foreseeable unintended outcomes
  • Pre-determined changes and how continuous compliance is maintained
  • Harmonised standards applied, or the solutions adopted where none were
  • The declaration of conformity, and the post-market monitoring plan

What to retain

The versioned pack, tied to the release it describes, plus the separate instructions for use given to deployers. Documentation that does not correspond to the version actually running is worse than none.

What usually goes wrong

Treating it as a launch deliverable. The pack must track the deployed version, which means the release process updates it or blocks. Note also that instructions for use are a distinct artefact under Article 13(3) — provider identity, intended purpose, the accuracy and robustness metrics validated against, foreseeable-misuse risks, oversight measures, expected lifetime and maintenance, and how to interpret the logs.

Where these fields come from

  • EU AI Act Annex IV, nine points (sub-point lettering omitted deliberately: sources conflict on the letters, not the content)
  • EU AI Act Article 13(3) for instructions for use — not Annex IX, which covers registration for real-world testing
  • Model Cards for Model Reporting (Mitchell et al., arXiv:1810.03993): Model Details, Intended Use, Factors, Metrics, Evaluation Data, Training Data, Quantitative Analyses, Ethical Considerations, Caveats and Recommendations

Fairness test record

Does this system treat groups differently, and is that justified? · required by 22 frameworks

Records the test, not the intention. Bias duties are met by measurement against defined groups and thresholds, repeated over time, with action when a threshold is crossed. New York City's bias audit is the only regime in force anywhere that prescribes the actual calculation, so its structure is worth following even where it does not bind you.

Typically owned by
Technical owner, reviewed by compliance; an independent auditor where the law requires one
Cadence
Pre-deployment, then at least annually — the NYC audit must be no more than one year old at the time of use

The template — 9 fields

  • Groups tested: sex categories, race/ethnicity categories, and the intersectional combinations
  • Selection rate per category — those selected to advance, divided by applicants in that category
  • Scoring rate per category, where the system scores rather than selects — the rate of scores above the sample median
  • Impact ratio per category, against the most-selected or highest-scoring category
  • Number of individuals assessed in each category, and the number in an unknown category
  • Any category under 2% of the data excluded from the impact ratio, with the justification
  • Data used: historical use data, or test data with an explanation of why and how it was generated
  • Disparities found, the explanation for each, and the action taken or the reasoned decision to accept
  • Auditor independence, tester, date, and next scheduled test

What to retain

Dated results with the threshold stated in advance, and the record of what followed a failure. Where NYC Local Law 144 applies, a summary of results is published on the careers section of the website and kept there for at least six months after latest use, and candidates get at least ten business days' notice.

What usually goes wrong

Treating 0.80 as a pass mark. The four-fifths rule comes from the EEOC Uniform Guidelines; Local Law 144 requires you to calculate and publish the impact ratio, not to clear it. Building a pass/fail gate on that basis misreads the law in both directions. The other failure is choosing the threshold after seeing the results — set it in advance and record it, or the test proves nothing.

Where these fields come from

  • NYC Local Law 144, rules at 6 RCNY §§5-300 to 5-304: selection rate, scoring rate, impact ratio, intersectional categories, the 2% exclusion, published summary contents, and the ten-business-day candidate notice
  • EU AI Act Article 10(2) on examining and mitigating bias in datasets

Accuracy and robustness test record

Does it work as claimed, and how does it fail? · required by 15 frameworks

Substantiates the performance claims made in the documentation, and establishes behaviour under adversarial conditions and edge cases rather than on the happy path. The Act expects testing against metrics and thresholds defined in advance, which is what separates this from a demo.

Typically owned by
Technical owner
Cadence
Pre-deployment and per release; adversarial testing at least annually for generative and agentic systems

The template — 10 fields

  • Claimed performance, and the metric behind the claim
  • Metrics and probabilistic thresholds defined in advance, appropriate to the intended purpose
  • Test set, and how it differs from training data
  • Results against the claim, including the conditions where it does not hold
  • Performance for specific persons or groups the system is intended to be used on
  • Edge cases and out-of-distribution behaviour
  • Adversarial testing: prompt injection including indirect injection, jailbreak, model extraction, data poisoning
  • Testing under real-world conditions, where undertaken
  • Failure modes found, and whether each is mitigated or accepted as residual risk
  • Tester, date, version tested; test logs dated and signed

What to retain

Test reports tied to versions, and the record of issues found, triaged and resolved. Red-team findings retained even where no action followed, with the reasoning. Residual risk judged acceptable, per hazard and overall.

What usually goes wrong

Testing the model and not the system. Most real failures come from retrieval, tool use, prompt construction and integration, none of which a model benchmark exercises. Running the test after the threshold is known is the other way to make the record worthless.

Where these fields come from

  • EU AI Act Article 9(6)–(7) on testing against prior-defined metrics and probabilistic thresholds, and Article 60 on real-world testing
  • EU AI Act Article 15 on accuracy, robustness and cybersecurity

Human oversight plan

Who can stop this, and are they actually able to? · required by 12 frameworks

Makes oversight real rather than nominal. The Act does not ask whether a human is in the loop; it asks whether that person can understand the system's limitations, resist over-relying on it, interpret its output, decide not to use it, and bring it to a halt safely.

Typically owned by
Deployer's business owner
Cadence
Defined before deployment; tested annually

The template — 9 fields

  • Oversight role, who holds it, and the competence, training and authority they have
  • What that person sees, and whether it is enough to judge an output
  • How they are made aware of the system's capacities and limitations, and how anomalies are detected
  • Measures addressing automation bias — the tendency to over-rely on the output
  • Interpretation tools and methods available to them
  • Their ability to decide not to use the system, or to disregard, override or reverse its output
  • The stop mechanism, and how the system halts in a safe state
  • Expected decision volume and the time available per decision
  • Escalation route and out-of-hours cover

What to retain

The plan, training and competence records, and a log of interventions actually made. A system with oversight designed in but never exercised should prompt a question about whether it is real.

What usually goes wrong

Automation bias, which the Act names explicitly. A reviewer approving hundreds of recommendations an hour is not exercising oversight, and the throughput figures in this artefact are what reveal it. Certain remote biometric identification uses additionally require two competent people to verify before any action is taken.

Where these fields come from

  • EU AI Act Article 14(4)(a)–(e), including automation bias and the stop function; Article 14(5) on two-person verification for remote biometric identification
  • EU AI Act Article 26(2): the deployer assigns oversight to people with the necessary competence, training, authority and support

Disclosure and content-marking record

Do the people affected know they are dealing with AI? · required by 27 frameworks

Covers three duties that get conflated: telling people they are interacting with an AI system, labelling deepfakes and synthetic content visibly, and marking generated output so it is machine-detectable. These duties applied from 2 August 2026 and were not deferred.

Typically owned by
Product owner, with communications
Cadence
At launch; reviewed on any interface or model change

The template — 9 fields

  • Disclosure shown to people interacting with the system, its exact wording, and where it appears
  • Whether it is given at the latest at the time of first interaction or exposure, clearly and distinguishably
  • Whether the AI nature is instead obvious to a reasonably well-informed, observant person, and the basis for that view
  • Content types generated or manipulated, and the visible or audible label applied to each
  • Machine-readable marking method, and whether it survives normal editing
  • For deepfakes: the disclosure, and any limitation where the work is artistic, satirical or fictional
  • For text published to inform the public on matters of public interest: the disclosure, or the human review and named editorial responsibility relied on instead
  • Emotion recognition or biometric categorisation: how exposed people are informed
  • Accessibility and languages of the disclosure

What to retain

Screenshots or recordings of the disclosure as users actually see it, the technical specification of the marking, and verification that it is detectable.

What usually goes wrong

Assuming the provider's machine-readable marking discharges your duty. It does not: the Commission's own guidance is that a deployer's deepfake disclosure must be a visible or audible label understandable without a detection tool. Marking applied upstream may also not survive your pipeline, and the duty sits with whoever puts the content into the world.

Where these fields come from

  • EU AI Act Article 50: (1) interaction disclosure, (2) machine-readable marking of synthetic content, (3) emotion recognition and biometric categorisation, (4) deepfakes and public-interest text, (5) form and timing
  • California AB 2013 for generative training-data disclosure published on the developer's website

Logging and retention specification

Could we reconstruct what happened, months later? · required by 7 frameworks

Defines what is logged, for how long, and who can read it — decided deliberately in advance, because the log you did not keep cannot be recreated once an incident or complaint arrives. The Act requires automatic recording of events over the system's lifetime.

Typically owned by
Technical owner, with security
Cadence
Defined at design; reviewed annually

The template — 9 fields

  • Events logged: inputs, outputs, decisions, overrides, configuration and version changes
  • Fields captured per event, and what is deliberately excluded
  • Whether the logs identify situations that may present a risk or lead to a substantial modification
  • Retention period and the provision setting it — at least six months for high-risk systems, for provider and deployer alike
  • Storage, access control, and who may read the logs
  • Tamper-evidence measures
  • How one decision is reconstructed end to end from the logs
  • Deletion process at end of retention
  • For remote biometric identification: start and end time of each use, the reference database, the input data that produced a match, and the people who verified the result

What to retain

The specification, plus a periodic test that a specific past decision can in fact be reconstructed. Retention quietly exceeding what is lawful is its own exposure.

What usually goes wrong

Logging everything, including personal data, indefinitely. Over-retention converts a compliance control into a data protection breach waiting to happen — and the Act's own six-month floor is expressly subject to data protection law.

Where these fields come from

  • EU AI Act Article 12 on automatic event logging, and Article 12(3) for the remote biometric identification minimum log set
  • EU AI Act Article 19 (providers) and Article 26(6) (deployers): at least six months

Post-market monitoring plan and incident log

No platform provides this

Is it still behaving, and what do we do when it is not? · required by 19 frameworks

Deployment is not the end of the obligation. Two distinct things live here: a documented plan for systematically collecting and analysing performance data across the system's lifetime, and a log of incidents with hard reporting deadlines attached. No commercial governance platform was found to produce an incident report as a named artefact, despite it being an explicit duty.

Typically owned by
Technical owner for monitoring; compliance for reporting
Cadence
Continuous monitoring; incident review per event; plan reviewed annually

The template — 10 fields

  • Metrics monitored, and the threshold that triggers action
  • How performance data is actively and systematically collected, documented and analysed across the lifetime
  • Data from deployers and other sources, and how continuous compliance is evaluated
  • Interaction with other AI systems, where relevant
  • Drift detection method and frequency
  • Who is alerted, how, and out of hours
  • Incident severity definitions, tied to the statutory categories
  • Per incident: what happened, when detected, when the causal link was established, who was affected, action taken
  • Reporting deadline met, to which authority, and whether an initial report preceded the complete one
  • Post-incident investigation, risk assessment, corrective action, and changes made

What to retain

Monitoring output over time, the incident log including incidents judged not reportable with the reasoning, and evidence of reports actually filed. The monitoring plan itself forms part of the technical documentation.

What usually goes wrong

Having no pre-agreed definition of a serious incident. The clock runs from establishing a causal link or its reasonable likelihood, and debating severity while it runs is how deadlines get missed. Note the deadlines differ: 15 days generally, 10 days where a person has died, and 2 days for a widespread infringement or a serious and irreversible disruption of critical infrastructure. A deployer who spots a serious incident tells the provider immediately, then the distributor and the market surveillance authority.

Where these fields come from

  • EU AI Act Article 72: documented post-market monitoring system, proportionate, based on a plan that forms part of Annex IV
  • EU AI Act Article 3(49) defines a serious incident; Article 73 sets the 15-day, 10-day and 2-day deadlines and permits an initial then complete report
  • California SB 53: critical safety incidents reported to Cal OES within 15 days, or 24 hours where there is imminent risk of death or serious physical injury
  • OECD AI Incidents Monitor: harm type, severity, scope and geographic scale as the structure of a harm record

Data provenance record

Where did the data come from, and are we allowed to use it? · required by 19 frameworks

Traces training, validation and test data to a source and a lawful basis, and records what is known about quality and representativeness. The hardest artefact to produce retrospectively, and now the one most often demanded — California requires much of it to be published outright.

Typically owned by
Data owner, with legal
Cadence
At dataset creation; revisited whenever data is added or the system is substantially modified

The template — 15 fields

  • Dataset identifier and version, and the design choices behind it
  • Collection process and the origin of the data; for personal data, the original purpose of collection
  • Sources or owners of each constituent dataset, including scraped and purchased data
  • Whether datasets were purchased or licensed, and any restriction attached
  • Whether they contain material protected by copyright, trademark or patent, or are public domain
  • Whether they contain personal information or aggregate consumer information
  • Number and types of data points, and the labels used
  • Preparation: annotation, labelling, cleaning, updating, enrichment, aggregation, de-identification
  • Assumptions about what the data is supposed to measure and represent
  • Availability, quantity, suitability, and statistical properties relative to the people the system will be used on
  • Geographical, contextual, behavioural or functional setting the data reflects
  • Examination for biases likely to affect health, safety, fundamental rights or to cause unlawful discrimination, and the measures taken
  • Time period of collection, whether collection is ongoing, and dates of first use
  • Whether synthetic data generation was or is used
  • For third-party models: what the provider does and does not disclose

What to retain

The record itself, licences and contracts for acquired data, the published disclosure where one is required, and the trail linking a deployed model version to the dataset versions behind it.

What usually goes wrong

Stopping at the boundary of your own organisation. A fine-tuned open-weight model inherits the provenance questions of its base model, and 'the provider does not disclose' is a finding to record, not a blank to leave. Special category personal data may be processed for bias detection and correction only where strictly necessary and with safeguards — it is an exception, not a licence.

Where these fields come from

  • EU AI Act Article 10(2)(a)–(g) on data governance practices, 10(3) on relevance and representativeness, 10(4) on setting, and 10(5) on special category data for bias correction
  • California AB 2013: twelve enumerated disclosures per training dataset, published on the developer's website
  • Datasheets for Datasets (Gebru et al., arXiv:1803.09010): Motivation, Composition, Collection Process, Preprocessing, Uses, Distribution, Maintenance

AI security control record

What stops someone attacking or stealing this system? · required by 14 frameworks

Records the controls specific to AI, which conventional security programmes usually miss, alongside the standard ones, and who is accountable for each.

Typically owned by
Security function, with the technical owner
Cadence
At deployment; reviewed annually and after any incident

The template — 9 fields

  • Access control over models, weights, prompts, training data and logs
  • Protection against data poisoning in training data and in retrieval sources
  • Defences against model extraction and inversion
  • Prompt injection handling, including indirect injection through retrieved content
  • Constraints on the tools and data an agent may reach
  • Secrets handling in prompts, logs and traces
  • Resilience to errors, faults and inconsistencies, and behaviour on failure
  • Dependencies on third-party models and their security posture
  • Penetration testing scope, and whether it included the model surface

What to retain

The control record, test results demonstrating the controls work, and penetration test results scoped to include the AI surface.

What usually goes wrong

Scoping a penetration test to the application and excluding the model. The interesting attacks are increasingly through the model, not around it.

Where these fields come from

  • EU AI Act Article 15 on accuracy, robustness and cybersecurity, including resilience to attempts to alter use, outputs or performance
  • OWASP GenAI Top 10 for LLM applications as the working taxonomy of attack classes

Third-party AI assessment

What are we relying on from our suppliers, and can we prove it? · required by 10 frameworks

Most organisations are deployers rather than providers. This records what each supplier commits to, what they disclose, and which obligations remain yours regardless of the contract.

Typically owned by
Procurement, with compliance
Cadence
Before contract; on renewal; on any material change by the supplier

The template — 9 fields

  • Supplier, product, and the AI inside it — including features added by vendor update
  • Your role and theirs under each applicable regime
  • Documentation the supplier provides, and what they refuse to provide
  • Instructions for use received, and whether they are sufficient to operate the system correctly
  • Contractual commitments on compliance, notification, audit rights and information access
  • Their sub-processors and upstream model providers
  • Whether anything you do makes you the provider: your branding, substantial modification, or a changed intended purpose
  • Obligations that remain yours despite the supplier relationship
  • Exit plan if the supplier becomes non-compliant, withdraws the product, or fails

What to retain

Completed assessments, the contract clauses relied on, and supplier documentation retained where given. Record refusals explicitly — an unanswered question is a finding, not a gap.

What usually goes wrong

Assessing at purchase and never again. AI features get switched on by vendor updates, which is how an assessed low-risk tool quietly becomes something else. Regimes borrowed from model risk management expect you to validate a vendor model in your own context rather than accepting the vendor's validation.

Where these fields come from

  • EU AI Act Article 25 on when a deployer becomes a provider; Article 26(1) on using a system per its instructions
  • US interagency model risk guidance (SR 26-2) and the NAIC AI model bulletin, both of which put third-party models and vendor due diligence expressly in scope

Conformity file and Statement of Applicability

No platform provides this

Can we demonstrate compliance formally, to someone external? · required by 5 frameworks

Assembles the other artefacts into the file an assessment, certification or registration actually requires. The Statement of Applicability — which control you applied, which you excluded, and why — is the central document of any management-system certification, and notably not one that any commercial AI governance platform was found to generate.

Typically owned by
Compliance lead
Cadence
Before placing on the market; on substantial modification; at each surveillance audit

The template — 11 fields

  • Assessment route: self-assessment or notified body, and the basis for choosing it
  • Standards and specifications applied, and where harmonised standards were not applied in full, the alternative solutions adopted
  • For each control in the standard's control set: included or excluded
  • Justification for each inclusion, and for each exclusion
  • Implementation status of each included control, and the evidence supporting it
  • Evidence type per control: a quantitative test result, a document, or a human attestation
  • Index of the artefacts relied on and where each lives
  • Gaps identified, with an owner and a date for closing each
  • Declarations and markings issued, with dates
  • Registrations filed, and in which database
  • Surveillance schedule and re-assessment triggers

What to retain

The file, the declaration of conformity, certificates where issued, and the audit trail of findings and their resolution.

What usually goes wrong

Building the file only when an audit is scheduled. If the underlying artefacts are maintained, this is an index; if they are not, it is an archaeology project. The quality management system behind it is itself a documented artefact — written policies, procedures and instructions covering the compliance strategy, design and development control, testing, data management, post-market monitoring, incident reporting, record-keeping and an accountability framework.

Where these fields come from

  • EU AI Act Article 17 on the quality management system, Article 43 on conformity assessment, Article 47 on the declaration of conformity, Article 49 and Annex VIII on registration
  • ISO/IEC 42001 requires a Statement of Applicability; we do not reproduce its clause text or control titles, which are paywalled and which we have not read

AI policy and literacy record

Do our people know the rules, and can they follow them? · required by 8 frameworks

Covers the written policy, the named roles carrying it, and the AI literacy duty. That duty is an obligation of means rather than result: no certification is mandated, and the required level is proportionate to role, context and the people the system is used on.

Typically owned by
AI governance lead, with HR
Cadence
Policy reviewed annually; training on joining and refreshed annually

The template — 10 fields

  • Policy version, approval date, and approving body
  • Scope: which systems, which people, which suppliers
  • Accountable roles, named, with their decision rights
  • Who is covered beyond employees — contractors, service providers, and others operating systems on your behalf
  • Training delivered by role, with content, method and date
  • How competence is checked, not merely attendance
  • How the level was matched to each group's technical knowledge, experience and context of use
  • Route for staff to raise a concern about an AI system
  • Where workers are subject to an AI system at work: how they and their representatives were informed
  • Exceptions granted, by whom, and for how long

What to retain

The approved policy with version history, training and competence records by role, and the log of concerns raised and how each was handled.

What usually goes wrong

One generic course for everyone. The duty is proportionate to role and risk: a developer, a procurement officer and a customer-service reviewer need materially different things. The Digital Omnibus softened the wording from ensuring a sufficient level of AI literacy to supporting its development — which lowers the bar but does not remove the need to show what you did.

Where these fields come from

  • EU AI Act Article 4 (as amended by Regulation (EU) 2026/1744) and Article 3(56) defining AI literacy
  • EU AI Act Article 26(7): employers inform workers' representatives and affected workers before using a high-risk system at the workplace

Decommissioning record

No platform provides this

How do we switch this off without losing the ability to answer for it?

Every governance platform reviewed handles intake and lifecycle; not one was found to produce a retirement artefact. That is a real gap, because obligations outlive the system: logs must be retained for a minimum period, records for years under some regimes, and a complaint about a decision made last year still has to be answerable after the system is gone.

Typically owned by
Business owner, with the technical owner and compliance
Cadence
At retirement, and reviewed when retention periods expire

The template — 10 fields

  • System identifier and the register entry being closed
  • Reason for retirement, and the decision-maker
  • Date of last production use, and date of shutdown
  • What replaced it, if anything, and whether the replacement inherited the same obligations
  • Decisions made by the system that remain live or appealable, and who now answers for them
  • Logs, models, prompts and datasets retained; where they are held; who can read them
  • Retention period for each, and the provision or rationale setting it
  • Scheduled deletion date for each retained item, and who executes it
  • Notifications made: registration withdrawn, deployers informed, authority notified where required
  • Where documentation and evidence for the retired system now live

What to retain

The completed record, the retained artefacts at their stated locations, and eventually the evidence of deletion at end of retention. Keep the register entry, marked retired — deleting it removes your own ability to show what you once ran.

What usually goes wrong

Deleting everything on shutdown, or keeping everything forever. Both are failures: the first destroys your ability to answer a later complaint, the second is an open-ended data protection exposure. Decide per artefact, in writing, before the system goes off.

Where these fields come from

  • EU AI Act Articles 19 and 26(6): logs retained at least six months, subject to data protection law
  • Colorado SB 26-189: developers and deployers keep records demonstrating compliance for at least three years

These templates are our own drafting, following the source provisions where a regime prescribes contents and established open schemas where one exists. They are not legal advice: whether an obligation binds your organisation, and any conformity assessment, needs qualified counsel. The obligations behind them are set out in the compliance library, where each regime links to its official source.

Phase 1: Assess

7 playbooks

Evaluate your organisation's AI readiness and identify opportunities.

Phase 2: Pilot

8 playbooks

Run a contained proof-of-concept to validate and learn.

Phase 3: Scale

14 playbooks

Expand from pilot to production with governance and change management.

Phase 4: Govern

8 playbooks

Establish ongoing oversight, compliance, and continuous improvement.