AI Consulting Toolkit
37 implementation playbooks and 16 governance artefact templates, with 252+ fill-in checklists across assessment, pilot, scale, and governance. Free, vendor-neutral, and structured to be completed rather than read.
Start from the job in front of you
- Enterprise compliance and regulatory alignment
Regulatory alignment is mostly evidence: a written policy, a classified inventory, assessed risks, named owners, monitoring, and an audit trail. These playbooks produce those artefacts rather than describing them.
- Long-term scalability past the pilot
Most AI programmes stall between a working pilot and reliable production. Scaling is an operating-model problem before it is a technical one, so these cover both the platform and the organisation around it.
- Ethical and lawful data usage
Ethical data handling turns on two questions: are we permitted to use this data this way, and does the result treat people fairly. These playbooks work through provenance and consent, then through measured fairness.
- Risk management templates
Risk work needs written criteria applied consistently, not case-by-case judgement. These supply the register, the classification criteria, the autonomy tiers for agents, and the runbook for when something goes wrong.
- Starting a first AI project
Start by establishing where you actually stand and which use case is worth funding, before committing to a build. Note that these playbooks assume a professional audience: if you are new to AI itself, the learning resources are the better first stop.
- Small teams and limited capacity
A small team cannot run a four-phase programme, and should not try. Take the readiness assessment, prioritise ruthlessly to one use case, charter it properly so the go/no-go is decidable, and pick the sourcing path with the lowest ongoing burden.
- Corporate digital transformation
Transformation fails when it funds models instead of workflow change. These start from where value is actually released: roughly a tenth on algorithms, a fifth on technology and data, and the rest on people and process.
Contents
Assess
7 playbooksEvaluate your organisation's AI readiness and identify opportunities.
- Stakeholder Alignment WorkshopTemplates for getting leadership buy-in.Executive6 items
- Enterprise AI Maturity & Capability ScorecardScore organisational AI maturity across six capability themes and set a realistic target state.Executive4 items
- AI Readiness AssessmentScore your org's data maturity, talent, and infrastructure.Manager12 items
- Use Case Prioritization MatrixRank AI opportunities by impact vs. feasibility.Manager8 items
- Competitor AI Landscape AnalysisMap competitor AI capabilities and identify strategic gaps.Manager4 items
- Data Estate & AI Readiness BlueprintInventory, classify, and permission the data estate before AI is switched on across it.Manager4 items
- AI for Enterprise Architecture PlaybookApply AI to the architecture function itself - roadmaps, asset discovery, solution assurance, and scenario planning.Manager4 items
Pilot
8 playbooksRun a contained proof-of-concept to validate and learn.
- Pilot Project Charter TemplateDefine scope, success metrics, and risk controls.Manager10 items
- Vendor Evaluation ScorecardCompare AI vendors across 15 dimensions.Manager15 items
- Build vs Buy vs Partner Decision FrameworkDecide per capability whether to call an API, fine-tune a smaller model, or build in-house - with three-year TCO.Manager4 items
- AI Ethics ChecklistEnsure fairness, transparency, and accountability.Practitioner9 items
- Prompt Engineering & LLM Integration PlaybookDesign, test, and govern LLM prompts for production use cases.Practitioner8 items
- LLM Fine-Tuning & Model Adaptation PlaybookDecide when to fine-tune and run adaptation experiments that beat prompting.Practitioner6 items
- AI Platform & Reference Architecture BlueprintDesign the shared platform layer - environments, model gateway, retrieval, security perimeter, and operations.Practitioner6 items
- AI Infrastructure & Compute Capacity PlaybookPlan the physical and platform layers - facilities, accelerators, networking, scheduling, quota, and tenancy isolation.Practitioner4 items
Scale
14 playbooksExpand from pilot to production with governance and change management.
- AI ROI Measurement FrameworkTrack and report business value of AI investments.Executive7 items
- AI-Led Business Transformation BlueprintDirect effort where value actually comes from: roughly 10% algorithms, 20% technology and data, 70% people and process.Executive4 items
- Change Management PlaybookDrive adoption with training and communication plans.Manager12 items
- Enterprise GenAI Rollout & Workforce Adoption PlaybookDeploy generative AI assistants across the workforce and convert licences into measured usage and value.Manager4 items
- AI-Powered Customer Experience PlaybookApply AI across the customer journey with deliberate containment, escalation, and experience measurement.Manager4 items
- AI Talent, Skills & Operating Model PlaybookSequence AI hires against capabilities to ship, choose an operating model, and build workforce fluency.Manager4 items
- MLOps Maturity RoadmapFrom manual deployments to automated ML pipelines.Practitioner18 items
- Responsible AI & Bias Testing PlaybookDetect, measure, and mitigate bias before and after deployment.Practitioner5 items
- Data Pipeline & Feature Engineering PlaybookBuild reliable, reproducible ML data pipelines from raw data to model-ready features.Practitioner6 items
- LLM Observability & Evaluation in ProductionInstrument, trace, and continuously evaluate LLM applications at scale.Practitioner5 items
- RAG System Optimization PlaybookTune retrieval quality, chunking, and grounding to make RAG applications accurate and reliable at scale.Practitioner5 items
- LLM Inference Cost & Latency Optimization PlaybookCut serving cost and tail latency for production LLM workloads without sacrificing quality.Practitioner6 items
- AI Incident Response PlaybookDetect, triage, and resolve production AI incidents with a repeatable process.Practitioner4 items
- AI Agent Orchestration & Tool-Use PlaybookDesign reliable multi-step agents with tools, guardrails, and fallbacks.Practitioner2 items
Govern
8 playbooksEstablish ongoing oversight, compliance, and continuous improvement.
- AI Governance FrameworkPolicies, accountability structures, and review processes.Executive20 items
- C-Suite AI Operating Model & Role CharterAssign AI accountability across the executive team and run a standing decision agenda with board oversight.Executive4 items
- Agentic AI Governance & Autonomy FrameworkTier AI agents by autonomy and impact, give each an identity and permissions, and scale oversight with autonomy.Executive4 items
- AI Incident Governance & Regulatory ResponseHandle bias incidents, compliance breaches, and regulator or customer notification obligations.Manager8 items
- EU AI Act Compliance ChecklistStep-by-step compliance checklist for the EU Artificial Intelligence Act.Manager4 items
- AI Management System ImplementationStand up a certifiable AI management system using a Plan-Do-Check-Act cycle and the Govern/Map/Measure/Manage functions.Manager4 items
- Data Governance Operating Model & Lifecycle ControlsEstablish data ownership roles, knowledge-area coverage, maturity targets, and stage controls across the data and model lifecycles.Manager4 items
- Model Risk Management PolicyMonitor, audit, and manage deployed AI systems.Practitioner14 items
Governance artefacts
16 templatesThe documents a regulator or auditor asks to see, each with the fields it carries and the provision that prescribes them.
- AI system register
- Risk classification record
- AI impact assessment
- Technical documentation pack
- Fairness test record
- Accuracy and robustness test record
- Human oversight plan
- Disclosure and content-marking record
- Logging and retention specification
- Post-market monitoring plan and incident log
- Data provenance record
- AI security control record
- Third-party AI assessment
- Conformity file and Statement of Applicability
- AI policy and literacy record
- Decommissioning record
What the templates cover
- AI assessment templates
- Readiness assessments, enterprise maturity scorecards, data-estate audits, use-case prioritisation matrices, and capability gap analyses to establish where your organisation actually stands before committing budget.
- AI pilot templates
- Pilot project charters, success-criteria definitions, vendor evaluation scorecards, build-vs-buy decision frameworks, and AI ethics checklists to run a first deployment that produces a defensible go/no-go decision.
- Scaling & operations templates
- MLOps maturity roadmaps, change management plans, ROI measurement frameworks, workforce adoption playbooks, and observability and incident-response runbooks for moving from pilot to production.
- AI governance templates
- AI governance frameworks, EU AI Act compliance checklists, model risk management policies, management-system implementation guides, and agentic-AI autonomy and oversight controls.
Governance artefacts
16 templatesRegulators do not ask whether you have a governance programme. They ask to see artefacts — a register, a classification, a test record, a log specification. These 16 templates turn each obligation into a document someone maintains, with the fields it has to carry and the provision that prescribes them. Free, and structured to be filled in rather than read.
Three of these, no governance platform produces
Reviewing every vendor in Gartner's AI Governance Platforms Magic Quadrant turned up three artefacts none of them was found to generate: Post-market monitoring plan and incident log, Conformity file and Statement of Applicability, Decommissioning record. The first is the central document of any management-system certification, the second is an explicit duty with a deadline attached, and the third closes a lifecycle every platform leaves open.
AI system register
What AI are we actually running? · required by 10 frameworks
AI system register
What AI are we actually running? · required by 10 frameworks
Every other obligation is downstream of this one. An organisation that cannot list its AI systems cannot classify, assess, monitor or document them, and cannot answer the first question any regulator or auditor asks. It is also the artefact every commercial platform builds first, which tells you something.
The template — 10 fields
- System name and internal identifier
- Business owner and technical owner — the role, and the person holding it
- What it does, in one sentence a non-specialist understands
- Built in-house, bought, or embedded in a purchased product
- Your role per regime: provider, deployer, importer, distributor
- Component objects governed in their own right: models, prompts, datasets, agents
- Third-party models and services it depends on
- Lifecycle stage: proposed, in development, live, suspended, retired
- Date entered, date last reviewed, date retired
- Link to the classification, assessment and documentation for this entry
What to retain
The register and its change history, showing when each system was added and by whom. Retain retired entries: a system switched off last year can still be the subject of a complaint, and several regimes set multi-year record-retention periods that outlive the system.
What usually goes wrong
Counting only the models the data team built. Shadow AI, AI embedded in SaaS products, and AI features switched on by a vendor update are the entries most often missing, and increasingly the ones carrying obligations. Registering prompts, datasets and agents as objects in their own right — as the stronger platforms do — catches what a system-level list misses.
Where these fields come from
- Presupposed across regimes; explicit in the NAIC AI model bulletin's AIS Program and in US interagency model risk guidance (SR 26-2), both of which expect a model inventory covering models in use, recently retired, and under development
Risk classification record
Which rules apply to this system, and why? · required by 14 frameworks
Risk classification record
Which rules apply to this system, and why? · required by 14 frameworks
The tier decides the duties. Classification has to be a written, reasoned decision rather than an assumption, because the reasoning is what gets challenged — and because a wrong answer with recorded reasoning fares far better than an undocumented right one.
The template — 8 fields
- System identifier, linked to the register entry
- Regimes considered, and why each does or does not apply
- Tier assigned under each applicable regime
- The specific criterion, annex entry or use-case category relied on
- Reasoning, including why adjacent tiers were rejected
- Whether anything makes you a provider rather than a deployer
- Residual uncertainty, and where legal advice was taken
- Classifier, reviewer, and date
What to retain
The dated classification with its reasoning, and the trail of any reclassification.
What usually goes wrong
Classifying once at launch and never again. A new purpose, a swapped foundation model or a new affected population can move a system between tiers, and nothing prompts you to notice. Putting your own name or trade mark on a third-party system, substantially modifying it, or changing its intended purpose to a high-risk one can each turn a deployer into a provider, with the full provider obligation set attached.
Where these fields come from
- EU AI Act Article 6 and Annex III for high-risk classification; Article 25 for when a deployer becomes a provider
AI impact assessment
Who could this harm, and what have we done about it? · required by 9 frameworks
AI impact assessment
Who could this harm, and what have we done about it? · required by 9 frameworks
Turns a rights-based obligation into an examinable document: who is affected, how, how severely, and what changed as a result. The EU AI Act's fundamental rights impact assessment, a GDPR data protection impact assessment and an algorithmic impact assessment overlap heavily, and the Act now permits a FRIA to incorporate or cross-refer to the relevant parts of a DPIA rather than duplicating it.
The template — 8 fields
- The deployer's processes in which the system will be used, in line with its intended purpose
- The period over which, and frequency with which, the system will be used
- Categories of natural persons and groups likely to be affected
- Specific risks of harm to those categories, drawing on the provider's Article 13 information
- How human oversight measures will be implemented, per the instructions for use
- Measures if those risks materialise, including internal governance and complaint mechanisms
- Where a DPIA exists: which parts are incorporated rather than repeated
- Consultation undertaken, decision, decision-maker, and date
What to retain
The completed assessment, records of consultation, and the decision to proceed or not. Where the FRIA duty applies, the results are notified to the market surveillance authority. Where a mitigation was promised, keep evidence it was implemented.
What usually goes wrong
Assuming it applies to you, or assuming it does not. The FRIA duty is narrower than commonly believed — public bodies, private entities providing public services, and deployers of creditworthiness and life/health insurance pricing systems. Colorado's annual impact assessment, widely planned for, was eliminated when SB 24-205 was repealed and replaced by SB 26-189. Writing it after go-live is the other failure: an assessment that never changed anything invites the question of what it was for.
Where these fields come from
- EU AI Act Article 27(1)(a)–(f) sets the six required elements; scope is public bodies, private entities providing public services, and Annex III 5(b) creditworthiness and 5(c) life and health insurance pricing
- GDPR Article 35 for the DPIA; Regulation (EU) 2026/1744 permits cross-reference between the two
Technical documentation pack
Can we explain how this system was built and what it does? · required by 18 frameworks
Technical documentation pack
Can we explain how this system was built and what it does? · required by 18 frameworks
The reference record handed to a regulator, an auditor or a downstream deployer. Assembled once and maintained, rather than reconstructed under deadline. The EU AI Act enumerates its contents in nine points, and the open model card schema covers much of the same ground in a form engineers will actually complete.
The template — 14 fields
- General description: intended purpose, provider, version and its relation to previous versions
- How it interacts with hardware, software and other AI systems not part of it
- The forms in which it is placed on the market — embedded, download, API — and the hardware it runs on
- Development process: methods and steps, including any pre-trained third-party models and how they were used, integrated or modified
- Design specifications: general logic, key design choices and their rationale, what it optimises for, expected output and output quality, trade-offs accepted
- System architecture, and the computational resources used to develop, train, test and validate
- Training methodologies and datasets: provenance, scope, main characteristics, labelling and cleaning
- Human oversight measures, and the technical measures helping deployers interpret outputs
- Validation and testing: procedures, data, metrics for accuracy and robustness, potentially discriminatory impacts, and dated signed test logs
- Cybersecurity measures
- Capabilities and limitations, including accuracy for specific groups, and foreseeable unintended outcomes
- Pre-determined changes and how continuous compliance is maintained
- Harmonised standards applied, or the solutions adopted where none were
- The declaration of conformity, and the post-market monitoring plan
What to retain
The versioned pack, tied to the release it describes, plus the separate instructions for use given to deployers. Documentation that does not correspond to the version actually running is worse than none.
What usually goes wrong
Treating it as a launch deliverable. The pack must track the deployed version, which means the release process updates it or blocks. Note also that instructions for use are a distinct artefact under Article 13(3) — provider identity, intended purpose, the accuracy and robustness metrics validated against, foreseeable-misuse risks, oversight measures, expected lifetime and maintenance, and how to interpret the logs.
Where these fields come from
- EU AI Act Annex IV, nine points (sub-point lettering omitted deliberately: sources conflict on the letters, not the content)
- EU AI Act Article 13(3) for instructions for use — not Annex IX, which covers registration for real-world testing
- Model Cards for Model Reporting (Mitchell et al., arXiv:1810.03993): Model Details, Intended Use, Factors, Metrics, Evaluation Data, Training Data, Quantitative Analyses, Ethical Considerations, Caveats and Recommendations
Fairness test record
Does this system treat groups differently, and is that justified? · required by 22 frameworks
Fairness test record
Does this system treat groups differently, and is that justified? · required by 22 frameworks
Records the test, not the intention. Bias duties are met by measurement against defined groups and thresholds, repeated over time, with action when a threshold is crossed. New York City's bias audit is the only regime in force anywhere that prescribes the actual calculation, so its structure is worth following even where it does not bind you.
The template — 9 fields
- Groups tested: sex categories, race/ethnicity categories, and the intersectional combinations
- Selection rate per category — those selected to advance, divided by applicants in that category
- Scoring rate per category, where the system scores rather than selects — the rate of scores above the sample median
- Impact ratio per category, against the most-selected or highest-scoring category
- Number of individuals assessed in each category, and the number in an unknown category
- Any category under 2% of the data excluded from the impact ratio, with the justification
- Data used: historical use data, or test data with an explanation of why and how it was generated
- Disparities found, the explanation for each, and the action taken or the reasoned decision to accept
- Auditor independence, tester, date, and next scheduled test
What to retain
Dated results with the threshold stated in advance, and the record of what followed a failure. Where NYC Local Law 144 applies, a summary of results is published on the careers section of the website and kept there for at least six months after latest use, and candidates get at least ten business days' notice.
What usually goes wrong
Treating 0.80 as a pass mark. The four-fifths rule comes from the EEOC Uniform Guidelines; Local Law 144 requires you to calculate and publish the impact ratio, not to clear it. Building a pass/fail gate on that basis misreads the law in both directions. The other failure is choosing the threshold after seeing the results — set it in advance and record it, or the test proves nothing.
Where these fields come from
- NYC Local Law 144, rules at 6 RCNY §§5-300 to 5-304: selection rate, scoring rate, impact ratio, intersectional categories, the 2% exclusion, published summary contents, and the ten-business-day candidate notice
- EU AI Act Article 10(2) on examining and mitigating bias in datasets
Accuracy and robustness test record
Does it work as claimed, and how does it fail? · required by 15 frameworks
Accuracy and robustness test record
Does it work as claimed, and how does it fail? · required by 15 frameworks
Substantiates the performance claims made in the documentation, and establishes behaviour under adversarial conditions and edge cases rather than on the happy path. The Act expects testing against metrics and thresholds defined in advance, which is what separates this from a demo.
The template — 10 fields
- Claimed performance, and the metric behind the claim
- Metrics and probabilistic thresholds defined in advance, appropriate to the intended purpose
- Test set, and how it differs from training data
- Results against the claim, including the conditions where it does not hold
- Performance for specific persons or groups the system is intended to be used on
- Edge cases and out-of-distribution behaviour
- Adversarial testing: prompt injection including indirect injection, jailbreak, model extraction, data poisoning
- Testing under real-world conditions, where undertaken
- Failure modes found, and whether each is mitigated or accepted as residual risk
- Tester, date, version tested; test logs dated and signed
What to retain
Test reports tied to versions, and the record of issues found, triaged and resolved. Red-team findings retained even where no action followed, with the reasoning. Residual risk judged acceptable, per hazard and overall.
What usually goes wrong
Testing the model and not the system. Most real failures come from retrieval, tool use, prompt construction and integration, none of which a model benchmark exercises. Running the test after the threshold is known is the other way to make the record worthless.
Where these fields come from
- EU AI Act Article 9(6)–(7) on testing against prior-defined metrics and probabilistic thresholds, and Article 60 on real-world testing
- EU AI Act Article 15 on accuracy, robustness and cybersecurity
Human oversight plan
Who can stop this, and are they actually able to? · required by 12 frameworks
Human oversight plan
Who can stop this, and are they actually able to? · required by 12 frameworks
Makes oversight real rather than nominal. The Act does not ask whether a human is in the loop; it asks whether that person can understand the system's limitations, resist over-relying on it, interpret its output, decide not to use it, and bring it to a halt safely.
The template — 9 fields
- Oversight role, who holds it, and the competence, training and authority they have
- What that person sees, and whether it is enough to judge an output
- How they are made aware of the system's capacities and limitations, and how anomalies are detected
- Measures addressing automation bias — the tendency to over-rely on the output
- Interpretation tools and methods available to them
- Their ability to decide not to use the system, or to disregard, override or reverse its output
- The stop mechanism, and how the system halts in a safe state
- Expected decision volume and the time available per decision
- Escalation route and out-of-hours cover
What to retain
The plan, training and competence records, and a log of interventions actually made. A system with oversight designed in but never exercised should prompt a question about whether it is real.
What usually goes wrong
Automation bias, which the Act names explicitly. A reviewer approving hundreds of recommendations an hour is not exercising oversight, and the throughput figures in this artefact are what reveal it. Certain remote biometric identification uses additionally require two competent people to verify before any action is taken.
Where these fields come from
- EU AI Act Article 14(4)(a)–(e), including automation bias and the stop function; Article 14(5) on two-person verification for remote biometric identification
- EU AI Act Article 26(2): the deployer assigns oversight to people with the necessary competence, training, authority and support
Disclosure and content-marking record
Do the people affected know they are dealing with AI? · required by 27 frameworks
Disclosure and content-marking record
Do the people affected know they are dealing with AI? · required by 27 frameworks
Covers three duties that get conflated: telling people they are interacting with an AI system, labelling deepfakes and synthetic content visibly, and marking generated output so it is machine-detectable. These duties applied from 2 August 2026 and were not deferred.
The template — 9 fields
- Disclosure shown to people interacting with the system, its exact wording, and where it appears
- Whether it is given at the latest at the time of first interaction or exposure, clearly and distinguishably
- Whether the AI nature is instead obvious to a reasonably well-informed, observant person, and the basis for that view
- Content types generated or manipulated, and the visible or audible label applied to each
- Machine-readable marking method, and whether it survives normal editing
- For deepfakes: the disclosure, and any limitation where the work is artistic, satirical or fictional
- For text published to inform the public on matters of public interest: the disclosure, or the human review and named editorial responsibility relied on instead
- Emotion recognition or biometric categorisation: how exposed people are informed
- Accessibility and languages of the disclosure
What to retain
Screenshots or recordings of the disclosure as users actually see it, the technical specification of the marking, and verification that it is detectable.
What usually goes wrong
Assuming the provider's machine-readable marking discharges your duty. It does not: the Commission's own guidance is that a deployer's deepfake disclosure must be a visible or audible label understandable without a detection tool. Marking applied upstream may also not survive your pipeline, and the duty sits with whoever puts the content into the world.
Where these fields come from
- EU AI Act Article 50: (1) interaction disclosure, (2) machine-readable marking of synthetic content, (3) emotion recognition and biometric categorisation, (4) deepfakes and public-interest text, (5) form and timing
- California AB 2013 for generative training-data disclosure published on the developer's website
Logging and retention specification
Could we reconstruct what happened, months later? · required by 7 frameworks
Logging and retention specification
Could we reconstruct what happened, months later? · required by 7 frameworks
Defines what is logged, for how long, and who can read it — decided deliberately in advance, because the log you did not keep cannot be recreated once an incident or complaint arrives. The Act requires automatic recording of events over the system's lifetime.
The template — 9 fields
- Events logged: inputs, outputs, decisions, overrides, configuration and version changes
- Fields captured per event, and what is deliberately excluded
- Whether the logs identify situations that may present a risk or lead to a substantial modification
- Retention period and the provision setting it — at least six months for high-risk systems, for provider and deployer alike
- Storage, access control, and who may read the logs
- Tamper-evidence measures
- How one decision is reconstructed end to end from the logs
- Deletion process at end of retention
- For remote biometric identification: start and end time of each use, the reference database, the input data that produced a match, and the people who verified the result
What to retain
The specification, plus a periodic test that a specific past decision can in fact be reconstructed. Retention quietly exceeding what is lawful is its own exposure.
What usually goes wrong
Logging everything, including personal data, indefinitely. Over-retention converts a compliance control into a data protection breach waiting to happen — and the Act's own six-month floor is expressly subject to data protection law.
Where these fields come from
- EU AI Act Article 12 on automatic event logging, and Article 12(3) for the remote biometric identification minimum log set
- EU AI Act Article 19 (providers) and Article 26(6) (deployers): at least six months
Post-market monitoring plan and incident log
No platform provides thisIs it still behaving, and what do we do when it is not? · required by 19 frameworks
Post-market monitoring plan and incident log
No platform provides thisIs it still behaving, and what do we do when it is not? · required by 19 frameworks
Deployment is not the end of the obligation. Two distinct things live here: a documented plan for systematically collecting and analysing performance data across the system's lifetime, and a log of incidents with hard reporting deadlines attached. No commercial governance platform was found to produce an incident report as a named artefact, despite it being an explicit duty.
The template — 10 fields
- Metrics monitored, and the threshold that triggers action
- How performance data is actively and systematically collected, documented and analysed across the lifetime
- Data from deployers and other sources, and how continuous compliance is evaluated
- Interaction with other AI systems, where relevant
- Drift detection method and frequency
- Who is alerted, how, and out of hours
- Incident severity definitions, tied to the statutory categories
- Per incident: what happened, when detected, when the causal link was established, who was affected, action taken
- Reporting deadline met, to which authority, and whether an initial report preceded the complete one
- Post-incident investigation, risk assessment, corrective action, and changes made
What to retain
Monitoring output over time, the incident log including incidents judged not reportable with the reasoning, and evidence of reports actually filed. The monitoring plan itself forms part of the technical documentation.
What usually goes wrong
Having no pre-agreed definition of a serious incident. The clock runs from establishing a causal link or its reasonable likelihood, and debating severity while it runs is how deadlines get missed. Note the deadlines differ: 15 days generally, 10 days where a person has died, and 2 days for a widespread infringement or a serious and irreversible disruption of critical infrastructure. A deployer who spots a serious incident tells the provider immediately, then the distributor and the market surveillance authority.
Where these fields come from
- EU AI Act Article 72: documented post-market monitoring system, proportionate, based on a plan that forms part of Annex IV
- EU AI Act Article 3(49) defines a serious incident; Article 73 sets the 15-day, 10-day and 2-day deadlines and permits an initial then complete report
- California SB 53: critical safety incidents reported to Cal OES within 15 days, or 24 hours where there is imminent risk of death or serious physical injury
- OECD AI Incidents Monitor: harm type, severity, scope and geographic scale as the structure of a harm record
Data provenance record
Where did the data come from, and are we allowed to use it? · required by 19 frameworks
Data provenance record
Where did the data come from, and are we allowed to use it? · required by 19 frameworks
Traces training, validation and test data to a source and a lawful basis, and records what is known about quality and representativeness. The hardest artefact to produce retrospectively, and now the one most often demanded — California requires much of it to be published outright.
The template — 15 fields
- Dataset identifier and version, and the design choices behind it
- Collection process and the origin of the data; for personal data, the original purpose of collection
- Sources or owners of each constituent dataset, including scraped and purchased data
- Whether datasets were purchased or licensed, and any restriction attached
- Whether they contain material protected by copyright, trademark or patent, or are public domain
- Whether they contain personal information or aggregate consumer information
- Number and types of data points, and the labels used
- Preparation: annotation, labelling, cleaning, updating, enrichment, aggregation, de-identification
- Assumptions about what the data is supposed to measure and represent
- Availability, quantity, suitability, and statistical properties relative to the people the system will be used on
- Geographical, contextual, behavioural or functional setting the data reflects
- Examination for biases likely to affect health, safety, fundamental rights or to cause unlawful discrimination, and the measures taken
- Time period of collection, whether collection is ongoing, and dates of first use
- Whether synthetic data generation was or is used
- For third-party models: what the provider does and does not disclose
What to retain
The record itself, licences and contracts for acquired data, the published disclosure where one is required, and the trail linking a deployed model version to the dataset versions behind it.
What usually goes wrong
Stopping at the boundary of your own organisation. A fine-tuned open-weight model inherits the provenance questions of its base model, and 'the provider does not disclose' is a finding to record, not a blank to leave. Special category personal data may be processed for bias detection and correction only where strictly necessary and with safeguards — it is an exception, not a licence.
Where these fields come from
- EU AI Act Article 10(2)(a)–(g) on data governance practices, 10(3) on relevance and representativeness, 10(4) on setting, and 10(5) on special category data for bias correction
- California AB 2013: twelve enumerated disclosures per training dataset, published on the developer's website
- Datasheets for Datasets (Gebru et al., arXiv:1803.09010): Motivation, Composition, Collection Process, Preprocessing, Uses, Distribution, Maintenance
AI security control record
What stops someone attacking or stealing this system? · required by 14 frameworks
AI security control record
What stops someone attacking or stealing this system? · required by 14 frameworks
Records the controls specific to AI, which conventional security programmes usually miss, alongside the standard ones, and who is accountable for each.
The template — 9 fields
- Access control over models, weights, prompts, training data and logs
- Protection against data poisoning in training data and in retrieval sources
- Defences against model extraction and inversion
- Prompt injection handling, including indirect injection through retrieved content
- Constraints on the tools and data an agent may reach
- Secrets handling in prompts, logs and traces
- Resilience to errors, faults and inconsistencies, and behaviour on failure
- Dependencies on third-party models and their security posture
- Penetration testing scope, and whether it included the model surface
What to retain
The control record, test results demonstrating the controls work, and penetration test results scoped to include the AI surface.
What usually goes wrong
Scoping a penetration test to the application and excluding the model. The interesting attacks are increasingly through the model, not around it.
Where these fields come from
- EU AI Act Article 15 on accuracy, robustness and cybersecurity, including resilience to attempts to alter use, outputs or performance
- OWASP GenAI Top 10 for LLM applications as the working taxonomy of attack classes
Third-party AI assessment
What are we relying on from our suppliers, and can we prove it? · required by 10 frameworks
Third-party AI assessment
What are we relying on from our suppliers, and can we prove it? · required by 10 frameworks
Most organisations are deployers rather than providers. This records what each supplier commits to, what they disclose, and which obligations remain yours regardless of the contract.
The template — 9 fields
- Supplier, product, and the AI inside it — including features added by vendor update
- Your role and theirs under each applicable regime
- Documentation the supplier provides, and what they refuse to provide
- Instructions for use received, and whether they are sufficient to operate the system correctly
- Contractual commitments on compliance, notification, audit rights and information access
- Their sub-processors and upstream model providers
- Whether anything you do makes you the provider: your branding, substantial modification, or a changed intended purpose
- Obligations that remain yours despite the supplier relationship
- Exit plan if the supplier becomes non-compliant, withdraws the product, or fails
What to retain
Completed assessments, the contract clauses relied on, and supplier documentation retained where given. Record refusals explicitly — an unanswered question is a finding, not a gap.
What usually goes wrong
Assessing at purchase and never again. AI features get switched on by vendor updates, which is how an assessed low-risk tool quietly becomes something else. Regimes borrowed from model risk management expect you to validate a vendor model in your own context rather than accepting the vendor's validation.
Where these fields come from
- EU AI Act Article 25 on when a deployer becomes a provider; Article 26(1) on using a system per its instructions
- US interagency model risk guidance (SR 26-2) and the NAIC AI model bulletin, both of which put third-party models and vendor due diligence expressly in scope
Conformity file and Statement of Applicability
No platform provides thisCan we demonstrate compliance formally, to someone external? · required by 5 frameworks
Conformity file and Statement of Applicability
No platform provides thisCan we demonstrate compliance formally, to someone external? · required by 5 frameworks
Assembles the other artefacts into the file an assessment, certification or registration actually requires. The Statement of Applicability — which control you applied, which you excluded, and why — is the central document of any management-system certification, and notably not one that any commercial AI governance platform was found to generate.
The template — 11 fields
- Assessment route: self-assessment or notified body, and the basis for choosing it
- Standards and specifications applied, and where harmonised standards were not applied in full, the alternative solutions adopted
- For each control in the standard's control set: included or excluded
- Justification for each inclusion, and for each exclusion
- Implementation status of each included control, and the evidence supporting it
- Evidence type per control: a quantitative test result, a document, or a human attestation
- Index of the artefacts relied on and where each lives
- Gaps identified, with an owner and a date for closing each
- Declarations and markings issued, with dates
- Registrations filed, and in which database
- Surveillance schedule and re-assessment triggers
What to retain
The file, the declaration of conformity, certificates where issued, and the audit trail of findings and their resolution.
What usually goes wrong
Building the file only when an audit is scheduled. If the underlying artefacts are maintained, this is an index; if they are not, it is an archaeology project. The quality management system behind it is itself a documented artefact — written policies, procedures and instructions covering the compliance strategy, design and development control, testing, data management, post-market monitoring, incident reporting, record-keeping and an accountability framework.
Where these fields come from
- EU AI Act Article 17 on the quality management system, Article 43 on conformity assessment, Article 47 on the declaration of conformity, Article 49 and Annex VIII on registration
- ISO/IEC 42001 requires a Statement of Applicability; we do not reproduce its clause text or control titles, which are paywalled and which we have not read
AI policy and literacy record
Do our people know the rules, and can they follow them? · required by 8 frameworks
AI policy and literacy record
Do our people know the rules, and can they follow them? · required by 8 frameworks
Covers the written policy, the named roles carrying it, and the AI literacy duty. That duty is an obligation of means rather than result: no certification is mandated, and the required level is proportionate to role, context and the people the system is used on.
The template — 10 fields
- Policy version, approval date, and approving body
- Scope: which systems, which people, which suppliers
- Accountable roles, named, with their decision rights
- Who is covered beyond employees — contractors, service providers, and others operating systems on your behalf
- Training delivered by role, with content, method and date
- How competence is checked, not merely attendance
- How the level was matched to each group's technical knowledge, experience and context of use
- Route for staff to raise a concern about an AI system
- Where workers are subject to an AI system at work: how they and their representatives were informed
- Exceptions granted, by whom, and for how long
What to retain
The approved policy with version history, training and competence records by role, and the log of concerns raised and how each was handled.
What usually goes wrong
One generic course for everyone. The duty is proportionate to role and risk: a developer, a procurement officer and a customer-service reviewer need materially different things. The Digital Omnibus softened the wording from ensuring a sufficient level of AI literacy to supporting its development — which lowers the bar but does not remove the need to show what you did.
Where these fields come from
- EU AI Act Article 4 (as amended by Regulation (EU) 2026/1744) and Article 3(56) defining AI literacy
- EU AI Act Article 26(7): employers inform workers' representatives and affected workers before using a high-risk system at the workplace
Decommissioning record
No platform provides thisHow do we switch this off without losing the ability to answer for it?
Decommissioning record
No platform provides thisHow do we switch this off without losing the ability to answer for it?
Every governance platform reviewed handles intake and lifecycle; not one was found to produce a retirement artefact. That is a real gap, because obligations outlive the system: logs must be retained for a minimum period, records for years under some regimes, and a complaint about a decision made last year still has to be answerable after the system is gone.
The template — 10 fields
- System identifier and the register entry being closed
- Reason for retirement, and the decision-maker
- Date of last production use, and date of shutdown
- What replaced it, if anything, and whether the replacement inherited the same obligations
- Decisions made by the system that remain live or appealable, and who now answers for them
- Logs, models, prompts and datasets retained; where they are held; who can read them
- Retention period for each, and the provision or rationale setting it
- Scheduled deletion date for each retained item, and who executes it
- Notifications made: registration withdrawn, deployers informed, authority notified where required
- Where documentation and evidence for the retired system now live
What to retain
The completed record, the retained artefacts at their stated locations, and eventually the evidence of deletion at end of retention. Keep the register entry, marked retired — deleting it removes your own ability to show what you once ran.
What usually goes wrong
Deleting everything on shutdown, or keeping everything forever. Both are failures: the first destroys your ability to answer a later complaint, the second is an open-ended data protection exposure. Decide per artefact, in writing, before the system goes off.
Where these fields come from
- EU AI Act Articles 19 and 26(6): logs retained at least six months, subject to data protection law
- Colorado SB 26-189: developers and deployers keep records demonstrating compliance for at least three years
These templates are our own drafting, following the source provisions where a regime prescribes contents and established open schemas where one exists. They are not legal advice: whether an obligation binds your organisation, and any conformity assessment, needs qualified counsel. The obligations behind them are set out in the compliance library, where each regime links to its official source.
Phase 1: Assess
Evaluate your organisation's AI readiness and identify opportunities.
- Executive
Stakeholder Alignment Workshop
Templates for getting leadership buy-in.
6 itemsOpen - Executive
Enterprise AI Maturity & Capability Scorecard
Score organisational AI maturity across six capability themes and set a realistic target state.
4 itemsOpen - Manager
AI Readiness Assessment
Score your org's data maturity, talent, and infrastructure.
12 itemsOpen - Manager
Use Case Prioritization Matrix
Rank AI opportunities by impact vs. feasibility.
8 itemsOpen - Manager
Competitor AI Landscape Analysis
Map competitor AI capabilities and identify strategic gaps.
4 itemsOpen - Manager
Data Estate & AI Readiness Blueprint
Inventory, classify, and permission the data estate before AI is switched on across it.
4 itemsOpen - Manager
AI for Enterprise Architecture Playbook
Apply AI to the architecture function itself - roadmaps, asset discovery, solution assurance, and scenario planning.
4 itemsOpen
Phase 2: Pilot
Run a contained proof-of-concept to validate and learn.
- Manager
Pilot Project Charter Template
Define scope, success metrics, and risk controls.
10 itemsOpen - Manager
Vendor Evaluation Scorecard
Compare AI vendors across 15 dimensions.
15 itemsOpen - Manager
Build vs Buy vs Partner Decision Framework
Decide per capability whether to call an API, fine-tune a smaller model, or build in-house - with three-year TCO.
4 itemsOpen - Practitioner
AI Ethics Checklist
Ensure fairness, transparency, and accountability.
9 itemsOpen - Practitioner
Prompt Engineering & LLM Integration Playbook
Design, test, and govern LLM prompts for production use cases.
8 itemsOpen - Practitioner
LLM Fine-Tuning & Model Adaptation Playbook
Decide when to fine-tune and run adaptation experiments that beat prompting.
6 itemsOpen - Practitioner
AI Platform & Reference Architecture Blueprint
Design the shared platform layer - environments, model gateway, retrieval, security perimeter, and operations.
6 itemsOpen - Practitioner
AI Infrastructure & Compute Capacity Playbook
Plan the physical and platform layers - facilities, accelerators, networking, scheduling, quota, and tenancy isolation.
4 itemsOpen
Phase 3: Scale
Expand from pilot to production with governance and change management.
- Executive
AI ROI Measurement Framework
Track and report business value of AI investments.
7 itemsOpen - Executive
AI-Led Business Transformation Blueprint
Direct effort where value actually comes from: roughly 10% algorithms, 20% technology and data, 70% people and process.
4 itemsOpen - Manager
Change Management Playbook
Drive adoption with training and communication plans.
12 itemsOpen - Manager
Enterprise GenAI Rollout & Workforce Adoption Playbook
Deploy generative AI assistants across the workforce and convert licences into measured usage and value.
4 itemsOpen - Manager
AI-Powered Customer Experience Playbook
Apply AI across the customer journey with deliberate containment, escalation, and experience measurement.
4 itemsOpen - Manager
AI Talent, Skills & Operating Model Playbook
Sequence AI hires against capabilities to ship, choose an operating model, and build workforce fluency.
4 itemsOpen - Practitioner
MLOps Maturity Roadmap
From manual deployments to automated ML pipelines.
18 itemsOpen - Practitioner
Responsible AI & Bias Testing Playbook
Detect, measure, and mitigate bias before and after deployment.
5 itemsOpen - Practitioner
Data Pipeline & Feature Engineering Playbook
Build reliable, reproducible ML data pipelines from raw data to model-ready features.
6 itemsOpen - Practitioner
LLM Observability & Evaluation in Production
Instrument, trace, and continuously evaluate LLM applications at scale.
5 itemsOpen - Practitioner
RAG System Optimization Playbook
Tune retrieval quality, chunking, and grounding to make RAG applications accurate and reliable at scale.
5 itemsOpen - Practitioner
LLM Inference Cost & Latency Optimization Playbook
Cut serving cost and tail latency for production LLM workloads without sacrificing quality.
6 itemsOpen - Practitioner
AI Incident Response Playbook
Detect, triage, and resolve production AI incidents with a repeatable process.
4 itemsOpen - Practitioner
AI Agent Orchestration & Tool-Use Playbook
Design reliable multi-step agents with tools, guardrails, and fallbacks.
2 itemsOpen
Phase 4: Govern
Establish ongoing oversight, compliance, and continuous improvement.
- Executive
AI Governance Framework
Policies, accountability structures, and review processes.
20 itemsOpen - Executive
C-Suite AI Operating Model & Role Charter
Assign AI accountability across the executive team and run a standing decision agenda with board oversight.
4 itemsOpen - Executive
Agentic AI Governance & Autonomy Framework
Tier AI agents by autonomy and impact, give each an identity and permissions, and scale oversight with autonomy.
4 itemsOpen - Manager
AI Incident Governance & Regulatory Response
Handle bias incidents, compliance breaches, and regulator or customer notification obligations.
8 itemsOpen - Manager
EU AI Act Compliance Checklist
Step-by-step compliance checklist for the EU Artificial Intelligence Act.
4 itemsOpen - Manager
AI Management System Implementation
Stand up a certifiable AI management system using a Plan-Do-Check-Act cycle and the Govern/Map/Measure/Manage functions.
4 itemsOpen - Manager
Data Governance Operating Model & Lifecycle Controls
Establish data ownership roles, knowledge-area coverage, maturity targets, and stage controls across the data and model lifecycles.
4 itemsOpen - Practitioner
Model Risk Management Policy
Monitor, audit, and manage deployed AI systems.
14 itemsOpen