If you run a regulated or risk-sensitive business, don’t start with the model. Start with the controls. The goal is to leave with an approved, tested, and auditable AI implementation plan, not just an AI policy sitting in a folder. If you’re moving AI into a live workflow, use the seven steps below before anything goes live.
The hard part is not “can AI do the task?” It’s “can you still explain, review, stop, and defend what happens next?” That’s where AI governance changes AI implementation in practice.
Pro tip: if nobody can say who can stop the system, it is not ready to launch.
Step 1: Start with the business problem, not the model
Write a one-page problem statement before you look at tools. Keep it plain. Name the process, the users, the current manual steps, the data touched, the systems connected, the expected productivity benefit, and the customer or employee impact. Then say why AI is the right fit instead of a deterministic rule or ordinary workflow automation.
Also define who the system is for and who it affects. An invoice-extraction assistant, an internal drafting assistant, and a recruitment-ranking tool are not the same governance case. They should not be treated like they are.
You also need a baseline. Capture transaction volume, staff hours, error or rework rate, cycle time, missed follow-ups, and financial impact before you claim any return. Do not guess at a universal ROI percentage. There isn’t one in the dossier, and your own numbers matter more anyway.
Done when: you have a signed problem statement, a use-case brief, a baseline, and a clear statement of what the system is not allowed to do.
Common mistake: picking a fashionable model first and then searching for a task later. Governance cannot assess an undefined system.
Step 2: Build an AI inventory and risk gate before you approve the use case
Create a live inventory for every approved and discovered AI use, including shadow use by staff. This is one of the first places AI governance has to become concrete. Record, at minimum, the business owner, technical owner, provider and model, model version and access mode, purpose, users and affected groups, inputs and outputs, personal or confidential data, connected systems, location and retention, human decision points, known limitations, legal or regulatory requirements, approval status, test status, monitoring owner, incident route, and planned end of life.
For generative AI, also record data provenance, foundation model and version, retrieval or fine-tuning method, known issues, and human-oversight responsibilities.
Then apply a proportionate risk gate:
-
Low impact: internal assistance with no sensitive data and no consequential decision
-
Elevated impact: customer-facing content, personal data, confidential information, automated routing, financial records, or material operational decisions
-
High impact: employment, access to essential services, credit or insurance-like decisions, health or safety, fundamental rights, regulated advice, or decisions about individuals
-
Prohibited or unacceptable: do not proceed where the use is unlawful or cannot be made safe enough
The key point is context. Risk depends on the business setting, the data, the people affected, and the cost of error. A vendor label like “enterprise-grade” does not tell you that.
Done when: the inventory and risk register cover every use case, the risk tier is explained, and a named senior owner and risk owner have signed the decision.
Common mistake: treating the vendor’s label as the risk assessment.
Step 3: Put approvals, accountability, and ethics ownership in writing
Small companies do not need a giant committee. They do need named ownership. A monthly governance meeting can be enough if it is documented and real. Assign, by name or role:
-
senior accountable owner with authority to approve, pause, and fund controls
-
business or process owner
-
technical and security owner
-
privacy or data-protection lead
-
compliance or legal adviser where needed
-
human reviewers and escalation contacts
-
vendor manager
-
independent reviewer or internal-audit contact for elevated and high-impact systems
Write the policy too. It should cover approved tools, prohibited inputs, acceptable uses, review obligations, records, incident reporting, customer or worker notices, procurement, and sanctions for bypassing controls. Keep a decision log with the proposal, evidence reviewed, conditions of approval, residual risks, approver, date, review date, and trigger for reapproval.
Here’s where people stall: they assign the IT buyer and assume that covers everything. It doesn’t. A vendor can supply a component. It does not own the business decision.
Done when: your org chart or policy names owners, the approval record shows who accepted residual risk, and staff know how to report an issue and reach a human reviewer.
Common mistake: making one person the default owner for customer, privacy, compliance, and operational risk.
Step 4: Map data, privacy, security, and supplier dependencies before you connect systems
Before you wire AI into core systems, map the data flow from source to model or provider, then to output, then to the downstream system. Classify personal, sensitive, confidential, financial, privileged, and proprietary data. Use minimisation. Only send the fields needed for the task, and do not use live customer data for experimentation when representative test data will do.
If personal data is involved, keep the relevant record of processing activities. If a Data Protection Impact Assessment is required, do it before processing begins. For high-risk processing, sign it off before you start, act on the mitigations, and revisit it when the processing changes.
Then assess the supplier as part of implementation, not after purchase. Ask whether inputs or outputs are used for training, where data and backups are processed, how long data is retained, who can access prompts and logs, what security assurance and incident notification are available, whether you can get logs and test evidence, what happens when the model changes or the service fails, and whether you can export data and switch provider.
Put those answers into contract and SLA controls: confidentiality, processing instructions, sub-processors, security, incident cooperation, audit access, service levels, version and change notification, deletion, exit, and continuity. Security should run through secure design, secure development, secure deployment, and secure operation.
Done when: a reviewer can trace every sensitive field, access path, provider, sub-processor, retention rule, contract control, and exit plan.
Common mistake: assuming an existing CRM or accounting integration is safe just because it’s familiar.
Step 5: Define human oversight, transparency, and user recourse as workflow controls
Don’t write “human in the loop” and call it done. Define the workflow. Say what the AI may recommend or execute, which outputs require review, what the reviewer sees, what authority they have, how fast they must review, and when they can edit, reject, override, or stop the system.
Also define escalation for uncertainty, bias, hallucination, or safety concerns. Record the decision and the intervention. Give customers, workers, or users a way to challenge an outcome.
For high-impact decisions, human oversight has to be meaningful. That means competence, information, time, and authority. A reviewer who can’t disagree is not a reviewer. A reviewer who is too rushed to understand the output is not a control.
Tell people when AI is creating content, interacting with them, or contributing to a decision that affects them. Keep the notice clear and tied to the context. Transparency is not about exposing model weights. It’s about giving people enough information to know AI is involved, what role it plays, what the limits are, and how to challenge an outcome.
Done when: a reviewer can demonstrate a real reject, override, or stop path in a test, notices are ready, and the audit record captures the AI output, human action, final outcome, and reason.
Common mistake: calling someone a reviewer when they don’t have time or authority to disagree.
Step 6: Test against known-good outcomes, harmful outcomes, and change
Build the test pack before production. You need accuracy and completeness checks, false positives and false negatives, known-good responses and rejection cases, bias and performance across relevant groups, privacy leakage and sensitive-data extraction checks, prompt injection and malicious file testing where relevant, hallucination and unsupported advice checks, resilience and failure handling, integration errors, duplicate writes and partial failures, human-review time and override rates, and provider or model version changes.
Set acceptance criteria for the use case. There is no universal threshold in the dossier, and that’s the right way to think about it. The bar for an internal drafting tool is not the bar for a customer-facing or regulated workflow.
Repeat testing after prompt, retrieval, model, data, workflow, integration, or security changes. Keep the test scope, data description, results, defects, mitigations, approver, and release decision. Treat assurance as an ongoing process, not a one-time certificate.
For invoice work, for example, route low-confidence or mismatched invoices to a person rather than posting them silently. That is the kind of control that keeps productivity and auditability in the same workflow.
Done when: the test pack shows scope, cases, thresholds, results, unresolved risks, sign-off, and release conditions.
Common mistake: testing only a happy path from the vendor demo and assuming that proves production safety.
Step 7: Operate, monitor, audit, and safely change or stop the system
Launch is not the end. It’s the start of monitoring. Build a live control loop that tracks quality, drift, bias, hallucination, complaints, overrides, security events, latency, availability, cost, and downstream errors. Define alert thresholds and owners. Keep incident, complaint, and remediation records. Log access, changes, releases, patches, reconfiguration, and version updates.
Review risks when the context or model changes. Review contracts after material changes. Run periodic risk-based internal audits and share the findings with senior management. Test rollback, fallback, and business continuity. If you cannot stop or reverse the system in practice, you do not have control.
Set stop conditions in advance. These can include unacceptable harm, repeated material errors, privacy or security breach, unexplainable change, failed monitoring, provider incident, or loss of required human control. If you need to decommission the system, do it deliberately. Preserve required records, revoke access, manage data leakage and retention, address dependencies, communicate the change, and keep critical work running through a fallback process.
This is the other place AI implementation goes wrong. Teams treat approval as permanent. AI behavior, providers, threats, users, and expectations change.
Done when: the system has a monitoring log or dashboard, an incident register, a change history, a review date, a tested rollback or fallback, a deactivation authority, and an audit pack that can reconstruct what happened.
Common mistake: treating launch approval as permanent approval.
What to do next
Start with one bounded use case, not five. Create the inventory, appoint the accountable owner, classify the risk, block sensitive inputs until they’ve been reviewed, and run a documented pilot with known-good tests, human approval, and a rollback path. Then decide whether the broader deployment is worth it.
If you want help turning that into a practical operating model, that’s where Governance for AI fits. The service is built to review current systems, build and implement governance structures, and train the team on the new policies and workflows. It starts with an AI Readiness & Systems Audit, turns that into a custom Innovation Plan, and then moves into a bespoke governance framework and workforce training.
If you want help turning that into a practical operating model, that’s where Governance for AI fits. The service is built to review current systems, build and implement governance structures, and train the team on the new policies and workflows. It starts with an AI Readiness & Systems Audit, turns that into a custom Innovation Plan, and then moves into a bespoke governance framework and workforce training.
FAQ
When should governance start?
Before you select the model or sign the vendor contract. Start at problem definition and use-case approval. Retrofitting ownership, data restrictions, human review, and audit logs after launch is slower and can make the workflow harder to use.
Does every AI tool need a governance board or DPIA?
Every use should be inventoried and given proportionate controls. A board can be a lightweight accountable forum for an SME. A DPIA depends on the processing and risk, especially for high-risk personal-data processing, and it should be revisited when the processing changes.
Can a regulated business use generative AI?
Yes, but the implementation has to define permitted data, provider terms, review, transparency, testing, monitoring, and stop conditions. Generative AI should not make consequential decisions alone just because a person can see the output.
Is a vendor compliance certificate enough?
No. Vendor assurance helps with procurement, but it does not transfer accountability. You still need to understand the data flow, contract terms, limitations, human oversight, monitoring, changes, incidents, and exit.
What should be logged?
Log enough to reconstruct the event, including system, model, version, input or reference to it where lawful, output, user, time, downstream action, human review or override, incidents, and changes. Retention depends on the use and applicable law.
