AI risk assessment is the step most organisations skip on the way from “this tool looks useful” to “forty people now depend on it”. The pilot works, the licence gets bought, and nobody writes down what could go wrong, who owns it, or what would make you switch the thing off. Six months later somebody in legal asks a reasonable question and there is no document to point at.
This guide gives you the document. It sets out what an AI risk assessment is, when you actually need one, a template you can fill in section by section, the seven risk categories worth scoring, a scoring scale that produces comparable numbers, the six steps of the exercise itself, the controls that genuinely move a score, and how to map the whole thing to NIST, ISO/IEC 42001 and the EU AI Act without hiring a consultancy.
None of it requires a governance function. A competent operations manager with a spreadsheet and two hours can complete a first-pass AI risk assessment for a single use case, and that first pass is worth more than a forty-page framework nobody opens. If you are building the wider programme, this pairs with our AI strategy work and sits alongside your AI acceptable use policy, which tells staff what they may do; the assessment tells you what the organisation is exposed to when they do it.
Table of contents
- What an AI Risk Assessment Actually Is
- When You Need an AI Risk Assessment — and When You Do Not
- The AI Risk Assessment Template, Section by Section
- The Seven Risk Categories Every AI Risk Assessment Must Score
- How to Score an AI Risk Assessment: Likelihood, Impact and Exposure
- The Six Steps of a Working AI Risk Assessment
- Controls That Actually Move an AI Risk Assessment Score
- Mapping Your AI Risk Assessment to NIST, ISO 42001 and the EU AI Act
- Who Signs Off an AI Risk Assessment, and How Often to Re-Run It
- AI Risk Assessment Mistakes That Waste the Whole Exercise
- Frequently Asked Questions About AI Risk Assessment
- References
What an AI Risk Assessment Actually Is
The phrase gets applied to four different documents, which is why so many of them end up duplicating each other. Being precise about the scope is what keeps the exercise short.
It is a decision record, not a compliance artefact
An AI risk assessment records a decision: we are deploying this capability, here is what could go wrong, here is what we have done about it, here is the exposure we accepted, and here is who accepted it. Everything else in the document exists to support those five facts.
Written that way, it stays useful. Written as a compliance exercise, it becomes a form nobody reads and nobody updates, and it protects nothing when the incident arrives.
The unit of assessment is the use case, not the tool
This is the single most common structural error. “We assessed Copilot” is not an assessment, because the risk of drafting internal meeting notes and the risk of drafting customer-facing contract summaries are not remotely the same, even in the same product.
Assess the use case: this team, doing this task, on this data, with this level of human review. One tool may need four assessments, and three of them will take fifteen minutes each.
It sits between the policy and the deployment
Your acceptable use policy sets the rules for everybody. Your AI risk assessment applies those rules to one specific deployment and works out whether the residual exposure is acceptable. The deployment then inherits the controls the assessment specified.
That ordering matters, because an assessment written after go-live can only document risk. Written before, it can change the design — which is the only point at which controls are cheap.
| Document | Question it answers | Scope | Owner |
|---|---|---|---|
| AI risk assessment | Is this use case safe enough to run? | One use case | Use case owner |
| Acceptable use policy | What may staff do with AI? | Whole organisation | HR and IT |
| DPIA | Is the personal data processing lawful and proportionate? | One processing activity | Data protection lead |
| Vendor due diligence | Is this supplier safe to buy from? | One supplier | Procurement |
| Risk register | What are we tracking across the business? | Everything | Risk or finance lead |
An AI risk assessment feeds the register rather than replacing it; if you keep a cybersecurity risk register, the residual scores from each assessment are what belong in it.
When You Need an AI Risk Assessment — and When You Do Not
A rule that demands a full assessment for every use of every tool gets ignored within a month. A rule that demands nothing produces the shadow deployments you find out about during an audit. The workable answer is a screening question with two outcomes.
The triggers that always require one
Run a full AI risk assessment when the use case touches any of the following: personal or special-category data; customer-facing output that is not reviewed by a person; an automated decision that affects someone’s money, employment, access or entitlement; regulated advice; safety-relevant instructions; or source code and credentials.
Any one of those is enough. They are the categories where a wrong output becomes somebody else’s problem rather than an internal annoyance.
The cases where a five-minute screening is enough
Internal drafting, summarising documents the author already has access to, brainstorming, rewriting text for tone, generating test data, or explaining an error message — these are low-exposure by construction, and a short screening record with a date and an owner is proportionate.
Keep the screening in the same register as your full AI risk assessment records. A “screened, low, no further action” line takes one minute and closes the question permanently.
Shadow AI is the case you will miss
The assessments you never write are the ones for tools nobody told you about: a browser extension, a free tier signed up with a work email, an AI feature switched on inside software you already own. That last category is the fastest-growing and the least visible.
Handle it with discovery rather than policy. Ask each team what they are using in a way that gets an honest answer, and assess what comes back instead of punishing the disclosure.
Re-assessment triggers
An assessment is valid until something changes. Write the triggers into the document: a new data source, a change in who uses it, removal of human review, a model version change, a supplier acquisition, an incident, or twelve months elapsed.
That last one matters least and gets followed most. The model version change is the one that quietly invalidates your original AI risk assessment, and it is the one nobody notices.
The AI Risk Assessment Template, Section by Section
Here is the structure that works. Six sections, roughly two pages, and it is deliberately short — the value is in the scoring and the sign-off, not the prose.
Section 1: Use case and owner
One paragraph in plain English: who uses it, for what task, how often, and what happens to the output. Name a single accountable owner, not a department. Record the tool, the model or service behind it, and whether it is a pilot or production.
If you cannot describe the use case in one paragraph, it is more than one use case — split it, and run a short AI risk assessment against each half rather than one vague document covering both.
Section 2: Data in, data out
List what goes in: categories of data, whether it includes personal data, whether it includes anything commercially sensitive, and where it comes from. List what comes out, who sees it, and whether it leaves the organisation.
Then record the three facts that decide most of your risk: is the data used to train the provider’s models, where is it processed, and how long is it retained. These belong in the AI risk assessment because they are the ones that change when a supplier updates its terms.
Section 3: Inherent risk scoring
Score each of the seven categories in the next section on likelihood and impact before any controls are applied. Inherent scoring feels artificial — you already have controls — but it is what tells you which risks the controls are actually carrying.
Section 4: Controls and owners
For each risk scored medium or above, record the control, the person responsible, and whether it is in place today or planned with a date. A planned control does not reduce a score until it exists.
Section 5: Residual risk and decision
Re-score with controls applied. Then write the decision explicitly: approved, approved with conditions, or declined. Name who approved it and when. This section is the reason the document exists.
Section 6: Review date, evidence and kill switch
Set the review date, list the re-assessment triggers, and record where the evidence lives — logs, evaluation results, the supplier’s security documentation. Finally, write down how you would turn the use case off and who can do it, because an AI risk assessment without a stop condition is a description, not a control.
| Section | Fields to capture | Typical time |
|---|---|---|
| 1. Use case and owner | Description, owner, tool, model, pilot or production | 10 min |
| 2. Data in, data out | Input categories, output audience, training use, location, retention | 25 min |
| 3. Inherent risk | Seven categories scored 1–5 for likelihood and impact | 30 min |
| 4. Controls and owners | Control, owner, in place or planned, target date | 30 min |
| 5. Residual risk and decision | Re-scored grid, decision, approver, date | 20 min |
| 6. Review and kill switch | Review date, triggers, evidence location, stop procedure | 15 min |
The Seven Risk Categories Every AI Risk Assessment Must Score
Seven is enough to be complete and few enough to finish. Score every category even when the answer is obviously low — the low scores are what make the high ones credible.
Accuracy and hallucination risk
The system produces something confident and wrong. Score it on how detectable the error is and how far it travels before a human sees it. A fabricated citation in an internal summary is an irritation; the same fabrication in advice sent to a client is a different category of event.
Data protection and confidentiality risk
Personal or confidential data enters a system you do not control, is retained longer than you expected, or is used to improve a shared model. This is where most UK organisations have a genuine legal exposure, and where the ICO’s AI guidance is the reference point rather than an opinion.
Security risk: prompt injection and model abuse
If the system reads untrusted content — emails, web pages, uploaded documents — that content can carry instructions. Prompt injection is not theoretical, and it is the risk that scales badly the moment the system is allowed to take actions rather than just produce text.
Bias and fairness risk
Outputs differ systematically across groups in a way that matters. This scores high wherever the output influences a decision about a person: shortlisting, pricing, eligibility, prioritisation of support.
Legal, intellectual property and contractual risk
Who owns the output, whether the input was yours to submit, whether the supplier’s terms permit your use, and whether any downstream customer contract prohibits AI processing. Most organisations discover the last one after the fact.
Operational and dependency risk
The capability becomes load-bearing. Availability, cost escalation, model deprecation and switching difficulty all belong here, and they connect directly to AI vendor lock-in — a dependency you cannot exit is a risk you cannot mitigate.
Reputational and disclosure risk
What happens if a customer, journalist or regulator learns exactly how this output is produced. If the honest answer is uncomfortable, the score is high regardless of the technical controls, and the remedy is usually disclosure rather than engineering.
How to Score an AI Risk Assessment: Likelihood, Impact and Exposure
Scoring is where most templates fall apart, because they offer a 1–5 scale without defining what the numbers mean. Two people then score the same risk three points apart and the totals become noise.
Use a 1–5 scale and define every number
Write one sentence per number, in your own language, before anybody scores anything. Likelihood 3 might be “we would expect this a few times a year across normal usage”. Impact 4 might be “a customer is materially affected and we have to tell them”.
The definitions matter more than the scale. Any consistent scale beats an undefined one, and a defined 1–5 makes two assessments comparable — which is the entire point of doing more than one.
Write impact in your own currency
Impact is not a colour. Express it as money, hours, customers affected or regulatory consequence, then map that to the number. A 4 that means “£25,000 to £100,000 or a reportable breach” survives a challenge; a 4 that means “quite bad” does not.
Exposure is what makes AI scoring different
Classic risk scoring is likelihood times impact. AI needs a third factor, because the same flawed output behaves very differently at different volumes and different levels of autonomy. Ten outputs a week reviewed by a specialist is not the same risk as ten thousand a week acting automatically.
Score exposure as volume multiplied by autonomy: how many outputs, and how far they travel without a person in the loop. It is the factor that most often turns an apparently modest AI risk assessment score into a genuine priority.
Set escalation thresholds before you score
Decide in advance which residual score requires which approval. Doing it afterwards invites the number to be negotiated down to fit the approval you wanted.
| Residual band | Score (L×I×E) | Decision right | Review cadence |
|---|---|---|---|
| Low | 1–15 | Use case owner | Annual |
| Moderate | 16–35 | Department head | Every 6 months |
| High | 36–74 | IT and data protection lead jointly | Quarterly |
| Severe | 75–125 | Board or equivalent, or do not proceed | Monthly until reduced |
The Six Steps of a Working AI Risk Assessment
The template is the artefact; these six steps are the exercise. Run an AI risk assessment in this order and a straightforward use case takes about two hours from blank page to signature.
Step 1: Describe the use case in one paragraph
Write it with the person who actually does the work, not with the person who bought the licence. Ask what they do today, what the tool changes, and what they would do if it were unavailable tomorrow. The gap between the sponsor’s description and the practitioner’s is usually where the risk lives.
Step 2: Map the data flow
Draw it if you can: source, what is sent, where it is processed, what comes back, who sees it, what is stored and for how long. Most of the data protection questions answer themselves once the flow is on paper, and the ones that do not are exactly the ones to raise with the supplier.
Step 3: Score the inherent risk
Work through the seven categories with controls switched off. Do it as a pair, not alone — the disagreements are informative, and a score two people argued about is more defensible than one person’s number. This is the half of the AI risk assessment that a reviewer will actually interrogate.
Step 4: Choose controls that change the score
For each medium-or-above risk, pick a control and say which score it moves and by how much. If you cannot articulate that, it is a good practice rather than a control, and it belongs somewhere else.
Step 5: Record residual risk and decide
Re-score, compare against your thresholds, and write the decision down with a name and a date. “Approved with conditions” is the most useful outcome available and the most underused — it lets you proceed while making the outstanding work visible.
Step 6: Set the review date and the kill switch
Fix the review date, list the triggers, and describe how the use case gets switched off in a hurry: who has the access, how long it takes, and what breaks. Rehearsing that once is worth more than another page of the AI risk assessment itself.
Controls That Actually Move an AI Risk Assessment Score
A control counts if you can name the score it changes. Everything else is hygiene. These five carry most of the reduction in practice.
Human review at the decision point
Not review in principle — review at the specific moment the output would otherwise become an action. Placed correctly, it is the single largest score reduction available in an AI risk assessment, converting a high accuracy score into a moderate one. Placed as a general instruction to “check the output”, it changes nothing, because nobody checks what already looks right.
Retrieval grounding and citation
Make the system answer from documents you control and show which one it used. This attacks hallucination at the source and makes errors detectable in seconds rather than never, which is why it moves both likelihood and impact.
Input filtering and data minimisation
Send less. Strip identifiers that the task does not need, block categories of data at the gateway, and separate the untrusted content the system reads from the instructions it follows. Minimisation is the cheapest control in any AI risk assessment and consistently the most underused.
Logging, evaluation sets and drift monitoring
Keep a small set of representative inputs with known-good outputs and re-run it after every model or prompt change. Twenty examples is enough to catch a regression that a subjective “seems fine” review will miss entirely, and it is the only control that detects silent model updates.
Contractual and supplier controls
Data processing terms, a commitment not to train on your data, notice of model changes, and a documented export path. These reduce legal and dependency scores and nothing else, so do not let them stand in for technical controls. Good vendor management makes them enforceable rather than aspirational, and the NCSC’s secure AI development guidelines are a reasonable baseline to hold suppliers to.
Mapping Your AI Risk Assessment to NIST, ISO 42001 and the EU AI Act
You do not need to adopt a framework to run a useful assessment. You do need your document to line up with one, so that the day somebody asks for a framework you are not starting again.
NIST AI RMF: Govern, Map, Measure, Manage
The NIST AI Risk Management Framework splits the work into four functions, and the six-step exercise above maps cleanly onto them: steps 1 and 2 are Map, step 3 is Measure, steps 4 and 5 are Manage, and step 6 plus your approval thresholds are Govern. It is voluntary, free, and the most practical starting point for an organisation with no existing programme.
ISO/IEC 42001 and the AI management system
ISO/IEC 42001 defines a management system rather than an assessment method — it wants evidence that assessments happen consistently, get reviewed, and feed improvement. Your completed assessments become the evidence. If certification is on the horizon, our guide to ISO 42001 implementation covers what the auditors look for.
The EU AI Act risk tiers
The Act classifies systems as prohibited, high-risk, limited-risk or minimal-risk, with obligations attached to each tier. If you sell into the EU, add one field to section 1 of the template recording your tier and your reasoning. Most internal productivity use cases land in minimal or limited risk, but the classification needs to be written down rather than assumed.
The UK position
The UK has taken a regulator-led approach rather than a single AI statute, which in practice means the ICO for personal data, the NCSC for security, and sector regulators for everything else. For UK organisations that makes the ICO’s guidance and a defensible DPIA the operative requirement, with the AI risk assessment sitting above it as the broader document.
| Template section | NIST AI RMF | ISO/IEC 42001 | EU AI Act |
|---|---|---|---|
| 1. Use case and owner | Map | Context and roles | Tier classification |
| 2. Data in, data out | Map | Data management | Data governance duties |
| 3. Inherent risk | Measure | Risk assessment | Risk management system |
| 4. Controls and owners | Manage | Risk treatment | Technical documentation |
| 5. Residual risk and decision | Manage and Govern | Management review | Conformity evidence |
| 6. Review and kill switch | Govern | Monitoring and improvement | Post-market monitoring |
Who Signs Off an AI Risk Assessment, and How Often to Re-Run It
An unsigned assessment is a draft, and a draft protects nobody. The governance around the document is short but non-negotiable.
The three roles you cannot skip
The use case owner, who is accountable for the risk and the controls. A data protection voice, wherever personal data is involved. And a technical voice who understands what the system can actually do — usually whoever would have to switch it off.
Three people, one meeting. Larger organisations add a committee; smaller ones should resist the temptation, because a committee converts a two-hour exercise into a two-month one.
Approval by score, not by seniority
Use the thresholds you set earlier. Low-band assessments are approved by the owner and filed. Severe-band assessments go up, and “do not proceed” has to be a genuinely available answer or the whole exercise is theatre.
Annual review versus event-driven review
Diary the annual review, but expect the event triggers to do the real work. A model version change, a new data source or the removal of human review should each force a re-score within days, not at the next anniversary. Treat it the way you treat change management for anything else in production.
Keep them somewhere findable
One folder, one index, one owner. An assessment nobody can find is worth precisely as much as one that was never written, and this is where sensible IT governance earns its keep.
AI Risk Assessment Mistakes That Waste the Whole Exercise
These six account for most of the wasted effort we see. All are avoidable in the first hour.
Assessing the tool instead of the use case
The most expensive error, because it produces a document that is simultaneously too vague to be useful and too broad to be wrong. Split by use case even when it means four short AI risk assessment records instead of one long one.
Stopping at inherent risk
A grid of red squares with no control column tells a reader that everything is dangerous and nothing is being done. Residual scoring is the half that carries the meaning.
Copying a template without defining the scale
The template is the easy part. If you take this one, spend the twenty minutes writing your own definitions for each number on the likelihood and impact scales before anybody uses it, or your scores will not be comparable to each other.
Making the assessment a gate nobody can pass
Set the bar so high that no realistic use case clears it and staff route around the process entirely. Shadow deployments are the direct consequence, and they are strictly worse than an imperfect approved one.
Never reviewing it again
An AI risk assessment from eighteen months ago describes a model that no longer exists. Unreviewed assessments do not merely go stale; they actively mislead, because they carry the authority of a signed document.
Leaving the evidence in an inbox
Screenshots, supplier emails and evaluation results scattered across personal mailboxes cannot support the document when it matters. Reference their location in section 6 and keep them with the assessment.
Frequently Asked Questions About AI Risk Assessment
How long should an AI risk assessment take?
About two hours for a first-pass on a straightforward internal use case, and half a day for something customer-facing or automated. Subsequent assessments on the same tool are much faster because sections 1 and 2 are largely reusable. If yours routinely takes longer than a day, the template is too heavy.
Do we need one if we only use ChatGPT or Copilot informally?
Yes, but a short one. Informal use is still use, and the informality is itself a risk factor because nobody knows what is being pasted in. A single screening record covering “general internal drafting” plus a clear acceptable use policy is proportionate for most organisations.
Is an AI risk assessment the same as a DPIA?
No. A DPIA is a specific UK GDPR obligation focused on personal data processing, and it may be legally required. An AI risk assessment is broader — accuracy, security, dependency and reputation are all in scope regardless of whether personal data is involved. Where both apply, run the assessment first and let it feed the DPIA.
Who should own the process in a small business?
Whoever owns IT decisions, with input from whoever owns customer relationships. An AI risk assessment does not need a dedicated risk function to be credible. What it does need is one named person who keeps the index and chases the review dates, because the process fails through drift rather than through bad judgement.
What if the supplier will not answer our questions?
Treat non-answers as findings and score them. A supplier who will not confirm whether your data trains their models has answered the question. If the use case is important enough, escalate it commercially; if it is not, pick a supplier who will answer.
How does this connect to our wider supplier checks?
Directly. The supplier-facing half of an AI risk assessment overlaps heavily with standard third-party review, so reuse the work — our supplier cyber-risk assessment checklist covers the security questions, and the AI-specific additions are training use, model change notice and output ownership.
Can we automate any of it?
The evidence collection, yes — logs, evaluation runs and inventory discovery all automate well. The scoring, not yet. The value of the exercise comes from two or three people disagreeing about a number in a room, and that conversation is the product.
References
NIST AI Risk Management Framework
NIST SP 800-30 Rev. 1: Guide for Conducting Risk Assessments
NIST AI 100-2 E2025: Adversarial Machine Learning
ICO Guidance on AI and Data Protection
ICO: Data Protection Impact Assessments
NCSC Guidelines for Secure AI System Development
NCSC Risk Management Collection
European Commission: Regulatory Framework for Artificial Intelligence
AI Regulation: A Pro-Innovation Approach
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.