Agent Risk Assessment Matrix
Eight questions place each AI agent on a four-by-four matrix of autonomy against consequence. The result is an oversight tier, the controls that tier requires, and the ISO/IEC 42001:2023 clause each control satisfies, printed as an auditor-ready register. Load the worked example to see five anonymised agents already filled in, then replace them with your own. No email required.
Listen to this briefing
The Agent Risk Assessment Matrix
Why agents need their own risk register
ISO/IEC 42001 requires an AI risk assessment and an impact assessment for every AI system in scope. The standard gives no method for rating an agent that reads supplier emails, plans actions across several systems and posts a journal entry without approval. Most registers I review rate that agent as a chatbot, because the template was designed for systems that answer questions.
This matrix rates every agent on two separate questions: how much it does without a person, and how much damage a wrong action causes. Scoring the two questions on separate axes gives an oversight decision the risk committee can defend and repeat for every agent in the register.
The clause references are to the published standard. The band thresholds and the tier matrix are my own calibration from applying the standard to agent deployments, and the basis table below records the source of each rule.
Your agent risk register
Each agent you add appears on the matrix and in the register below. All entries remain in your browser and are not transmitted.
The outlined marker is the agent in the form. Numbered markers are agents in the register.
| # | Agent | Autonomy | Consequence | Tier | Controls | Holds | Actions |
|---|---|---|---|---|---|---|---|
| No agents yet. Add one above, or load the worked example. | |||||||
The PDF register gives each agent one page recording the eight answers, the oversight tier, every required control with its ISO/IEC 42001 reference, and a sign-off block for the owner, the approver and residual-risk acceptance. The worked example is an illustrative composite and none of its five agents is a client record.
This tool produces a self-assessment based on your own description of each agent. It does not constitute an independent risk assessment, an audit or a legal opinion. The tier and the control set are the starting point for your own risk treatment plan and statement of applicability, which remain your responsibility. Your entries stay in your browser and are neither sent to Terence Kok nor reviewed by anyone.
The seventeen controls and where ISO/IEC 42001 asks for them
Seven controls apply to every agent at every tier. The remaining ten are required by the agent's tier or by a single answer that makes them necessary at any tier. The register states the reason each control applies.
Show all seventeen controls with their ISO/IEC 42001 references ▸
| Control | What the register row has to evidence | Applies | ISO/IEC 42001:2023 |
|---|---|---|---|
| C01 Register the agent with a named owner | An inventory entry stating purpose, intended use, the systems it touches, the model and platform behind it, and one accountable person by name. Also: Governance Baseline Q2. | All tiers | 4.3 Scope; A.3.2 AI roles and responsibilities; A.9.4 Intended use |
| C02 Enumerate the action allowlist | A dated list of every action the agent can take without approval, covering read, write, external communication and financial commitment. Anything not listed is denied. Also: Governance Baseline Q1; RA control B3. | All tiers | A.9.4 Intended use; A.4.4 Tooling resources; A.10.2 Allocating responsibilities |
| C03 Log every run end to end | Inputs, retrieved sources and their versions, each tool call, model and prompt version, approvals, and outputs, retained for the period your risk criteria set. Also: TRACE T; RA control E1. | All tiers | A.6.2.8 Recording of event logs; 9.1 Monitoring, measurement, analysis and evaluation |
| C04 Set acceptance criteria and test against them before release | Pre-agreed, measurable criteria independent of the agent’s own output, with the evaluation results filed before go-live. Also: TRACE A. | All tiers | A.6.2.2 Requirements and specification; A.6.2.4 Verification and validation |
| C05 Control changes to model, prompt and policy | Model, prompt and policy versions held in a register; every change is a release with an approver and a re-run of the acceptance tests. Also: RA control E3. | All tiers | A.6.2.7 Technical documentation; 8.1 Operational planning and control |
| C06 Tell users what it is and what it is for | Users and affected parties are told they are dealing with an AI system, what it is intended to do, and how to raise a concern. | All tiers | A.8.2 System documentation and information for users; A.3.3 Reporting of concerns |
| C07 Demonstrate the halt before go-live | A person has stopped the agent mid-task and the stop worked, with the date and the operator recorded. Repeat the exercise on a defined interval. Also: TRACE E; RA control E4. | All tiers | A.6.2.6 Operation and monitoring; A.8.4 Communication of incidents |
| C08 Put a human checkpoint before consequential actions | The actions that wait for a person, and the thresholds that define them (amount, record type, recipient), are documented and enforced in application code, since instructions in the prompt cannot be relied on. Also: Governance Baseline Q2; RA control D2. | Tier 2 and abovePartly or fully irreversible actions | A.9.2 Responsible use; A.3.2 AI roles and responsibilities |
| C09 Record an AI system impact assessment | A documented assessment of the effect on individuals, groups and society, reviewed when the scope or the affected population changes. | Tier 2 and aboveCustomers, the public or third parties affectedPersonal data | 6.1.4 and 8.4 AI system impact assessment; A.5.2, A.5.3, A.5.4 |
| C10 Write the incident procedure and name the suspension authority | Who is notified, what is reverted, how it is logged, who can suspend the agent, and what counts as a reportable event under the rules that apply. Also: Governance Baseline Q3. | Tier 2 and aboveStatutory or licensed activity | A.8.4 Communication of incidents; A.8.3 External reporting; 10.2 Nonconformity and corrective action |
| C11 Keep the manual fallback exercised | Staff can do the task by hand if the agent is stopped, suspended or changes behaviour after a model update, and have done so within a defined period. Also: Governance Baseline Q4. | Tier 2 and above | A.6.2.6 Operation and monitoring; A.4.6 Human resources |
| C12 Screen untrusted input and treat retrieved content as data | Inbound content is classified for injection patterns and stripped of secrets; retrieved text can never widen the agent’s permissions. Also: RA controls A2, C2. | Tier 2 and aboveReads or receives untrusted content | A.6.2.6 Operation and monitoring; 6.1.2 AI risk assessment |
| C13 Require a named approver per consequential action with the trace attached | Each consequential action waits for a specific person who sees the reasoning trace, and the approver queue is measured to ensure review remains substantive. Also: RA control D2. | Tier 3 and above | A.9.2 Responsible use; A.3.2 AI roles and responsibilities |
| C14 Enforce data access at retrieval and provenance on every source | Document-level access control applied at query time, every passage tagged with source, version and trust tier; access control itself under ISO/IEC 27001. Also: RA control C2. | Tier 3 and abovePersonal or regulated data | A.7.5 Data provenance; A.7.4 Quality of data; A.7.2 Data for AI systems |
| C15 Allocate supplier and platform responsibilities in writing | The model provider, the agent platform and any integrator each have documented responsibilities for security, availability, change notice and incident reporting. | Tier 3 and aboveStatutory or licensed activity | A.10.2 Allocating responsibilities; A.10.3 Suppliers |
| C16 Bound delegation between agents | An agent that instructs other agents passes on a narrower permission set than its own, and the full chain is visible in one log. Also: RA risk register, agent-to-agent escalation. | By trigger onlyInstructs other agents | A.6.2.2 Requirements and specification; A.4.4 Tooling resources; A.6.2.8 Event logs |
| C17 Independent review before go-live and after material change | Internal audit or an independent function reviews the register row, the evidence behind it and the residual-risk acceptance before deployment and after any material change. | Tier 3 and above | 9.2 Internal audit; 9.3 Management review; 6.1.3 AI risk treatment |
Annex A references are to the control objectives and controls of ISO/IEC 42001:2023. Clause references without an "A." prefix are to the management system requirements. Access control, logging integrity and supplier security are ISO/IEC 27001 controls, referenced and not duplicated.
What each rule in the matrix rests on
Every rule in the matrix is traced to a published standard, framework or reference architecture. The rules that rest on my own calibration are marked as such in the table.
Show the ten rules and their sources ▸
| Rule | Source |
|---|---|
| Risk is assessed per AI system, with likelihood and consequence, and treated through documented controls | ISO/IEC 42001:2023 clauses 6.1.2, 6.1.3, 8.2 and 8.3 Source → |
| An impact assessment on individuals, groups and society is recorded separately from the organisational risk assessment | ISO/IEC 42001:2023 clauses 6.1.4 and 8.4; Annex A.5 Source → |
| Each control in the register is referenced to the Annex A control or clause it evidences | ISO/IEC 42001:2023 Annex A and the statement of applicability (clause 6.1.3) Source → |
| Reversibility decides the human checkpoint: irreversible actions wait for a named person | TRACE Framework v1.1, criterion R and the three reversibility tiers Source → |
| An agent that cannot be halted mid-task is not deployed, whatever its model performance | TRACE Framework v1.1, criterion E Source → |
| The four register fields an auditor asks for first: allowed actions, named reviewer, error procedure, manual fallback | Four-Question AI Governance Baseline v1.1 Source → |
| Excessive agency is contained by least-privilege tool permissions, a human checkpoint by tier and output guardrails | OWASP Top 10 for LLM Applications 2025, LLM06; OWASP Agentic AI Threats T2, T3, T10 Source → |
| Controls C01 to C17 correspond to the control groups of the Governed Agentic RAG reference architecture | Governed Agentic RAG reference architecture v1.0, control set and risk register Source → |
| Map, measure and manage AI risk across the life cycle, with monitoring after deployment | NIST AI Risk Management Framework 1.0, MAP, MEASURE and MANAGE functions Source → |
| The band thresholds (Assisted 0–2, Supervised 3–5, Bounded 6–7, Autonomous 8–9; Low 0–4, Moderate 5–8, High 9–12, Severe 13–15) and the tier matrix | Practitioner calibration by Terence Kok from applying ISO/IEC 42001 to agent deployments. Not prescribed by the standard. Re-cut the bands to your own risk criteria under clause 6.1.1 if your appetite differs. |
How to read the matrix and use the register
Two axes instead of one score
Autonomy measures how much the agent does before a person sees the result, based on its action scope, the position of the human checkpoint and the reach of a single run. Consequence measures the damage a wrong action causes, based on reversibility, data sensitivity, blast radius, exposure to untrusted input and regulatory weight. The oversight tier is read from the cell where the two scores meet.
- Two agents with the same combined score can require opposite treatment, since a high-autonomy summariser with low consequence needs monitoring and a low-autonomy payment drafter with high consequence needs a named approver
- The least costly route out of Tier 4 is usually to reduce autonomy by placing the person earlier in the loop, because reducing consequence usually means redesigning the task
- The matrix records that trade-off explicitly, which is the evidence a risk committee needs to accept residual risk under clause 6.1.1
- Rate the agent as it will operate in production, based on the review that will take place in practice
- If the tier is higher than expected, change one answer at a time and observe how the marker moves
- Record any contested answers in the register notes so that the next review can test them
Where the register lands in the AIMS
Each register row is a clause 6.1.2 risk assessment and, where personal data or the public are involved, a clause 6.1.4 impact assessment for one agent. The control set is the input to the risk treatment plan under 6.1.3. The ISO reference against each control is carried into the statement of applicability.
- A certification auditor first asks for the list of AI systems in scope and the assessment behind each one, and a register with one row per agent satisfies that request
- A register row with a named owner, a dated halt test and an approver satisfies three of the four Governance Baseline questions on its own
- The standard does not distinguish agents from other AI systems, so your register must make that distinction to prevent chatbot controls being applied to an agent that moves money
- Print the register and file it with the AI system inventory under clause 4.3
- Copy each control's ISO reference into the statement of applicability with "applied" or "not applied" and the reason
- Re-run the row after any model, prompt, scope or supplier change, and date the re-run
Holds that override the tier
Three conditions hold deployment regardless of tier: the halt has never been demonstrated, no single person owns the agent, or a Tier 2 or higher agent has no named approver. The register records each as a separate hold so that a low score cannot offset it.
- A documented kill switch that has never been exercised is the most common finding I raise on agent deployments, and it takes an afternoon to close
- Ownership by a team does not satisfy the standard, which requires roles and authorities assigned to a named individual whom an auditor will ask to describe the agent
- Clearing a hold before go-live costs far less than clearing it after the first incident
- Clear every hold before the row goes to the risk committee
- Schedule the halt exercise on the same cadence as the tier review
- If a Tier 4 agent is already live, treat the row as an open nonconformity under clause 10.2 and start the redesign
About This Tool
This matrix was built by Terence Kok, a certified ISO/IEC 42001 Lead Auditor and AIGP, who has unified AI governance across three national regulatory environments, Singapore's IMDA among them, under a single standard. The control set follows the Governed Agentic RAG reference architecture, whose machine-readable form is in the public AI governance toolkit.
See the full certification list →This matrix rates agents you have already decided to build. To test whether a task should be handed to an agent at all, run it through the TRACE Agent Evaluation first. To check whether the management system around the register would survive certification, use theISO 42001 AIMS Readiness Checklist. To test whether your board could answer for the agents in the register, use theBoard AI Oversight Checklist.
Taking the register through to a statement of applicability
The matrix places each agent in its oversight tier. Agreeing the risk criteria, clearing the holds and writing the treatment plan is work I do directly with clients.
Book a private session
I built this tool after reviewing the same register three times in one quarter, each one a chatbot template relabelled for agents. The two-axis matrix is the method I use in my own engagements, and its value is that a summariser and a payment agent can never share a row. Load the worked example first. The public-sector agent in Tier 4 is deliberate, because that is where I most often find agents already in production.

