Agent Risk Assessment Matrix

Eight questions place each AI agent on a four-by-four matrix of autonomy against consequence. The result is an oversight tier, the controls that tier requires, and the ISO/IEC 42001:2023 clause each control satisfies, printed as an auditor-ready register. Load the worked example to see five anonymised agents already filled in, then replace them with your own. No email required.

Listen to this briefing

The Agent Risk Assessment Matrix

0:00

Why agents need their own risk register

ISO/IEC 42001 requires an AI risk assessment and an impact assessment for every AI system in scope. The standard gives no method for rating an agent that reads supplier emails, plans actions across several systems and posts a journal entry without approval. Most registers I review rate that agent as a chatbot, because the template was designed for systems that answer questions.

This matrix rates every agent on two separate questions: how much it does without a person, and how much damage a wrong action causes. Scoring the two questions on separate axes gives an oversight decision the risk committee can defend and repeat for every agent in the register.

Figure 1Two agents one autonomy band apart sit three oversight tiers apart once consequence is scored.
Two axes instead of one scoreA four-by-four matrix of autonomy (Assisted, Supervised, Bounded, Autonomous) against consequence (Low, Moderate, High, Severe). Each cell names an oversight tier from 1 Monitor to 4 Hold. Two example agents sit one autonomy band apart: a meeting summariser with no review lands at Supervised autonomy and Low consequence in Tier 1 Monitor; a benefit payment agent sampled after the fact lands at Bounded autonomy and Severe consequence in Tier 4 Hold.ConsequenceAutonomySevereTIER 2SuperviseTIER 3ApproveTIER 4HoldTIER 4HoldHighTIER 2SuperviseTIER 2SuperviseTIER 3ApproveTIER 4HoldModerateTIER 1MonitorTIER 2SuperviseTIER 2SuperviseTIER 3ApproveLowTIER 1MonitorTIER 1MonitorTIER 1MonitorTIER 2SuperviseAssistedSupervisedBoundedAutonomous22Meeting summariser, no reviewReads and reports across several systems; nobodychecks the output unless a problem is raised.Supervised × LowTier 1 Monitor11Benefit payment agent, sampled afterwardsPays claimants within a limit across several systems;money sent, regulated data, a decision on entitlements.Bounded × SevereTier 4 Hold
Figure 2Eight answers produce two banded scores, one matrix lookup and one oversight tier with its required controls and holds.
How the eight answers become a register rowThree answers about how much the agent does on its own are scored 0 to 3 each and summed to an autonomy score out of 9, banded Assisted, Supervised, Bounded or Autonomous. Five answers about how bad a wrong action is are summed to a consequence score out of 15, banded Low, Moderate, High or Severe. The two bands are looked up in a four-by-four matrix to give an oversight tier from 1 to 4. The register row records the tier and its oversight mode, the control set (7 controls for every agent plus more by tier or by a single answer), and any holds.1. Eight answersA. How much it does on its ownAction scopeHuman checkpointTask chainscored 0–3 eachB. How bad a wrong action isReversibilityData sensitivityBlast radiusExposure to untrusted inputRegulatory weightscored 0–3 each2. Two scores, bandedAutonomy 0–9AssistedSupervisedBoundedAutonomousConsequence 0–15LowModerateHighSevere3. Matrix lookup2344223412231112Autonomy →↑ Consequence1 Monitor2 Supervise3 Approve4 Hold4. Register rowOversight tierTier 1 to 4 and thereview it demandsControl set7 for every agent,more by tier or by asingle answerHoldsHalt never tested,no named owner,no approver at Tier 2+Each control carries its ISO/IEC 42001:2023 clause, so the row reads into the risk treatment plan and the statement of applicability.

The clause references are to the published standard. The band thresholds and the tier matrix are my own calibration from applying the standard to agent deployments, and the basis table below records the source of each rule.

Describe one agent

Answer the eight questions for the agent as it will operate in production, including the review that will take place in practice.

A How much it does on its own 0 / 9

The agent retrieves, summarises or answers, and no system is changed.

Nothing executes until a person approves it.

A single call into a single tool or system.

B How bad a wrong action is 0 / 15

The action is reversed immediately and no one outside the team is affected.

Published material or internal information with no access restriction.

The person using it sees the error and corrects it.

Structured data from systems you control, entered by staff.

General law only; no sector or activity rule names this task.

C Register fields

Your agent risk register

Each agent you add appears on the matrix and in the register below. All entries remain in your browser and are not transmitted.

Consequence
Severe13–15
T2 Supervise
T3 Approve
T4 Hold
T4 Hold
High9–12
T2 Supervise
T2 Supervise
T3 Approve
T4 Hold
Moderate5–8
T1 Monitor
T2 Supervise
T2 Supervise
T3 Approve
Low0–4
T1 Monitor
T1 Monitor
T1 Monitor
T2 Supervise
Assisted0–2
Supervised3–5
Bounded6–7
Autonomous8–9
Autonomy

The outlined marker is the agent in the form. Numbered markers are agents in the register.

#AgentAutonomyConsequenceTierControlsHoldsActions
No agents yet. Add one above, or load the worked example.

The PDF register gives each agent one page recording the eight answers, the oversight tier, every required control with its ISO/IEC 42001 reference, and a sign-off block for the owner, the approver and residual-risk acceptance. The worked example is an illustrative composite and none of its five agents is a client record.

This tool produces a self-assessment based on your own description of each agent. It does not constitute an independent risk assessment, an audit or a legal opinion. The tier and the control set are the starting point for your own risk treatment plan and statement of applicability, which remain your responsibility. Your entries stay in your browser and are neither sent to Terence Kok nor reviewed by anyone.

The seventeen controls and where ISO/IEC 42001 asks for them

Seven controls apply to every agent at every tier. The remaining ten are required by the agent's tier or by a single answer that makes them necessary at any tier. The register states the reason each control applies.

Figure 3How the control set stacks. Tier 4 carries the same controls as Tier 3 and adds a hold on deployment.
How the control set stacks by tierFour columns, one per oversight tier. Tier 1 carries the 7 controls that apply to every agent (C01–C07). Tier 2 adds 5 (C08–C12). Tier 3 adds 4 (C13, C14, C15, C17). Tier 4 carries the same 16 controls and deployment is held until the agent is redesigned. Below the columns, the answers that pull a control in at any tier: C08: partly or fully irreversible actions; C09: customers, the public or third parties affected, or personal data; C10: statutory or licensed activity; C12: reads or receives untrusted content; C14: personal or regulated data; C15: statutory or licensed activity; C16: instructs other agents (this answer alone, at any tier).C01–C077Tier 1 Monitor7 controlsC01–C077C08–C125Tier 2 Supervise12 controlsC01–C077C08–C125C13, C14, C15, C174Tier 3 Approve16 controlsC01–C077C08–C125C13, C14, C15, C174Do not deploy as scoredTier 4 Hold16 controlsEvery agent: 7 controlsTier 2 adds: 5 controlsTier 3 adds: 4 controlsPulled in at any tier by a single answerC08: partly or fully irreversible actionsC09: customers, the public or third parties affected, or personal dataC10: statutory or licensed activityC12: reads or receives untrusted contentC14: personal or regulated dataC15: statutory or licensed activityC16: instructs other agents (this answer alone, at any tier)
Show all seventeen controls with their ISO/IEC 42001 references
ControlWhat the register row has to evidenceAppliesISO/IEC 42001:2023
C01 Register the agent with a named ownerAn inventory entry stating purpose, intended use, the systems it touches, the model and platform behind it, and one accountable person by name. Also: Governance Baseline Q2.All tiers4.3 Scope; A.3.2 AI roles and responsibilities; A.9.4 Intended use
C02 Enumerate the action allowlistA dated list of every action the agent can take without approval, covering read, write, external communication and financial commitment. Anything not listed is denied. Also: Governance Baseline Q1; RA control B3.All tiersA.9.4 Intended use; A.4.4 Tooling resources; A.10.2 Allocating responsibilities
C03 Log every run end to endInputs, retrieved sources and their versions, each tool call, model and prompt version, approvals, and outputs, retained for the period your risk criteria set. Also: TRACE T; RA control E1.All tiersA.6.2.8 Recording of event logs; 9.1 Monitoring, measurement, analysis and evaluation
C04 Set acceptance criteria and test against them before releasePre-agreed, measurable criteria independent of the agent’s own output, with the evaluation results filed before go-live. Also: TRACE A.All tiersA.6.2.2 Requirements and specification; A.6.2.4 Verification and validation
C05 Control changes to model, prompt and policyModel, prompt and policy versions held in a register; every change is a release with an approver and a re-run of the acceptance tests. Also: RA control E3.All tiersA.6.2.7 Technical documentation; 8.1 Operational planning and control
C06 Tell users what it is and what it is forUsers and affected parties are told they are dealing with an AI system, what it is intended to do, and how to raise a concern.All tiersA.8.2 System documentation and information for users; A.3.3 Reporting of concerns
C07 Demonstrate the halt before go-liveA person has stopped the agent mid-task and the stop worked, with the date and the operator recorded. Repeat the exercise on a defined interval. Also: TRACE E; RA control E4.All tiersA.6.2.6 Operation and monitoring; A.8.4 Communication of incidents
C08 Put a human checkpoint before consequential actionsThe actions that wait for a person, and the thresholds that define them (amount, record type, recipient), are documented and enforced in application code, since instructions in the prompt cannot be relied on. Also: Governance Baseline Q2; RA control D2.Tier 2 and abovePartly or fully irreversible actionsA.9.2 Responsible use; A.3.2 AI roles and responsibilities
C09 Record an AI system impact assessmentA documented assessment of the effect on individuals, groups and society, reviewed when the scope or the affected population changes.Tier 2 and aboveCustomers, the public or third parties affectedPersonal data6.1.4 and 8.4 AI system impact assessment; A.5.2, A.5.3, A.5.4
C10 Write the incident procedure and name the suspension authorityWho is notified, what is reverted, how it is logged, who can suspend the agent, and what counts as a reportable event under the rules that apply. Also: Governance Baseline Q3.Tier 2 and aboveStatutory or licensed activityA.8.4 Communication of incidents; A.8.3 External reporting; 10.2 Nonconformity and corrective action
C11 Keep the manual fallback exercisedStaff can do the task by hand if the agent is stopped, suspended or changes behaviour after a model update, and have done so within a defined period. Also: Governance Baseline Q4.Tier 2 and aboveA.6.2.6 Operation and monitoring; A.4.6 Human resources
C12 Screen untrusted input and treat retrieved content as dataInbound content is classified for injection patterns and stripped of secrets; retrieved text can never widen the agent’s permissions. Also: RA controls A2, C2.Tier 2 and aboveReads or receives untrusted contentA.6.2.6 Operation and monitoring; 6.1.2 AI risk assessment
C13 Require a named approver per consequential action with the trace attachedEach consequential action waits for a specific person who sees the reasoning trace, and the approver queue is measured to ensure review remains substantive. Also: RA control D2.Tier 3 and aboveA.9.2 Responsible use; A.3.2 AI roles and responsibilities
C14 Enforce data access at retrieval and provenance on every sourceDocument-level access control applied at query time, every passage tagged with source, version and trust tier; access control itself under ISO/IEC 27001. Also: RA control C2.Tier 3 and abovePersonal or regulated dataA.7.5 Data provenance; A.7.4 Quality of data; A.7.2 Data for AI systems
C15 Allocate supplier and platform responsibilities in writingThe model provider, the agent platform and any integrator each have documented responsibilities for security, availability, change notice and incident reporting.Tier 3 and aboveStatutory or licensed activityA.10.2 Allocating responsibilities; A.10.3 Suppliers
C16 Bound delegation between agentsAn agent that instructs other agents passes on a narrower permission set than its own, and the full chain is visible in one log. Also: RA risk register, agent-to-agent escalation.By trigger onlyInstructs other agentsA.6.2.2 Requirements and specification; A.4.4 Tooling resources; A.6.2.8 Event logs
C17 Independent review before go-live and after material changeInternal audit or an independent function reviews the register row, the evidence behind it and the residual-risk acceptance before deployment and after any material change.Tier 3 and above9.2 Internal audit; 9.3 Management review; 6.1.3 AI risk treatment

Annex A references are to the control objectives and controls of ISO/IEC 42001:2023. Clause references without an "A." prefix are to the management system requirements. Access control, logging integrity and supplier security are ISO/IEC 27001 controls, referenced and not duplicated.

What each rule in the matrix rests on

Every rule in the matrix is traced to a published standard, framework or reference architecture. The rules that rest on my own calibration are marked as such in the table.

Show the ten rules and their sources
RuleSource
Risk is assessed per AI system, with likelihood and consequence, and treated through documented controlsISO/IEC 42001:2023 clauses 6.1.2, 6.1.3, 8.2 and 8.3 Source →
An impact assessment on individuals, groups and society is recorded separately from the organisational risk assessmentISO/IEC 42001:2023 clauses 6.1.4 and 8.4; Annex A.5 Source →
Each control in the register is referenced to the Annex A control or clause it evidencesISO/IEC 42001:2023 Annex A and the statement of applicability (clause 6.1.3) Source →
Reversibility decides the human checkpoint: irreversible actions wait for a named personTRACE Framework v1.1, criterion R and the three reversibility tiers Source →
An agent that cannot be halted mid-task is not deployed, whatever its model performanceTRACE Framework v1.1, criterion E Source →
The four register fields an auditor asks for first: allowed actions, named reviewer, error procedure, manual fallbackFour-Question AI Governance Baseline v1.1 Source →
Excessive agency is contained by least-privilege tool permissions, a human checkpoint by tier and output guardrailsOWASP Top 10 for LLM Applications 2025, LLM06; OWASP Agentic AI Threats T2, T3, T10 Source →
Controls C01 to C17 correspond to the control groups of the Governed Agentic RAG reference architectureGoverned Agentic RAG reference architecture v1.0, control set and risk register Source →
Map, measure and manage AI risk across the life cycle, with monitoring after deploymentNIST AI Risk Management Framework 1.0, MAP, MEASURE and MANAGE functions Source →
The band thresholds (Assisted 0–2, Supervised 3–5, Bounded 6–7, Autonomous 8–9; Low 0–4, Moderate 5–8, High 9–12, Severe 13–15) and the tier matrixPractitioner calibration by Terence Kok from applying ISO/IEC 42001 to agent deployments. Not prescribed by the standard. Re-cut the bands to your own risk criteria under clause 6.1.1 if your appetite differs.

How to read the matrix and use the register

Two axes instead of one score

Purpose

Autonomy measures how much the agent does before a person sees the result, based on its action scope, the position of the human checkpoint and the reach of a single run. Consequence measures the damage a wrong action causes, based on reversibility, data sensitivity, blast radius, exposure to untrusted input and regulatory weight. The oversight tier is read from the cell where the two scores meet.

Why It Matters
  • Two agents with the same combined score can require opposite treatment, since a high-autonomy summariser with low consequence needs monitoring and a low-autonomy payment drafter with high consequence needs a named approver
  • The least costly route out of Tier 4 is usually to reduce autonomy by placing the person earlier in the loop, because reducing consequence usually means redesigning the task
  • The matrix records that trade-off explicitly, which is the evidence a risk committee needs to accept residual risk under clause 6.1.1
How to Use
  1. Rate the agent as it will operate in production, based on the review that will take place in practice
  2. If the tier is higher than expected, change one answer at a time and observe how the marker moves
  3. Record any contested answers in the register notes so that the next review can test them

Where the register lands in the AIMS

Purpose

Each register row is a clause 6.1.2 risk assessment and, where personal data or the public are involved, a clause 6.1.4 impact assessment for one agent. The control set is the input to the risk treatment plan under 6.1.3. The ISO reference against each control is carried into the statement of applicability.

Why It Matters
  • A certification auditor first asks for the list of AI systems in scope and the assessment behind each one, and a register with one row per agent satisfies that request
  • A register row with a named owner, a dated halt test and an approver satisfies three of the four Governance Baseline questions on its own
  • The standard does not distinguish agents from other AI systems, so your register must make that distinction to prevent chatbot controls being applied to an agent that moves money
How to Use
  1. Print the register and file it with the AI system inventory under clause 4.3
  2. Copy each control's ISO reference into the statement of applicability with "applied" or "not applied" and the reason
  3. Re-run the row after any model, prompt, scope or supplier change, and date the re-run

Holds that override the tier

Purpose

Three conditions hold deployment regardless of tier: the halt has never been demonstrated, no single person owns the agent, or a Tier 2 or higher agent has no named approver. The register records each as a separate hold so that a low score cannot offset it.

Why It Matters
  • A documented kill switch that has never been exercised is the most common finding I raise on agent deployments, and it takes an afternoon to close
  • Ownership by a team does not satisfy the standard, which requires roles and authorities assigned to a named individual whom an auditor will ask to describe the agent
  • Clearing a hold before go-live costs far less than clearing it after the first incident
How to Use
  1. Clear every hold before the row goes to the risk committee
  2. Schedule the halt exercise on the same cadence as the tier review
  3. If a Tier 4 agent is already live, treat the row as an open nonconformity under clause 10.2 and start the redesign
ISO/IEC 42001 Lead Auditor badge

About This Tool

This matrix was built by Terence Kok, a certified ISO/IEC 42001 Lead Auditor and AIGP, who has unified AI governance across three national regulatory environments, Singapore's IMDA among them, under a single standard. The control set follows the Governed Agentic RAG reference architecture, whose machine-readable form is in the public AI governance toolkit.

See the full certification list →

This matrix rates agents you have already decided to build. To test whether a task should be handed to an agent at all, run it through the TRACE Agent Evaluation first. To check whether the management system around the register would survive certification, use theISO 42001 AIMS Readiness Checklist. To test whether your board could answer for the agents in the register, use theBoard AI Oversight Checklist.

Taking the register through to a statement of applicability

The matrix places each agent in its oversight tier. Agreeing the risk criteria, clearing the holds and writing the treatment plan is work I do directly with clients.

Book a private session
Terence Kok
Before You Go

I built this tool after reviewing the same register three times in one quarter, each one a chatbot template relabelled for agents. The two-axis matrix is the method I use in my own engagements, and its value is that a summariser and a payment agent can never share a row. Load the worked example first. The public-sector agent in Tier 4 is deliberate, because that is where I most often find agents already in production.

Terence Kok