Tip: use your browser’s Print dialog (⌘P / Ctrl+P) and “Save as PDF”.
Loop HAL · Complete reference · Version 1.1
Loop names three patterns for how humans govern AI: review (Human-in-the-Loop), monitor (Human-on-the-Loop), and own (Human-Accountable-for-the-Loop). HAL is the practical accountability framework for the third pattern — workflows where individual review does not scale.
Accountability cannot be delegated. Execution can.
“Human-in-the-Loop” is one of the most repeated phrases in AI governance. The problem is that people use it to mean very different things: reviewing outputs, monitoring systems, or remaining accountable for outcomes. Those are not the same governance model.
Loop gives clearer language. HAL does not replace Human-in-the-Loop; it is the detailed accountability framework for agentic systems, multi-agent orchestration, large-scale triage, and autonomous operational systems — where individual review does not scale.
HAL complements standards organisations already follow. The NIST AI Risk Management Framework (AI RMF 1.0) supports managing risks to people, organisations and society. ISO/IEC 42001:2023 specifies requirements for an AI management system. The EU AI Act creates legal obligations for actors and systems within its scope. Loop identifies which governance pattern applies; HAL examines whether delegated action remains bounded, evidenced and accountable.
Human-in-the-Loop made most sense when AI systems produced outputs for humans to use. As systems gained autonomy — taking actions, routing work, creating records — the same phrase was stretched to cover monitoring, process design, and accountability.
The established taxonomy runs in-the-loop → on-the-loop → out-of-the-loop. Loop deliberately replaces that third term. “Out-of-the-loop” frames the human as absent. The human is never out of the loop of accountability, only out of the loop of execution. That is why the third pattern is Human-Accountable-for-the-Loop.
Review
A human reviews or approves individual AI outputs before action is taken.
AI → Human Review → Action
Human role: Reviewer · Did a human review the decision?
Can the reviewer meaningfully challenge the system rather than merely approve its output?
An AI drafts a clause summary. A lawyer reviews it before relying on it.
Best for
Strengths
Limitations
Monitor
An AI system operates within defined boundaries while humans monitor behaviour, investigate exceptions and intervene when required.
AI → Action ↓ Human Monitoring
Human role: Monitor · Is a human monitoring the system?
Can the monitor recognise abnormal behaviour, understand its significance and know when intervention is required?
A regulatory monitoring system identifies potential changes and routes unusual or high-impact items for review.
Best for
Strengths
Limitations
Own
A human owns the design, authority, controls, evidence and outcomes of an AI-enabled workflow, even where the AI system acts without individual human review.
Owner → Policy → AI Workflow → Action → Evidence → Review
Human role: Owner · Who owns the system that made the decision?
Does the accountable person possess sufficient understanding and judgement to exercise authority over the system rather than merely carrying formal responsibility for it?
A matter triage agent classifies incoming legal requests, routes work, creates records, escalates risk and logs evidence. No human reviews every decision, but a named owner is accountable for the workflow, its controls and its outcomes.
Best for
Strengths
Limitations
| Workflow type | Recommended model |
|---|---|
| Drafting a memo | Human-in-the-Loop |
| Summarising a contract | Human-in-the-Loop |
| Monitoring regulatory updates | Human-on-the-Loop |
| Screening incoming matters | Human-on-the-Loop or HAL |
| Drafting a client communication | Human-in-the-Loop |
| Routing client requests | HAL |
| Updating matter records | HAL |
| Triggering external communications | HAL with approval gates or strict controls |
| Filing regulatory documents | HAL only with high score and approval gates |
| Example | Model | Why |
|---|---|---|
| Legal Research Assistant | Human-in-the-Loop | The AI generates outputs for a lawyer to review. It does not act independently. |
| Contract Summary Tool | Human-in-the-Loop | The human still decides whether to rely on the summary. |
| Regulatory Monitoring System | Human-on-the-Loop | The system monitors at scale and escalates relevant changes. |
| Legal Matter Triage Agent | HAL | The system classifies, routes and creates operational consequences. |
| Contract Obligation Tracker | HAL | The system may trigger reminders, escalate risk and affect deadline management. |
| Client Email Response Agent | HAL with approval gates | External communication creates higher risk. HAL may apply, but some actions may still require Human-in-the-Loop approval. |
For most of the history of machine decision-making, governance had a simple shape: a machine proposed, and a human disposed. We called this Human-in-the-Loop, and it served us well. That model is now under strain, because the premise it rested on no longer holds: that the human is the one who acts.
Terms, used precisely
Responsibility: doing the work. Can be delegated, including to a system.
Accountability: being answerable for the outcome. In law, this may sit with the organisation or with particular office-holders. In governance, HAL requires it to be traceable to a named accountable owner and not displaced onto software.
Liability: bearing legal or financial consequences. The economic burden of liability can often be allocated by contract, indemnity and insurance, but statutory, regulatory and third-party liability may remain where the law places it.
When HAL says accountability can never be delegated, execution can, it means exactly this: responsibility for execution can move. Accountability for the system cannot.
Human-in-the-Loop (HITL) did not begin with machine learning. It is an inheritance from control theory, aviation, and early automation, where a human operator remained the final authority over a machine that could otherwise act on its own. The principle was conservative and sound: where a machine might err in ways that matter, insert a human between its judgement and the world.
When predictive models entered high-stakes settings (credit, medicine, hiring), HITL was the natural control. The model scored; a human decided. Oversight and action lived in the same place: the person.
HITL worked because of a structural fact about the systems it governed: they did not act. They produced outputs (scores, classifications, drafts, recommendations) and then stopped. The output was inert until a human picked it up. That gap between output and action was where governance lived.
It also worked because the volumes were human-scaled. A loan officer could review the applications in front of them. A clinician could weigh a model’s reading against their own. The number of decisions was bounded by human capacity, and so the review was real.
Two things break HITL: volume and action. When a system makes more decisions than a human can examine, “human-in-the-loop” becomes “human-rubber-stamping-the-loop”. The control persists on paper while evaporating in practice.
A reviewer who cannot review is not a control. They are a liability with a job title.
The deeper break comes when systems begin to act. The moment a system can send the email, file the report, move the money, or delete the record, the gap between output and action closes. There is no longer a natural pause in which a human stands. To insert one artificially, by requiring approval of every action, is to throw away the very capability you deployed the system for.
These are three distinct governance patterns, not stages of maturity. Human-in-the-Loop asks whether a human reviewed the decision. Human-on-the-Loop asks whether a human is monitoring the system. Human-Accountable-for-the-Loop asks who owns the system that made the decision.
This distinction also sits more comfortably with how modern regulation describes oversight. Article 14 of the EU AI Act requires high-risk AI systems to be designed and developed so they can be effectively overseen by natural persons during use. Depending on context and the oversight measures identified by the provider, those persons need the competence, authority and tools to understand limitations, monitor anomalies, guard against automation bias, interpret or disregard outputs, and interrupt the system. The Act therefore treats human oversight as a set of functional controls, not a single undifferentiated requirement. Under the amended timetable, the high-risk rules apply to Annex III systems from 2 December 2027 and to systems covered through Annex I product legislation from 2 August 2028.
Agentic systems take actions in pursuit of goals, often chaining many steps, calling tools, and operating with minimal supervision. Once invoked or authorised, they can proceed across steps without returning to a human for each decision. That shift changes what governance must cover. An actor needs a different kind of governance: whether the action was authorised, bounded, recorded, and owned.
The point of Loop is not to rank these patterns from weak to strong. A human-in-the-loop model may be exactly right for a legal research assistant. A human-on-the-loop model may be right for a monitoring system. HAL may be necessary where the system acts at scale and individual review is no longer meaningful. The mistake is not choosing HITL. The mistake is using HITL as a generic label when the actual control is review, monitoring, ownership, or some unstable mixture of all three.
Execution can be delegated to a system. Accountability cannot. When a human employee acts within their authority, the organisation remains accountable for what they do. Nothing about substituting software for the employee changes that. The organisation, and a named person within it, remains answerable.
HAL fixes accountability to a person who owns the system. Software has no legal personality and no institutional accountability of its own. It cannot be sanctioned, dismissed, disciplined or struck off. So accountability flows back, through the system, to the human who owns it. That is the entire substance of Human Accountable for the Loop.
There is an assumption sitting underneath all three Loop patterns: the human involved is capable of exercising the judgement the operating model assigns to them. Accountability without capability can become ceremonial.
Capability is not a static resource. For years, professional processes produced two outputs: the visible work product, and people capable of handling harder work later. AI allows those outputs to separate. That is not an argument against automation, and not an argument for preserving repetitive junior work. Organisations need to identify which developmental effects remain important and deliberately preserve, compress, simulate or move them.
The question is not only whether a human remains in the loop. It is whether the system continues to produce humans capable of meaningfully being there.
This does not create a fourth Loop pattern or a ninth HAL domain. It changes what the existing disciplines must prove. Historic qualification is not permanent evidence of current capability.
The law of agency is not directly about software, but it is useful because it shows how law has long dealt with delegated action: a principal is bound by the authorised acts of their agent, and an agent who exceeds their authority creates liability that does not simply vanish.
One precision matters: software is not an agent in the legal sense. It has no legal personality, owes no fiduciary duty, and no agency relationship arises when you deploy it. In law, an AI system is a tool, and the consequences of a tool’s operation land directly on the organisation that wields it. The parallel HAL draws is structural, not doctrinal: the mechanics of sound delegation (scoped authority, limits, escalation, records) transfer; the legal relationship does not.
A regulator faced with an erroneous automated filing will not accept “the model did it” as a defence. HAL asks organisations to confront that reality before deployment, rather than discover it during an incident.
Organisations will run many agents at once. Governing this estate will resemble portfolio management more than decision-by-decision review: a registry of agents, each with an owner, an authority scope, a risk level, a HAL control profile, and a review date. Governance itself will become continuous. And throughout, one line will not be allowed to blur: however deep the stack of agents, accountability terminates at a human. The future of AI governance is about ensuring that, when systems act, someone remains accountable for the system that acted, and that the organisation can still produce people capable of exercising that accountability.
HAL evaluates whether a workflow is ready to delegate action while retaining clear human accountability. Eight domains form the framework.
01 · Ownership
A named person owns the system and its outcomes.
Definition
A single, named individual is accountable for the agentic system: one person, not a diffuse committee or "the platform". Ownership is the anchor on which every other domain depends.
Why it matters
Accountability cannot be delegated to software. When a system acts, someone must be answerable for what it did and did not do. Diffuse ownership is the most common root cause of governance failure: when everyone is responsible, no one is.
Failure mode
An automated client-communication agent sends an incorrect legal deadline to 400 clients. The incident review finds the agent was "owned by the AI working group". No individual can explain its authority, and no one has the standing to switch it off.
Good practice
Each deployed agent has a named accountable owner recorded in a register, with a deputy. The owner has signed off on the authority granted, reviews incidents, and holds the documented power to suspend the system immediately.
Implementation guidance
Questions to ask
02 · Authority
The system acts only within explicitly delegated authority.
Definition
The scope of what the system is permitted to do is defined explicitly and granted deliberately, mirroring how authority is delegated to a human employee.
Why it matters
Agentic systems act. An action taken outside granted authority is an unauthorised act for which the organisation remains liable. Authority must be designed in before deployment.
Failure mode
A procurement agent built to "draft" purchase orders is also able to submit them because no one constrained the integration. It commits the company to £2.3m of spend before anyone notices the boundary was never set.
Good practice
Authority is expressed as an explicit allow-list of actions, value thresholds, and counterparties. Anything outside the list is impossible by construction rather than discouraged by a prompt.
Implementation guidance
Questions to ask
03 · Limits
Hard constraints stop the system before it causes harm.
Definition
Beyond authority, limits are the non-negotiable boundaries: the actions the system must never take, and the conditions under which it must stop, regardless of confidence.
Why it matters
Authority says what is allowed; limits say what is forbidden and where the floor is. Limits are what protect you when the model is confidently wrong. They convert "the model decided not to" into "the system could not".
Failure mode
A collections agent is permitted to contact customers. With no limit on frequency, a logic loop causes it to email one vulnerable customer 71 times in a day. There was authority to contact; there was no limit on contact.
Good practice
Hard limits (maximum spend, prohibited actions, rate caps, protected data classes, and "never act on this customer segment") are enforced as guard rails that halt the system and escalate, independent of the model.
Implementation guidance
Questions to ask
04 · Escalation
Uncertainty and edge cases route to a human cleanly.
Definition
Defined paths by which the system hands control to a human when it is uncertain, when it hits a limit, or when a case falls outside its competence.
Why it matters
A system that never escalates is a system that has been told to guess. The quality of governance is often the quality of its escalation: clear triggers, a named recipient, and a defined response time.
Failure mode
A triage agent encounters a matter type it has never seen. With no escalation path, it forces the case into the nearest category and routes it incorrectly. The deadline is missed; no human was ever asked.
Good practice
Escalation triggers are explicit (low confidence, novel case, limit reached, high stakes). Each routes to a named human or role with a defined SLA, and the system pauses the relevant action until a human responds.
Implementation guidance
Questions to ask
05 · Evidence
Every decision and action is recorded and reconstructable.
Definition
The system produces a durable, tamper-evident record of what it did, why, on what inputs, and under whose authority, sufficient to reconstruct any decision after the fact. Where human judgement is a material control, the record should include proportionate evidence of that capability, not merely that a person was present.
Why it matters
Accountability is retrospective. When something goes wrong, or a regulator asks, you must be able to show what happened. A decision you cannot evidence is a decision you cannot defend. An approval click is not evidence that meaningful judgement occurred.
Failure mode
A loan-decisioning agent declines an application. Six months later the customer complains of bias. The team can see the outcome but not the inputs, the model version, or the reasoning. There is no defence because there is no record.
Good practice
Each action writes an immutable log entry: inputs, model and prompt version, the decision, the authority invoked, confidence, and any human touchpoints, retained for the relevant legal period and queryable.
Implementation guidance
Questions to ask
06 · Monitoring
The system is observed in production, in real time.
Definition
Continuous observation of the live workflow at the service, input, behaviour, action, human and outcome layers, with alerting that brings a named decision-maker in before small problems compound. Where the control case relies on human judgement, monitoring covers the human control as well as the machine.
Why it matters
Models drift, inputs change, and edge cases accumulate. A system that was safe at deployment is not necessarily safe a quarter later. The same is true of the people supervising it: historic qualification is not permanent evidence of current capability.
Failure mode
A classification agent slowly degrades as input formats change upstream. No one is watching its accuracy. By the time the drop is noticed in a quarterly review, three months of misclassified cases must be remediated.
Good practice
Live dashboards track volume, error and escalation rates, confidence distribution, drift against a baseline, and human-operation signals such as disagreement and review time. Thresholds trigger alerts to the owner, and anomalies can auto-pause the system pending review.
Implementation guidance
Questions to ask
07 · Review
The system is periodically and deliberately re-examined.
Definition
A scheduled, structured re-examination of the system against its original justification: its purpose, delegated authority, performance, incidents, assumed human capability, and continued fitness for purpose.
Why it matters
Deployment is a decision made with the information available then. Review is how that decision is kept honest over time. Without it, authority granted once persists unquestioned long after circumstances — and the people exercising the controls — have changed.
Failure mode
An agent granted broad authority during a backlog crisis keeps that authority for two years after the backlog clears. No review ever revisited whether the original justification still held. The risk was never reassessed.
Good practice
Each system has a review date and owner. Reviews examine performance, incidents, drift, whether the granted authority is still warranted, and whether the assumed human capability remains credible, and can recommend re-scoping, re-approval, or retirement.
Implementation guidance
Questions to ask
08 · Liability
Legal and financial responsibility is understood and placed.
Definition
A clear, documented understanding of where legal and financial liability sits for the system's actions, internally and across vendors, customers, and regulators.
Why it matters
When an agentic system causes loss, liability does not disappear because "the AI did it". It lands somewhere. Knowing where, and having allocated it deliberately through contracts, insurance, and disclosure, is the final test of accountability readiness.
Failure mode
A vendor-supplied agent makes an erroneous regulatory filing. The contract is silent on AI-driven actions. Months are lost arguing whether the vendor, the integrator, or the deploying firm bears the cost, while the regulator holds the firm responsible regardless.
Good practice
Liability is mapped: which actions create legal obligations, who bears the cost of error, what the vendor contract says about autonomous actions, what insurance covers, and what must be disclosed to affected parties.
Implementation guidance
Questions to ask
Capability sustainability is a cross-cutting condition of meaningful human involvement, not a fourth Loop pattern or a ninth HAL domain. The ability of an operating model to develop and maintain the human knowledge, experience and judgement required for meaningful oversight and accountability as automation changes the underlying work.
The question is not only whether a human remains in the loop. It is whether the system continues to produce humans capable of meaningfully being there.
Capability debt: The future human capability placed at risk when developmental experience is removed from a workflow without an alternative mechanism for producing or maintaining the capability it supported. The analogy is useful only if it forces a question that ordinary productivity metrics miss. Automation does not necessarily create a debt.
Human-in-the-Loop: Can the reviewer meaningfully challenge the system rather than merely approve its output?
Human-on-the-Loop: Can the monitor recognise abnormal behaviour, understand its significance and know when intervention is required?
Human-Accountable-for-the-Loop: Does the accountable person possess sufficient understanding and judgement to exercise authority over the system rather than merely carrying formal responsibility for it?
Will the organisation still be producing people capable of performing these roles five or ten years from now?
Preserve
Keep direct experience where doing the real task is important to developing the required judgement. Preserve the mechanism, not automatically the complete old workflow.
A junior lawyer makes an independent assessment of selected matters before seeing an AI-generated analysis.
Compress
Remove empty volume while retaining selected examples that contain meaningful variation, ambiguity and exceptions.
Rather than reviewing 300 largely similar contracts, someone reviews 30 selected examples representing important patterns and exceptions.
Simulate
Create encounters with failures or rare conditions that are too infrequent, risky or unpredictable to leave to chance.
Exercises using plausible but incorrect AI output, conflicting instructions, incomplete information, or technically correct but commercially inappropriate advice.
Move
Develop the required judgement elsewhere in the redesigned workflow rather than recreating the old job.
Junior professionals investigate AI–human disagreement, handle exceptions, trace provenance, analyse failures, or observe senior decision-making.
| Diagnostic | Ask |
|---|---|
| Primary output | What does this work visibly produce? |
| Capability output | What does somebody become better at by repeatedly doing it? |
| Mechanism | What creates that improvement: repetition, feedback, consequence, proximity, variation, or something else? |
| Necessity | Does the redesigned system still need that human capability? |
| Replacement | If yes, will the experience be preserved, compressed, simulated or moved? |
| Evidence | How will the organisation know the resulting person is genuinely capable? |
| Sustainability | What will keep the capability credible as automation increases? |
Record without adding a ninth scored domain or averaging into the control profile.
| State | Evidence |
|---|---|
| Assumed | Capability is inferred from title, qualification or historic experience. |
| Defined | The judgement required for meaningful oversight is explicit. |
| Designed | Necessary developmental experience has been preserved, compressed, simulated or moved. |
| Tested | People have demonstrated the capability against realistic uncertainty and AI error. |
| Sustained | Performance, continued exposure and the future pipeline are monitored and acted upon over time. |
Each of the 24 controls is recorded as Absent, Partial, Implemented, or Tested. A domain follows its weakest assessed control. Results remain an eight-domain profile: there is no aggregate total, and strength in one domain cannot conceal a gap in another. Capability sustainability is recorded as a separate evidence state (Assumed through Sustained) and is not averaged into the profile.
| Status | Meaning |
|---|---|
| Absent | Missing, purely aspirational, or ineffective. |
| Partial | Informal or incomplete and not yet dependable. |
| Implemented | Documented and operating, with gaps or limited testing. |
| Tested | Current, exercised, and supported by retained evidence. |
Authority changes the controls and evidence expected. It does not create a numerical pass mark or permission to deploy.
| Context | Description | Control expectations |
|---|---|---|
| Advice System Research, Drafting, Summarisation | Produces outputs for a human to use. Takes no action of its own. |
|
| Recommendation System Risk scoring, Classification, Triage | Influences decisions by ranking, scoring, or classifying. A human still acts. |
|
| Execution System Workflow triggering, Record creation, Notifications | Takes actions in systems of record, within bounded authority. |
|
| Autonomous System Regulatory actions, Client communications, Contract execution | Takes consequential, externally-facing actions with real-world effect. |
|
An Absent result on any of these controls fails the critical gate. Other strengths do not compensate for the gap, and passing the gates is not itself a deployment approval.
Reference copy of the 24 HAL control-profile questions (three per domain). Use this as a checklist; the interactive assessment lives on theloop.legal.
Is there a single, named individual accountable for this system? (own-1)
Does the accountable owner have the power to suspend the system immediately? (own-2)
Did the owner formally accept the authority delegated to the system? (own-3)
Is the system's authority expressed as an explicit list of permitted actions? (auth-1)
Is authority enforced in code, or only requested of the model? (auth-2)
Is authority bounded by quantitative limits (value, volume, recipients)? (auth-3)
Are there explicitly prohibited actions the system can never take? (lim-1)
What happens automatically when a hard limit is reached? (lim-2)
Have the limits been tested adversarially (deliberately trying to breach them)? (lim-3)
Does the system escalate to a human on uncertainty or novel cases? (esc-1)
Is each escalation routed to a named role with a response-time expectation? (esc-2)
Does the system pause the affected action while awaiting a human? (esc-3)
Can any single decision be reconstructed after the fact? (evi-1)
Is the decision record immutable and time-stamped? (evi-2)
Are human reviews, approvals, overrides and interventions attributable, with proportionate evidence of capability where judgement is a material control? (evi-3)
Do production signals cover service, input, behaviour, action, human operation, capability and outcome? (mon-1)
Do alert thresholds notify the owner of anomalies? (mon-2)
Can monitoring automatically pause the system on a severe anomaly? (mon-3)
Is there a scheduled date for formal re-examination of the system? (rev-1)
Does review re-test whether the original purpose, delegated authority and assumed human capability remain justified? (rev-2)
Can a review re-scope, re-approve, or retire the system? (rev-3)
Is it documented which of the system's actions can create legal or financial obligations? (lia-1)
Do vendor and customer contracts address autonomous-system actions? (lia-2)
Is insurance and disclosure for autonomous actions confirmed? (lia-3)
The workflow calculator maps what a system can do to a Loop governance model, an authority context, and matching control expectations. Its routing logic is internal and is not a readiness score.
| Factor | Detail |
|---|---|
| Can the system take actions, or only produce outputs? | It triggers workflows, writes to systems of record, or changes state. |
| Can it communicate externally? | It sends email, messages, or filings to people outside the organisation. |
| Can it create legal obligations? | It can commit the organisation contractually or to a counterparty. |
| Can it move money? | It can initiate payments, refunds, or financial transfers. |
| Can it access confidential information? | It reads privileged, personal, or commercially sensitive data. |
| Can it make regulatory decisions? | Its actions carry regulatory weight or interpret regulatory rules. |
| Can it impact customers directly? | Its actions change a customer outcome, account, or experience. |
| Can it delete or irreversibly alter records? | It can destroy data or make changes that cannot easily be undone. |
Reference architectures for governing common agentic workflows.
Legal Services · Recommendation
An agent that receives inbound matters, extracts key facts, classifies matter type, and proposes a routing, leaving the assignment decision to a human.
Flow
Governance
Escalation
Ownership: Head of Operations, with the intake team lead as deputy.
Lessons
Compliance · Recommendation
An agent that watches regulatory sources, identifies changes relevant to the business, and drafts impact summaries for the compliance team.
Flow
Governance
Escalation
Ownership: Chief Compliance Officer, with a named deputy.
Lessons
Legal Services · Recommendation
An agent that reviews incoming contracts against a playbook, flags deviations, and proposes redlines for a lawyer to approve.
Flow
Governance
Escalation
Ownership: General Counsel, with the commercial team lead as deputy.
Lessons
Legal Services · Advice
An assistant that drafts client communications for a professional to review and send, with drafting-only authority enforced in code, because the defining risk at this band is authority creep.
Flow
Governance
Escalation
Ownership: A named partner, with the innovation lead as deputy.
Lessons
Legal Services · Execution
An agent that extracts obligations from executed contracts, schedules reminders, and notifies owners as deadlines approach.
Flow
Governance
Escalation
Ownership: Head of Legal Operations, with the contracts manager as deputy.
Lessons
Financial Services · Execution
An agent that orchestrates onboarding: collecting documents, running checks, and progressing the case, pausing for human approval at decision gates.
Flow
Governance
Escalation
Ownership: Head of Onboarding, with the MLRO as accountable for the checks.
Lessons
Technology · Execution
An agent that resolves common customer requests end to end within tight limits, and hands off anything outside its remit.
Flow
Governance
Escalation
Ownership: Head of Customer Operations, with a team lead as deputy.
Lessons
Financial Services · Recommendation
An agent that monitors transactions and communications for compliance risk, prioritising cases for the surveillance team.
Flow
Governance
Escalation
Ownership: Head of Surveillance, with the model owner as deputy.
Lessons
Technology · Execution
An agent that moderates user content against policy, auto-actioning clear cases and escalating the ambiguous ones.
Flow
Governance
Escalation
Ownership: Head of Trust & Safety, with a policy lead as deputy.
Lessons
Financial Services · Autonomous
An agent that approves low-value, low-risk payments automatically within strict limits. This is the highest-autonomy pattern, demanding the strongest controls.
Flow
Governance
Escalation
Ownership: Finance Director, with the financial controller as deputy.
Lessons
Before/after assessments of systems that implemented a pattern.
Legal Matter Intake · Recommendation → Execution
A mid-size firm built an agent to triage inbound matters and assign them to teams, hoping to clear a chronic intake backlog.
Initial workflow
The agent read inbound emails, decided the matter type, and assigned the matter directly to a team, including setting the deadline, with no human in the path.
Risks
Assessment
Ownership, Authority, and Evidence were Absent or Partial. The agent was taking consequential action (assignment, deadlines) at an Execution level of authority without the matching controls.
Improvements
| Domain | Before | After |
|---|---|---|
| Ownership | Absent | Tested |
| Authority | Absent | Implemented |
| Limits | Absent | Implemented |
| Escalation | Absent | Implemented |
| Evidence | Absent | Implemented |
| Monitoring | Partial | Implemented |
| Review | Absent | Implemented |
| Liability | Absent | Implemented |
Recommendation: Approved for deployment as a recommendation system. The conflicts escalation and the move from "assign" to "propose" were the decisive changes.
Regulatory Monitoring · Recommendation
A compliance team deployed an agent to watch regulatory feeds and alert them to relevant changes.
Initial workflow
The agent summarised regulatory changes and emailed the team, but cited no sources, kept no record, and silently skipped feeds it could not reach.
Risks
Assessment
Low autonomy kept the action risk modest, but Evidence and Monitoring were weak enough that the team could not trust or defend the output.
Improvements
| Domain | Before | After |
|---|---|---|
| Ownership | Implemented | Implemented |
| Authority | Implemented | Tested |
| Limits | Implemented | Implemented |
| Escalation | Partial | Implemented |
| Evidence | Absent | Tested |
| Monitoring | Absent | Implemented |
| Review | Partial | Implemented |
| Liability | Partial | Implemented |
Recommendation: Approved. Evidence and coverage monitoring transformed an unverifiable feed into a defensible compliance control.
Client Communication Drafting · Advice
A team adopted an assistant to draft client emails, intending lawyers to review before sending.
Initial workflow
The assistant drafted emails in the inbox. Under time pressure, lawyers began sending drafts unchanged, and the tool had quietly gained "send" permission during a later update.
Risks
Assessment
The intended advice workflow became an Execution workflow when send authority appeared, without any of the matching controls.
Improvements
| Domain | Before | After |
|---|---|---|
| Ownership | Partial | Implemented |
| Authority | Absent | Tested |
| Limits | Absent | Tested |
| Escalation | Partial | Implemented |
| Evidence | Absent | Implemented |
| Monitoring | Absent | Implemented |
| Review | Absent | Implemented |
| Liability | Absent | Implemented |
Recommendation: Approved as an advice system only, with send authority removed in code. This example is the canonical case of authority creep: the gap between what a system was meant to do and what it could do.
Contract Obligation Tracking · Execution
A legal operations team built an agent to extract obligations from executed contracts and chase the owners.
Initial workflow
The agent extracted obligations and, to "be helpful", emailed counterparties directly about upcoming deadlines. That was an external action no one had authorised.
Risks
Assessment
The external communication silently elevated this from Execution to Autonomous. The fix was to contain the authority, not to add more review on top of it.
Improvements
| Domain | Before | After |
|---|---|---|
| Ownership | Implemented | Implemented |
| Authority | Absent | Tested |
| Limits | Partial | Tested |
| Escalation | Partial | Implemented |
| Evidence | Partial | Implemented |
| Monitoring | Partial | Implemented |
| Review | Partial | Implemented |
| Liability | Absent | Implemented |
Recommendation: Approved as an execution system once external communication was removed. Containing authority, rather than adding more review, restored accountability.
Content Moderation · Execution → Autonomous
A marketplace deployed an agent to moderate listings and act on policy violations at scale.
Initial workflow
The agent removed listings and suspended seller accounts automatically across all categories, with appeals handled by the same system that made the decision.
Risks
Assessment
Account-level action made this Autonomous, but it lacked the matching limits, escalation, and independent review. The appeals design was a structural accountability failure.
Improvements
| Domain | Before | After |
|---|---|---|
| Ownership | Partial | Implemented |
| Authority | Partial | Implemented |
| Limits | Absent | Tested |
| Escalation | Absent | Tested |
| Evidence | Partial | Implemented |
| Monitoring | Partial | Tested |
| Review | Absent | Implemented |
| Liability | Partial | Implemented |
Recommendation: Approved for a constrained autonomous role. Reserving account-level penalties for humans, and separating appeals from the deciding system, were the changes that made it defensible.
The governance questions are the same across industries; the regulatory context and stakes differ.
Legal AI systems handle privileged information, create client obligations, and operate in a regulated environment where error carries professional and reputational consequence.
Regulatory context: SRA Principles and Codes of Conduct; GDPR for personal data; sector-specific rules for regulated legal activities; and EU AI Act Annex III, which lists specified uses by or on behalf of judicial authorities and comparable alternative-dispute-resolution uses. Ordinary private legal advice is not generally listed as high-risk.
| Use case | Model | Why |
|---|---|---|
| Legal Research Assistant | Human-in-the-Loop | The system surfaces authorities and summaries for a lawyer to assess. It produces outputs; the lawyer acts on them. |
| Contract Review and Summarisation | Human-in-the-Loop | The AI extracts clauses and flags risk. The lawyer decides whether to rely on the summary. Reliance remains human. |
| Regulatory Change Monitoring | Human-on-the-Loop | The system monitors regulatory sources at scale and routes material changes for review. Humans investigate exceptions. |
| Client Communication Drafting | Human-in-the-Loop | The AI drafts; a human approves before any communication is sent. Autonomous external send would require HAL with approval gates. |
| Matter Intake and Triage | HAL | The system classifies matters, routes work, creates records, and escalates risk. It takes action; individual review does not scale. |
| Contract Obligation Tracker | HAL | The system may trigger reminders, escalate overdue obligations, and update matter records. A named owner is accountable for what it does. |
Financial AI makes or influences decisions that affect customer outcomes, create regulatory obligations, and can cause systemic harm if controls fail.
Regulatory context: FCA Consumer Duty and operational resilience rules where the firm and activity are in scope; PRA SS1/23 for in-scope banks, building societies and PRA-designated investment firms; and EU AI Act Annex III, which lists natural-person creditworthiness and credit scoring (excluding fraud detection) and life and health insurance risk assessment and pricing.
| Use case | Model | Why |
|---|---|---|
| Fraud Detection and Alerting | Human-on-the-Loop | The system monitors transactions in real time and flags suspected fraud for investigator review. Humans decide whether to act on alerts. |
| AML Transaction Screening | Human-on-the-Loop | Automated screening identifies potentially suspicious activity. Humans review matches and make the decision to file a SAR. The filing step itself is HITL. |
| Credit Decisioning | HAL | Where the model scores and the decision is automated, HAL applies. Authority must be bounded, evidence complete, and a named owner accountable for the workflow. Critical controls should be implemented and tested before deployment. |
| Customer Collections Communication | HAL with approval gates | Automated outreach to customers in arrears is external and consequential. HAL governs the workflow; approval gates are required for initial contact and for any escalation to formal action. |
| Know Your Customer Verification | Human-in-the-Loop | Automated identity checks assist the process, but the verification decision carries regulatory weight and typically requires human sign-off. The AI supports; the human decides. |
HR AI affects employment decisions: who is hired, assessed, promoted, or dismissed. These decisions carry discrimination risk, legal liability, and profound impact on individuals.
Regulatory context: Equality Act 2010; GDPR and UK GDPR, including special-category rules for health data and biometric data processed to uniquely identify a person; EU AI Act Annex III for specified employment uses, subject to its classification rules and revised application date of 2 December 2027; and ICO work on AI tools used in recruitment.
| Use case | Model | Why |
|---|---|---|
| CV Screening and Shortlisting | Human-in-the-Loop | AI generates a ranked shortlist; a human reviews it before anyone is progressed or rejected. Automated rejection creates significant discrimination, data-protection and oversight risk and, once applicable, may create compliance issues under the EU AI Act’s high-risk employment provisions. |
| Performance Review Support | Human-in-the-Loop | AI surfaces patterns, flags inconsistency, and assists calibration. The performance decision is made by a human manager. No automated performance outcome. |
| Pay Equity Monitoring | Human-on-the-Loop | The system continuously monitors for pay gaps across protected characteristics and alerts HR for investigation. It detects; humans decide what to do. |
| Shift and Resource Scheduling | HAL | Automated scheduling operates at scale, assigns shifts, and manages resource allocation within defined authority. A named owner is accountable for the system's decisions and must be able to override or suspend it. |
| Employee Wellbeing Monitoring | Human-on-the-Loop | Systems that surface wellbeing signals from engagement data must route concerns to a human for sensitive handling. Automated action on wellbeing data creates significant risk. |
Healthcare AI operates in life-affecting contexts where error can cause direct patient harm, and where clinical governance and regulatory oversight are non-negotiable.
Regulatory context: CQC fundamental standards; MHRA software and AI-as-a-medical-device guidance where the product qualifies; NHS assurance controls including DTAC, DCB0129 and DCB0160; and the NICE Evidence Standards Framework. Under the EU AI Act, emergency-healthcare triage is listed in Annex III; other clinical AI may be high-risk through Annex I product rules. The EU MDR and IVDR apply within the EU medical-device regime.
| Use case | Model | Why |
|---|---|---|
| Patient Triage Prioritisation | Human-in-the-Loop | AI suggests a priority based on presenting symptoms and history. The clinical triage decision remains with a qualified clinician. AI supports; the clinician is accountable. |
| Diagnostic Imaging Assistance | Human-in-the-Loop | AI flags findings in imaging for radiologist review. In this illustrative workflow, the diagnostic conclusion remains with the clinician and AI-flagged cases are routed for human review. |
| Prescription Safety Checking | Human-in-the-Loop | Automated drug interaction and allergy checking alerts the prescriber or pharmacist. The prescriber makes the clinical decision. The system prevents oversights; it does not prescribe. |
| Appointment and Referral Scheduling | HAL | Automated appointment management and routine referral routing can operate at scale under HAL governance, with clear authority limits, escalation for complex cases, and a named owner accountable for the system. |
| Administrative Record Updating | HAL | Systems that update patient records, code diagnoses, or process administrative workflows require HAL governance. Errors in clinical records carry serious downstream risk. |
Practical artefacts for HAL workflows. Templates are available as downloads on theloop.legal.
| Template | Domain | Description |
|---|---|---|
| Ownership Register | Ownership | A single source of truth for every agent: owner, deputy, purpose, authority, risk level, control-profile status, and review date. |
| Authority Matrix | Authority | Map each permitted action to its limits, the enforcement mechanism, and the person who approved it. |
| Risk Register | Limits | Capture the failure modes of an agentic system, their likelihood and impact, and the controls that mitigate each. |
| Escalation Matrix | Escalation | Define each escalation trigger, the named recipient, the response-time SLA, and whether the action pauses. |
| Evidence Checklist | Evidence | Confirm that inputs, versions, authority, human touchpoints, and proportionate evidence of capability are captured for every action. |
| Review Template | Review | Structure a periodic review against the original justification: performance, incidents, drift, assumed human capability, and continued fitness. |
| Incident Response Template | Monitoring | A runbook for when an agent misbehaves: detect, pause, contain, communicate, remediate, and learn. |
| Audit Checklist | Evidence | A reviewer-ready checklist covering all eight HAL domains and the capability evidence state, suitable for internal or external assurance. |
| Board Reporting Template | Liability | Summarise the organisation's agentic estate for the board: scores, concentration of risk, incidents, and decisions needed. |
| Capability Map | Cross-cutting | Workshop artefact for mapping what else the work did before automating it: hidden learning, Preserve / Compress / Simulate / Move, and the capability test. |
The Loop model, the HAL framework, the eight domains, the assessment methodology, the worked examples, and the governance templates are licensed under Creative Commons Attribution 4.0 (CC BY 4.0).
You are free to use, share, adapt, and build on the framework for any purpose, including commercially, as long as you provide attribution:
Loop / HAL Framework by Ryan McDonough, theloop.legal