Loop HAL · Complete reference · Version 1.1

Human governance of AI systems

Ryan McDonough · theloop.legal · Static reference (no interactive tools)

Loop names three patterns for how humans govern AI: review (Human-in-the-Loop), monitor (Human-on-the-Loop), and own (Human-Accountable-for-the-Loop). HAL is the practical accountability framework for the third pattern — workflows where individual review does not scale.

Accountability cannot be delegated. Execution can.

1. About Loop & HAL

“Human-in-the-Loop” is one of the most repeated phrases in AI governance. The problem is that people use it to mean very different things: reviewing outputs, monitoring systems, or remaining accountable for outcomes. Those are not the same governance model.

Loop gives clearer language. HAL does not replace Human-in-the-Loop; it is the detailed accountability framework for agentic systems, multi-agent orchestration, large-scale triage, and autonomous operational systems — where individual review does not scale.

HAL complements standards organisations already follow. The NIST AI Risk Management Framework (AI RMF 1.0) supports managing risks to people, organisations and society. ISO/IEC 42001:2023 specifies requirements for an AI management system. The EU AI Act creates legal obligations for actors and systems within its scope. Loop identifies which governance pattern applies; HAL examines whether delegated action remains bounded, evidenced and accountable.

Principles

2. The three Loop models

Human-in-the-Loop made most sense when AI systems produced outputs for humans to use. As systems gained autonomy — taking actions, routing work, creating records — the same phrase was stretched to cover monitoring, process design, and accountability.

The established taxonomy runs in-the-loop → on-the-loop → out-of-the-loop. Loop deliberately replaces that third term. “Out-of-the-loop” frames the human as absent. The human is never out of the loop of accountability, only out of the loop of execution. That is why the third pattern is Human-Accountable-for-the-Loop.

Review

Human-in-the-Loop

A human reviews or approves individual AI outputs before action is taken.

AI → Human Review → Action

Human role: Reviewer · Did a human review the decision?

Can the reviewer meaningfully challenge the system rather than merely approve its output?

An AI drafts a clause summary. A lawyer reviews it before relying on it.

Best for

  • Drafting
  • Summarisation
  • Legal research
  • Low-volume triage
  • Internal recommendations
  • Assistant-style tools

Strengths

  • Clear human decision point
  • Easy to explain
  • Familiar governance pattern
  • Strong fit for legal work where judgement remains central

Limitations

  • Does not scale well to high-volume workflows
  • Can become performative if review is rushed
  • May create false comfort if the reviewer lacks context or independent judgement
  • Assumes a capable reviewer exists and will remain available
  • Slows down workflows where automation is the point

Monitor

Human-on-the-Loop

An AI system operates within defined boundaries while humans monitor behaviour, investigate exceptions and intervene when required.

AI → Action ↓ Human Monitoring

Human role: Monitor · Is a human monitoring the system?

Can the monitor recognise abnormal behaviour, understand its significance and know when intervention is required?

A regulatory monitoring system identifies potential changes and routes unusual or high-impact items for review.

Best for

  • Alerting systems
  • Risk monitoring
  • Operational dashboards
  • Compliance screening
  • Exception-based workflows
  • Medium-volume automation

Strengths

  • Better suited to scale than Human-in-the-Loop
  • Allows automation while preserving oversight
  • Supports exception-based intervention
  • Useful where normal operations are predictable

Limitations

  • Requires strong monitoring
  • Humans may miss weak signals
  • Intervention thresholds must be well designed
  • Accountability can become unclear if ownership is not defined
  • Historic experience is not permanent evidence that monitors remain ready

Own

Human-Accountable-for-the-Loop (HAL)

A human owns the design, authority, controls, evidence and outcomes of an AI-enabled workflow, even where the AI system acts without individual human review.

Owner → Policy → AI Workflow → Action → Evidence → Review

Human role: Owner · Who owns the system that made the decision?

Does the accountable person possess sufficient understanding and judgement to exercise authority over the system rather than merely carrying formal responsibility for it?

A matter triage agent classifies incoming legal requests, routes work, creates records, escalates risk and logs evidence. No human reviews every decision, but a named owner is accountable for the workflow, its controls and its outcomes.

Best for

  • Agentic systems
  • Multi-agent workflows
  • Workflow orchestration
  • Large-scale triage
  • Obligation tracking
  • Autonomous operational systems
  • High-volume decision workflows

Strengths

  • Designed for scale
  • Focuses on real accountability
  • Supports agentic workflows
  • Forces clarity on ownership, authority, limits, evidence and escalation

Limitations

  • Requires mature governance
  • Not suitable where accountability is vague
  • Needs technical logging and review processes
  • Higher setup burden than basic review workflows
  • Accountability without capability can become ceremonial

3. When each model fits

Workflow typeRecommended model
Drafting a memoHuman-in-the-Loop
Summarising a contractHuman-in-the-Loop
Monitoring regulatory updatesHuman-on-the-Loop
Screening incoming mattersHuman-on-the-Loop or HAL
Drafting a client communicationHuman-in-the-Loop
Routing client requestsHAL
Updating matter recordsHAL
Triggering external communicationsHAL with approval gates or strict controls
Filing regulatory documentsHAL only with high score and approval gates

Decision guide

  1. Does the AI system only generate outputs for human use?
    If yes → Human-in-the-Loop likely fits.
  2. Does the AI system operate continuously or at scale, but humans intervene when needed?
    If yes → Human-on-the-Loop likely fits.
  3. Does the AI system take action, trigger workflows, create records, route work or affect outcomes without review of every action?
    If yes → Human-Accountable-for-the-Loop likely fits.
  4. Could failure create legal, financial, regulatory, reputational or client impact?
    If yes → Assess against HAL before deployment.

Illustrative worked examples (model mapping)

ExampleModelWhy
Legal Research AssistantHuman-in-the-LoopThe AI generates outputs for a lawyer to review. It does not act independently.
Contract Summary ToolHuman-in-the-LoopThe human still decides whether to rely on the summary.
Regulatory Monitoring SystemHuman-on-the-LoopThe system monitors at scale and escalates relevant changes.
Legal Matter Triage AgentHALThe system classifies, routes and creates operational consequences.
Contract Obligation TrackerHALThe system may trigger reminders, escalate risk and affect deadline management.
Client Email Response AgentHAL with approval gatesExternal communication creates higher risk. HAL may apply, but some actions may still require Human-in-the-Loop approval.

4. Who Is Actually in the Loop?

Long-form paper · Version 1.1 · ≈ 20 min read

For most of the history of machine decision-making, governance had a simple shape: a machine proposed, and a human disposed. We called this Human-in-the-Loop, and it served us well. That model is now under strain, because the premise it rested on no longer holds: that the human is the one who acts.

Terms, used precisely

Responsibility: doing the work. Can be delegated, including to a system.

Accountability: being answerable for the outcome. In law, this may sit with the organisation or with particular office-holders. In governance, HAL requires it to be traceable to a named accountable owner and not displaced onto software.

Liability: bearing legal or financial consequences. The economic burden of liability can often be allocated by contract, indemnity and insurance, but statutory, regulatory and third-party liability may remain where the law places it.

When HAL says accountability can never be delegated, execution can, it means exactly this: responsibility for execution can move. Accountability for the system cannot.

01 · The rise of Human-in-the-Loop

Human-in-the-Loop (HITL) did not begin with machine learning. It is an inheritance from control theory, aviation, and early automation, where a human operator remained the final authority over a machine that could otherwise act on its own. The principle was conservative and sound: where a machine might err in ways that matter, insert a human between its judgement and the world.

When predictive models entered high-stakes settings (credit, medicine, hiring), HITL was the natural control. The model scored; a human decided. Oversight and action lived in the same place: the person.

02 · Why the phrase became comforting

HITL worked because of a structural fact about the systems it governed: they did not act. They produced outputs (scores, classifications, drafts, recommendations) and then stopped. The output was inert until a human picked it up. That gap between output and action was where governance lived.

It also worked because the volumes were human-scaled. A loan officer could review the applications in front of them. A clinician could weigh a model’s reading against their own. The number of decisions was bounded by human capacity, and so the review was real.

03 · Why it breaks down at scale

Two things break HITL: volume and action. When a system makes more decisions than a human can examine, “human-in-the-loop” becomes “human-rubber-stamping-the-loop”. The control persists on paper while evaporating in practice.

A reviewer who cannot review is not a control. They are a liability with a job title.

The deeper break comes when systems begin to act. The moment a system can send the email, file the report, move the money, or delete the record, the gap between output and action closes. There is no longer a natural pause in which a human stands. To insert one artificially, by requiring approval of every action, is to throw away the very capability you deployed the system for.

04 · Review, monitoring and accountability

These are three distinct governance patterns, not stages of maturity. Human-in-the-Loop asks whether a human reviewed the decision. Human-on-the-Loop asks whether a human is monitoring the system. Human-Accountable-for-the-Loop asks who owns the system that made the decision.

This distinction also sits more comfortably with how modern regulation describes oversight. Article 14 of the EU AI Act requires high-risk AI systems to be designed and developed so they can be effectively overseen by natural persons during use. Depending on context and the oversight measures identified by the provider, those persons need the competence, authority and tools to understand limitations, monitor anomalies, guard against automation bias, interpret or disregard outputs, and interrupt the system. The Act therefore treats human oversight as a set of functional controls, not a single undifferentiated requirement. Under the amended timetable, the high-risk rules apply to Annex III systems from 2 December 2027 and to systems covered through Annex I product legislation from 2 August 2028.

05 · Why agents create a new governance problem

Agentic systems take actions in pursuit of goals, often chaining many steps, calling tools, and operating with minimal supervision. Once invoked or authorised, they can proceed across steps without returning to a human for each decision. That shift changes what governance must cover. An actor needs a different kind of governance: whether the action was authorised, bounded, recorded, and owned.

06 · The Loop model

The point of Loop is not to rank these patterns from weak to strong. A human-in-the-loop model may be exactly right for a legal research assistant. A human-on-the-loop model may be right for a monitoring system. HAL may be necessary where the system acts at scale and individual review is no longer meaningful. The mistake is not choosing HITL. The mistake is using HITL as a generic label when the actual control is review, monitoring, ownership, or some unstable mixture of all three.

07 · Why HAL matters

Execution can be delegated to a system. Accountability cannot. When a human employee acts within their authority, the organisation remains accountable for what they do. Nothing about substituting software for the employee changes that. The organisation, and a named person within it, remains answerable.

HAL fixes accountability to a person who owns the system. Software has no legal personality and no institutional accountability of its own. It cannot be sanctioned, dismissed, disciplined or struck off. So accountability flows back, through the system, to the human who owns it. That is the entire substance of Human Accountable for the Loop.

08 · Where the reviewer comes from

There is an assumption sitting underneath all three Loop patterns: the human involved is capable of exercising the judgement the operating model assigns to them. Accountability without capability can become ceremonial.

Capability is not a static resource. For years, professional processes produced two outputs: the visible work product, and people capable of handling harder work later. AI allows those outputs to separate. That is not an argument against automation, and not an argument for preserving repetitive junior work. Organisations need to identify which developmental effects remain important and deliberately preserve, compress, simulate or move them.

The question is not only whether a human remains in the loop. It is whether the system continues to produce humans capable of meaningfully being there.

This does not create a fourth Loop pattern or a ninth HAL domain. It changes what the existing disciplines must prove. Historic qualification is not permanent evidence of current capability.

09 · Implications for legal AI

The law of agency is not directly about software, but it is useful because it shows how law has long dealt with delegated action: a principal is bound by the authorised acts of their agent, and an agent who exceeds their authority creates liability that does not simply vanish.

One precision matters: software is not an agent in the legal sense. It has no legal personality, owes no fiduciary duty, and no agency relationship arises when you deploy it. In law, an AI system is a tool, and the consequences of a tool’s operation land directly on the organisation that wields it. The parallel HAL draws is structural, not doctrinal: the mechanics of sound delegation (scoped authority, limits, escalation, records) transfer; the legal relationship does not.

A regulator faced with an erroneous automated filing will not accept “the model did it” as a defence. HAL asks organisations to confront that reality before deployment, rather than discover it during an incident.

10 · Conclusion

Organisations will run many agents at once. Governing this estate will resemble portfolio management more than decision-by-decision review: a registry of agents, each with an owner, an authority scope, a risk level, a HAL control profile, and a review date. Governance itself will become continuous. And throughout, one line will not be allowed to blur: however deep the stack of agents, accountability terminates at a human. The future of AI governance is about ensuring that, when systems act, someone remains accountable for the system that acted, and that the organisation can still produce people capable of exercising that accountability.

5. The eight HAL domains

HAL evaluates whether a workflow is ready to delegate action while retaining clear human accountability. Eight domains form the framework.

01 · Ownership

Ownership

A named person owns the system and its outcomes.

Definition

A single, named individual is accountable for the agentic system: one person, not a diffuse committee or "the platform". Ownership is the anchor on which every other domain depends.

Why it matters

Accountability cannot be delegated to software. When a system acts, someone must be answerable for what it did and did not do. Diffuse ownership is the most common root cause of governance failure: when everyone is responsible, no one is.

Failure mode

An automated client-communication agent sends an incorrect legal deadline to 400 clients. The incident review finds the agent was "owned by the AI working group". No individual can explain its authority, and no one has the standing to switch it off.

Good practice

Each deployed agent has a named accountable owner recorded in a register, with a deputy. The owner has signed off on the authority granted, reviews incidents, and holds the documented power to suspend the system immediately.

Implementation guidance

Questions to ask

02 · Authority

Authority

The system acts only within explicitly delegated authority.

Definition

The scope of what the system is permitted to do is defined explicitly and granted deliberately, mirroring how authority is delegated to a human employee.

Why it matters

Agentic systems act. An action taken outside granted authority is an unauthorised act for which the organisation remains liable. Authority must be designed in before deployment.

Failure mode

A procurement agent built to "draft" purchase orders is also able to submit them because no one constrained the integration. It commits the company to £2.3m of spend before anyone notices the boundary was never set.

Good practice

Authority is expressed as an explicit allow-list of actions, value thresholds, and counterparties. Anything outside the list is impossible by construction rather than discouraged by a prompt.

Implementation guidance

Questions to ask

03 · Limits

Limits

Hard constraints stop the system before it causes harm.

Definition

Beyond authority, limits are the non-negotiable boundaries: the actions the system must never take, and the conditions under which it must stop, regardless of confidence.

Why it matters

Authority says what is allowed; limits say what is forbidden and where the floor is. Limits are what protect you when the model is confidently wrong. They convert "the model decided not to" into "the system could not".

Failure mode

A collections agent is permitted to contact customers. With no limit on frequency, a logic loop causes it to email one vulnerable customer 71 times in a day. There was authority to contact; there was no limit on contact.

Good practice

Hard limits (maximum spend, prohibited actions, rate caps, protected data classes, and "never act on this customer segment") are enforced as guard rails that halt the system and escalate, independent of the model.

Implementation guidance

Questions to ask

04 · Escalation

Escalation

Uncertainty and edge cases route to a human cleanly.

Definition

Defined paths by which the system hands control to a human when it is uncertain, when it hits a limit, or when a case falls outside its competence.

Why it matters

A system that never escalates is a system that has been told to guess. The quality of governance is often the quality of its escalation: clear triggers, a named recipient, and a defined response time.

Failure mode

A triage agent encounters a matter type it has never seen. With no escalation path, it forces the case into the nearest category and routes it incorrectly. The deadline is missed; no human was ever asked.

Good practice

Escalation triggers are explicit (low confidence, novel case, limit reached, high stakes). Each routes to a named human or role with a defined SLA, and the system pauses the relevant action until a human responds.

Implementation guidance

Questions to ask

05 · Evidence

Evidence

Every decision and action is recorded and reconstructable.

Definition

The system produces a durable, tamper-evident record of what it did, why, on what inputs, and under whose authority, sufficient to reconstruct any decision after the fact. Where human judgement is a material control, the record should include proportionate evidence of that capability, not merely that a person was present.

Why it matters

Accountability is retrospective. When something goes wrong, or a regulator asks, you must be able to show what happened. A decision you cannot evidence is a decision you cannot defend. An approval click is not evidence that meaningful judgement occurred.

Failure mode

A loan-decisioning agent declines an application. Six months later the customer complains of bias. The team can see the outcome but not the inputs, the model version, or the reasoning. There is no defence because there is no record.

Good practice

Each action writes an immutable log entry: inputs, model and prompt version, the decision, the authority invoked, confidence, and any human touchpoints, retained for the relevant legal period and queryable.

Implementation guidance

Questions to ask

06 · Monitoring

Monitoring

The system is observed in production, in real time.

Definition

Continuous observation of the live workflow at the service, input, behaviour, action, human and outcome layers, with alerting that brings a named decision-maker in before small problems compound. Where the control case relies on human judgement, monitoring covers the human control as well as the machine.

Why it matters

Models drift, inputs change, and edge cases accumulate. A system that was safe at deployment is not necessarily safe a quarter later. The same is true of the people supervising it: historic qualification is not permanent evidence of current capability.

Failure mode

A classification agent slowly degrades as input formats change upstream. No one is watching its accuracy. By the time the drop is noticed in a quarterly review, three months of misclassified cases must be remediated.

Good practice

Live dashboards track volume, error and escalation rates, confidence distribution, drift against a baseline, and human-operation signals such as disagreement and review time. Thresholds trigger alerts to the owner, and anomalies can auto-pause the system pending review.

Implementation guidance

Questions to ask

07 · Review

Review

The system is periodically and deliberately re-examined.

Definition

A scheduled, structured re-examination of the system against its original justification: its purpose, delegated authority, performance, incidents, assumed human capability, and continued fitness for purpose.

Why it matters

Deployment is a decision made with the information available then. Review is how that decision is kept honest over time. Without it, authority granted once persists unquestioned long after circumstances — and the people exercising the controls — have changed.

Failure mode

An agent granted broad authority during a backlog crisis keeps that authority for two years after the backlog clears. No review ever revisited whether the original justification still held. The risk was never reassessed.

Good practice

Each system has a review date and owner. Reviews examine performance, incidents, drift, whether the granted authority is still warranted, and whether the assumed human capability remains credible, and can recommend re-scoping, re-approval, or retirement.

Implementation guidance

Questions to ask

08 · Liability

Liability

Legal and financial responsibility is understood and placed.

Definition

A clear, documented understanding of where legal and financial liability sits for the system's actions, internally and across vendors, customers, and regulators.

Why it matters

When an agentic system causes loss, liability does not disappear because "the AI did it". It lands somewhere. Knowing where, and having allocated it deliberately through contracts, insurance, and disclosure, is the final test of accountability readiness.

Failure mode

A vendor-supplied agent makes an erroneous regulatory filing. The contract is silent on AI-driven actions. Months are lost arguing whether the vendor, the integrator, or the deploying firm bears the cost, while the regulator holds the firm responsible regardless.

Good practice

Liability is mapped: which actions create legal obligations, who bears the cost of error, what the vendor contract says about autonomous actions, what insurance covers, and what must be disclosed to affected parties.

Implementation guidance

Questions to ask

6. What else did the work do?

Capability sustainability is a cross-cutting condition of meaningful human involvement, not a fourth Loop pattern or a ninth HAL domain. The ability of an operating model to develop and maintain the human knowledge, experience and judgement required for meaningful oversight and accountability as automation changes the underlying work.

The question is not only whether a human remains in the loop. It is whether the system continues to produce humans capable of meaningfully being there.

Capability debt: The future human capability placed at risk when developmental experience is removed from a workflow without an alternative mechanism for producing or maintaining the capability it supported. The analogy is useful only if it forces a question that ordinary productivity metrics miss. Automation does not necessarily create a debt.

Human-in-the-Loop: Can the reviewer meaningfully challenge the system rather than merely approve its output?

Human-on-the-Loop: Can the monitor recognise abnormal behaviour, understand its significance and know when intervention is required?

Human-Accountable-for-the-Loop: Does the accountable person possess sufficient understanding and judgement to exercise authority over the system rather than merely carrying formal responsibility for it?

Will the organisation still be producing people capable of performing these roles five or ten years from now?

The capability preservation cycle

  1. Map the work
  2. Map the hidden learning
  3. Preserve / Compress / Simulate / Move
  4. Test the capability
  5. Monitor capability over time

Map the hidden learning

  1. What does repetition teach? Which patterns, variations and exceptions become recognisable only through repeated exposure? Separate useful variation from unnecessary volume. Repetition is not automatically development.
  2. Where does judgement develop? Where do rules stop being sufficient and people interpret ambiguity, context or competing considerations? Identify the points at which an allow-list or playbook cannot settle the decision.
  3. Where does feedback happen? How does someone discover that their original judgement was wrong, and who explains why? A correction without reasoning may change the next click without forming later judgement.
  4. What does real responsibility add? Does deciding on live work create learning that a simulation cannot completely reproduce? Consider uncertainty, consequences, time pressure and responsibility to other people.
  5. What does proximity teach? What is learnt by seeing experienced practitioners handle disagreement, incidents, trade-offs and recovery? This includes client conversations, escalation, and watching how seniors decide when the answer is not obvious.
  6. What capability exists afterwards? What can somebody recognise, understand or decide after two years in the work that they could not do when they started? If the honest answer is nothing the redesigned operating model needs, that supports automation.

Preserve / Compress / Simulate / Move

Preserve

Keep direct experience where doing the real task is important to developing the required judgement. Preserve the mechanism, not automatically the complete old workflow.

A junior lawyer makes an independent assessment of selected matters before seeing an AI-generated analysis.

Compress

Remove empty volume while retaining selected examples that contain meaningful variation, ambiguity and exceptions.

Rather than reviewing 300 largely similar contracts, someone reviews 30 selected examples representing important patterns and exceptions.

Simulate

Create encounters with failures or rare conditions that are too infrequent, risky or unpredictable to leave to chance.

Exercises using plausible but incorrect AI output, conflicting instructions, incomplete information, or technically correct but commercially inappropriate advice.

Move

Develop the required judgement elsewhere in the redesigned workflow rather than recreating the old job.

Junior professionals investigate AI–human disagreement, handle exceptions, trace provenance, analyse failures, or observe senior decision-making.

Before automating a task

DiagnosticAsk
Primary outputWhat does this work visibly produce?
Capability outputWhat does somebody become better at by repeatedly doing it?
MechanismWhat creates that improvement: repetition, feedback, consequence, proximity, variation, or something else?
NecessityDoes the redesigned system still need that human capability?
ReplacementIf yes, will the experience be preserved, compressed, simulated or moved?
EvidenceHow will the organisation know the resulting person is genuinely capable?
SustainabilityWhat will keep the capability credible as automation increases?

Capability evidence states

Record without adding a ninth scored domain or averaging into the control profile.

StateEvidence
AssumedCapability is inferred from title, qualification or historic experience.
DefinedThe judgement required for meaningful oversight is explicit.
DesignedNecessary developmental experience has been preserved, compressed, simulated or moved.
TestedPeople have demonstrated the capability against realistic uncertainty and AI error.
SustainedPerformance, continued exposure and the future pipeline are monitored and acted upon over time.

7. Control-profile methodology

Each of the 24 controls is recorded as Absent, Partial, Implemented, or Tested. A domain follows its weakest assessed control. Results remain an eight-domain profile: there is no aggregate total, and strength in one domain cannot conceal a gap in another. Capability sustainability is recorded as a separate evidence state (Assumed through Sustained) and is not averaged into the profile.

Evidence states

StatusMeaning
AbsentMissing, purely aspirational, or ineffective.
PartialInformal or incomplete and not yet dependable.
ImplementedDocumented and operating, with gaps or limited testing.
TestedCurrent, exercised, and supported by retained evidence.

Authority context

Authority changes the controls and evidence expected. It does not create a numerical pass mark or permission to deploy.

ContextDescriptionControl expectations
Advice System
Research, Drafting, Summarisation
Produces outputs for a human to use. Takes no action of its own.
  • Enforce the drafting or advisory boundary.
  • Preserve material sources and outputs.
  • Make the human decision point real and attributable.
Recommendation System
Risk scoring, Classification, Triage
Influences decisions by ranking, scoring, or classifying. A human still acts.
  • Test how recommendations influence decisions and affected groups.
  • Support meaningful challenge rather than ceremonial review.
  • Provide escalation, appeal, and correction routes.
Execution System
Workflow triggering, Record creation, Notifications
Takes actions in systems of record, within bounded authority.
  • Enforce authority through identities, permissions, and workflow state.
  • Apply effective limits and hold exceptions safely.
  • Reconstruct actions and test timely containment.
Autonomous System
Regulatory actions, Client communications, Contract execution
Takes consequential, externally-facing actions with real-world effect.
  • Apply all execution controls with narrower reserved actions and tighter aggregate limits.
  • Provide rapid or automatic containment where safe and justified.
  • Require independent challenge and workable remedy.

Critical (non-compensable) controls

An Absent result on any of these controls fails the critical gate. Other strengths do not compensate for the gap, and passing the gates is not itself a deployment approval.

8. Assessment questions

Reference copy of the 24 HAL control-profile questions (three per domain). Use this as a checklist; the interactive assessment lives on theloop.legal.

Ownership

Is there a single, named individual accountable for this system? (own-1)

Does the accountable owner have the power to suspend the system immediately? (own-2)

Did the owner formally accept the authority delegated to the system? (own-3)

Authority

Is the system's authority expressed as an explicit list of permitted actions? (auth-1)

Is authority enforced in code, or only requested of the model? (auth-2)

Is authority bounded by quantitative limits (value, volume, recipients)? (auth-3)

Limits

Are there explicitly prohibited actions the system can never take? (lim-1)

What happens automatically when a hard limit is reached? (lim-2)

Have the limits been tested adversarially (deliberately trying to breach them)? (lim-3)

Escalation

Does the system escalate to a human on uncertainty or novel cases? (esc-1)

Is each escalation routed to a named role with a response-time expectation? (esc-2)

Does the system pause the affected action while awaiting a human? (esc-3)

Evidence

Can any single decision be reconstructed after the fact? (evi-1)

Is the decision record immutable and time-stamped? (evi-2)

Are human reviews, approvals, overrides and interventions attributable, with proportionate evidence of capability where judgement is a material control? (evi-3)

Monitoring

Do production signals cover service, input, behaviour, action, human operation, capability and outcome? (mon-1)

Do alert thresholds notify the owner of anomalies? (mon-2)

Can monitoring automatically pause the system on a severe anomaly? (mon-3)

Review

Is there a scheduled date for formal re-examination of the system? (rev-1)

Does review re-test whether the original purpose, delegated authority and assumed human capability remain justified? (rev-2)

Can a review re-scope, re-approve, or retire the system? (rev-3)

Liability

Is it documented which of the system's actions can create legal or financial obligations? (lia-1)

Do vendor and customer contracts address autonomous-system actions? (lia-2)

Is insurance and disclosure for autonomous actions confirmed? (lia-3)

9. Workflow risk factors

The workflow calculator maps what a system can do to a Loop governance model, an authority context, and matching control expectations. Its routing logic is internal and is not a readiness score.

FactorDetail
Can the system take actions, or only produce outputs?It triggers workflows, writes to systems of record, or changes state.
Can it communicate externally?It sends email, messages, or filings to people outside the organisation.
Can it create legal obligations?It can commit the organisation contractually or to a counterparty.
Can it move money?It can initiate payments, refunds, or financial transfers.
Can it access confidential information?It reads privileged, personal, or commercially sensitive data.
Can it make regulatory decisions?Its actions carry regulatory weight or interpret regulatory rules.
Can it impact customers directly?Its actions change a customer outcome, account, or experience.
Can it delete or irreversibly alter records?It can destroy data or make changes that cannot easily be undone.

10. Reference patterns

Reference architectures for governing common agentic workflows.

Compliance · Recommendation

Regulatory Monitoring

An agent that watches regulatory sources, identifies changes relevant to the business, and drafts impact summaries for the compliance team.

Flow

  1. Agent polls regulatory feeds and publications
  2. Agent filters for relevance to the firm
  3. Agent drafts a plain-language impact summary
  4. Agent tags affected policies and owners
  5. Compliance officer validates and actions

Governance

Escalation

Ownership: Chief Compliance Officer, with a named deputy.

Lessons

Legal Services · Recommendation

Contract Review

An agent that reviews incoming contracts against a playbook, flags deviations, and proposes redlines for a lawyer to approve.

Flow

  1. Contract uploaded for first-pass review
  2. Agent compares clauses against the playbook
  3. Agent flags deviations and missing protections
  4. Agent drafts proposed redlines with rationale
  5. Lawyer reviews, edits, and accepts

Governance

Escalation

Ownership: General Counsel, with the commercial team lead as deputy.

Lessons

Legal Services · Advice

Client Communication Drafting

An assistant that drafts client communications for a professional to review and send, with drafting-only authority enforced in code, because the defining risk at this band is authority creep.

Flow

  1. Professional requests a draft, or the assistant proposes one
  2. Assistant drafts the communication, citing sources for factual claims
  3. Draft is tagged as AI-generated in the workflow
  4. Professional reviews, edits, and sends
  5. Sent communication retained alongside its draft history

Governance

Escalation

Ownership: A named partner, with the innovation lead as deputy.

Lessons

Legal Services · Execution

Contract Obligation Tracking

An agent that extracts obligations from executed contracts, schedules reminders, and notifies owners as deadlines approach.

Flow

  1. Executed contract ingested
  2. Agent extracts obligations, dates, and owners
  3. Agent creates tracked tasks with reminders
  4. Agent notifies owners ahead of deadlines
  5. Owner confirms or reassigns each obligation

Governance

Escalation

Ownership: Head of Legal Operations, with the contracts manager as deputy.

Lessons

Financial Services · Execution

Client Onboarding

An agent that orchestrates onboarding: collecting documents, running checks, and progressing the case, pausing for human approval at decision gates.

Flow

  1. New client initiates onboarding
  2. Agent requests and validates required documents
  3. Agent runs KYC/AML checks via integrations
  4. Agent assembles the case for approval
  5. Compliance approves before account activation

Governance

Escalation

Ownership: Head of Onboarding, with the MLRO as accountable for the checks.

Lessons

Technology · Execution

Customer Service

An agent that resolves common customer requests end to end within tight limits, and hands off anything outside its remit.

Flow

  1. Customer raises a request
  2. Agent identifies intent and eligibility
  3. Agent resolves within permitted action set
  4. Agent confirms resolution to the customer
  5. Anything out of scope is handed to a human agent

Governance

Escalation

Ownership: Head of Customer Operations, with a team lead as deputy.

Lessons

Financial Services · Recommendation

Compliance Monitoring

An agent that monitors transactions and communications for compliance risk, prioritising cases for the surveillance team.

Flow

  1. Agent ingests transactions and communications
  2. Agent scores cases against risk indicators
  3. Agent prioritises and explains each alert
  4. Surveillance analyst investigates
  5. Outcomes feed back to tune the model

Governance

Escalation

Ownership: Head of Surveillance, with the model owner as deputy.

Lessons

Technology · Execution

Content Moderation

An agent that moderates user content against policy, auto-actioning clear cases and escalating the ambiguous ones.

Flow

  1. Content submitted or reported
  2. Agent classifies against the policy taxonomy
  3. Clear violations actioned automatically
  4. Borderline cases queued for human review
  5. Appeals routed to a separate human track

Governance

Escalation

Ownership: Head of Trust & Safety, with a policy lead as deputy.

Lessons

Financial Services · Autonomous

Financial Approval

An agent that approves low-value, low-risk payments automatically within strict limits. This is the highest-autonomy pattern, demanding the strongest controls.

Flow

  1. Payment request received
  2. Agent validates against policy and limits
  3. Within-limit, low-risk payments approved automatically
  4. Anything else escalated for human approval
  5. All approvals reconciled and reviewed daily

Governance

Escalation

Ownership: Finance Director, with the financial controller as deputy.

Lessons

11. Worked examples

Before/after assessments of systems that implemented a pattern.

Regulatory Monitoring · Recommendation

Regulatory Change Monitoring

A compliance team deployed an agent to watch regulatory feeds and alert them to relevant changes.

Initial workflow

The agent summarised regulatory changes and emailed the team, but cited no sources, kept no record, and silently skipped feeds it could not reach.

Risks

Assessment

Low autonomy kept the action risk modest, but Evidence and Monitoring were weak enough that the team could not trust or defend the output.

Improvements

DomainBeforeAfter
OwnershipImplementedImplemented
AuthorityImplementedTested
LimitsImplementedImplemented
EscalationPartialImplemented
EvidenceAbsentTested
MonitoringAbsentImplemented
ReviewPartialImplemented
LiabilityPartialImplemented

Recommendation: Approved. Evidence and coverage monitoring transformed an unverifiable feed into a defensible compliance control.

Client Communication Drafting · Advice

Client Email Drafting Assistant

A team adopted an assistant to draft client emails, intending lawyers to review before sending.

Initial workflow

The assistant drafted emails in the inbox. Under time pressure, lawyers began sending drafts unchanged, and the tool had quietly gained "send" permission during a later update.

Risks

Assessment

The intended advice workflow became an Execution workflow when send authority appeared, without any of the matching controls.

Improvements

DomainBeforeAfter
OwnershipPartialImplemented
AuthorityAbsentTested
LimitsAbsentTested
EscalationPartialImplemented
EvidenceAbsentImplemented
MonitoringAbsentImplemented
ReviewAbsentImplemented
LiabilityAbsentImplemented

Recommendation: Approved as an advice system only, with send authority removed in code. This example is the canonical case of authority creep: the gap between what a system was meant to do and what it could do.

Contract Obligation Tracking · Execution

Contract Obligation Tracker

A legal operations team built an agent to extract obligations from executed contracts and chase the owners.

Initial workflow

The agent extracted obligations and, to "be helpful", emailed counterparties directly about upcoming deadlines. That was an external action no one had authorised.

Risks

Assessment

The external communication silently elevated this from Execution to Autonomous. The fix was to contain the authority, not to add more review on top of it.

Improvements

DomainBeforeAfter
OwnershipImplementedImplemented
AuthorityAbsentTested
LimitsPartialTested
EscalationPartialImplemented
EvidencePartialImplemented
MonitoringPartialImplemented
ReviewPartialImplemented
LiabilityAbsentImplemented

Recommendation: Approved as an execution system once external communication was removed. Containing authority, rather than adding more review, restored accountability.

Content Moderation · Execution → Autonomous

Marketplace Moderation Agent

A marketplace deployed an agent to moderate listings and act on policy violations at scale.

Initial workflow

The agent removed listings and suspended seller accounts automatically across all categories, with appeals handled by the same system that made the decision.

Risks

Assessment

Account-level action made this Autonomous, but it lacked the matching limits, escalation, and independent review. The appeals design was a structural accountability failure.

Improvements

DomainBeforeAfter
OwnershipPartialImplemented
AuthorityPartialImplemented
LimitsAbsentTested
EscalationAbsentTested
EvidencePartialImplemented
MonitoringPartialTested
ReviewAbsentImplemented
LiabilityPartialImplemented

Recommendation: Approved for a constrained autonomous role. Reserving account-level penalties for humans, and separating appeals from the deciding system, were the changes that made it defensible.

12. Sector examples

The governance questions are the same across industries; the regulatory context and stakes differ.

Financial Services

Financial AI makes or influences decisions that affect customer outcomes, create regulatory obligations, and can cause systemic harm if controls fail.

Regulatory context: FCA Consumer Duty and operational resilience rules where the firm and activity are in scope; PRA SS1/23 for in-scope banks, building societies and PRA-designated investment firms; and EU AI Act Annex III, which lists natural-person creditworthiness and credit scoring (excluding fraud detection) and life and health insurance risk assessment and pricing.

Use caseModelWhy
Fraud Detection and AlertingHuman-on-the-LoopThe system monitors transactions in real time and flags suspected fraud for investigator review. Humans decide whether to act on alerts.
AML Transaction ScreeningHuman-on-the-LoopAutomated screening identifies potentially suspicious activity. Humans review matches and make the decision to file a SAR. The filing step itself is HITL.
Credit DecisioningHALWhere the model scores and the decision is automated, HAL applies. Authority must be bounded, evidence complete, and a named owner accountable for the workflow. Critical controls should be implemented and tested before deployment.
Customer Collections CommunicationHAL with approval gatesAutomated outreach to customers in arrears is external and consequential. HAL governs the workflow; approval gates are required for initial contact and for any escalation to formal action.
Know Your Customer VerificationHuman-in-the-LoopAutomated identity checks assist the process, but the verification decision carries regulatory weight and typically requires human sign-off. The AI supports; the human decides.

Human Resources

HR AI affects employment decisions: who is hired, assessed, promoted, or dismissed. These decisions carry discrimination risk, legal liability, and profound impact on individuals.

Regulatory context: Equality Act 2010; GDPR and UK GDPR, including special-category rules for health data and biometric data processed to uniquely identify a person; EU AI Act Annex III for specified employment uses, subject to its classification rules and revised application date of 2 December 2027; and ICO work on AI tools used in recruitment.

Use caseModelWhy
CV Screening and ShortlistingHuman-in-the-LoopAI generates a ranked shortlist; a human reviews it before anyone is progressed or rejected. Automated rejection creates significant discrimination, data-protection and oversight risk and, once applicable, may create compliance issues under the EU AI Act’s high-risk employment provisions.
Performance Review SupportHuman-in-the-LoopAI surfaces patterns, flags inconsistency, and assists calibration. The performance decision is made by a human manager. No automated performance outcome.
Pay Equity MonitoringHuman-on-the-LoopThe system continuously monitors for pay gaps across protected characteristics and alerts HR for investigation. It detects; humans decide what to do.
Shift and Resource SchedulingHALAutomated scheduling operates at scale, assigns shifts, and manages resource allocation within defined authority. A named owner is accountable for the system's decisions and must be able to override or suspend it.
Employee Wellbeing MonitoringHuman-on-the-LoopSystems that surface wellbeing signals from engagement data must route concerns to a human for sensitive handling. Automated action on wellbeing data creates significant risk.

Healthcare

Healthcare AI operates in life-affecting contexts where error can cause direct patient harm, and where clinical governance and regulatory oversight are non-negotiable.

Regulatory context: CQC fundamental standards; MHRA software and AI-as-a-medical-device guidance where the product qualifies; NHS assurance controls including DTAC, DCB0129 and DCB0160; and the NICE Evidence Standards Framework. Under the EU AI Act, emergency-healthcare triage is listed in Annex III; other clinical AI may be high-risk through Annex I product rules. The EU MDR and IVDR apply within the EU medical-device regime.

Use caseModelWhy
Patient Triage PrioritisationHuman-in-the-LoopAI suggests a priority based on presenting symptoms and history. The clinical triage decision remains with a qualified clinician. AI supports; the clinician is accountable.
Diagnostic Imaging AssistanceHuman-in-the-LoopAI flags findings in imaging for radiologist review. In this illustrative workflow, the diagnostic conclusion remains with the clinician and AI-flagged cases are routed for human review.
Prescription Safety CheckingHuman-in-the-LoopAutomated drug interaction and allergy checking alerts the prescriber or pharmacist. The prescriber makes the clinical decision. The system prevents oversights; it does not prescribe.
Appointment and Referral SchedulingHALAutomated appointment management and routine referral routing can operate at scale under HAL governance, with clear authority limits, escalation for complex cases, and a named owner accountable for the system.
Administrative Record UpdatingHALSystems that update patient records, code diagnoses, or process administrative workflows require HAL governance. Errors in clinical records carry serious downstream risk.

13. Governance toolkit

Practical artefacts for HAL workflows. Templates are available as downloads on theloop.legal.

TemplateDomainDescription
Ownership RegisterOwnershipA single source of truth for every agent: owner, deputy, purpose, authority, risk level, control-profile status, and review date.
Authority MatrixAuthorityMap each permitted action to its limits, the enforcement mechanism, and the person who approved it.
Risk RegisterLimitsCapture the failure modes of an agentic system, their likelihood and impact, and the controls that mitigate each.
Escalation MatrixEscalationDefine each escalation trigger, the named recipient, the response-time SLA, and whether the action pauses.
Evidence ChecklistEvidenceConfirm that inputs, versions, authority, human touchpoints, and proportionate evidence of capability are captured for every action.
Review TemplateReviewStructure a periodic review against the original justification: performance, incidents, drift, assumed human capability, and continued fitness.
Incident Response TemplateMonitoringA runbook for when an agent misbehaves: detect, pause, contain, communicate, remediate, and learn.
Audit ChecklistEvidenceA reviewer-ready checklist covering all eight HAL domains and the capability evidence state, suitable for internal or external assurance.
Board Reporting TemplateLiabilitySummarise the organisation's agentic estate for the board: scores, concentration of risk, incidents, and decisions needed.
Capability MapCross-cuttingWorkshop artefact for mapping what else the work did before automating it: hidden learning, Preserve / Compress / Simulate / Move, and the capability test.

14. Licence

The Loop model, the HAL framework, the eight domains, the assessment methodology, the worked examples, and the governance templates are licensed under Creative Commons Attribution 4.0 (CC BY 4.0).

You are free to use, share, adapt, and build on the framework for any purpose, including commercially, as long as you provide attribution:

Loop / HAL Framework by Ryan McDonough, theloop.legal