Skip to content

HAL · Capability sustainability

What else did the work do?

An accountable AI operating model cannot depend indefinitely on human judgement while simultaneously removing the experiences through which that judgement develops. This is a system-design question, not a training programme.

The question is not only whether a human remains in the loop. It is whether the system continues to produce humans capable of meaningfully being there.

This is not a fourth Loop pattern and not a ninth HAL domain. It is a cross-cutting condition of meaningful human involvement. For Human-in-the-Loop, Human-on-the-Loop and Human-Accountable-for-the-Loop alike, the operating model assumes a capable human. The method below asks whether that assumption will still be true after the work is redesigned.

It is also not an argument against automation, and not an argument for preserving repetitive junior work because previous generations had to do it. Traditional professional work often combined production and development in the same activity. AI allows those functions to separate. Organisations therefore need to identify which developmental effects remain important and deliberately reproduce, relocate or redesign them.

Two outputs of professional work

Work product is not the only product

For years, professional processes produced both work and workers capable of doing more difficult work later. We measured the first because the second happened gradually in the background. AI allows us to separate them.

Output A

Work product

The visible organisational output.

  • Contract reviewed
  • Advice produced
  • Code written
  • Account reconciled
  • Research conducted

Output B

Human capability

The less visible output created through participation.

  • Pattern recognition and exception judgement
  • Understanding of normality
  • Source scepticism
  • Escalation judgement
  • Calibration of confidence

Capability debt is the future human capability placed at risk when developmental experience is removed from a workflow without an alternative mechanism for producing or maintaining the capability it supported. The analogy to technical debt is useful only if it forces a question that ordinary productivity metrics miss. Automation does not necessarily create a debt; some repetitive work adds little, and well-designed AI use can accelerate feedback. The point is that organisations currently have poor visibility of whether it is happening.

The assumption underneath every Loop pattern

The human involved is capable of the judgement assigned to them

Human-in-the-Loop
Can the reviewer meaningfully challenge the system rather than merely approve its output?
Human-on-the-Loop
Can the monitor recognise abnormal behaviour, understand its significance and know when intervention is required?
Human-Accountable-for-the-Loop
Does the accountable person possess sufficient understanding and judgement to exercise authority over the system rather than merely carrying formal responsibility for it?

Will the organisation still be producing people capable of performing these roles five or ten years from now?

The method

The capability preservation cycle

Map the workflow twice: once for what it produces, and once for what participation develops. Then decide deliberately what happens to the capability, test whether the replacement works, and keep watching it. These are not levels and they do not need an acronym.

  1. 01

    Map the work

  2. 02

    Map the hidden learning

  3. 03

    Preserve / Compress / Simulate / Move

  4. 04

    Test the capability

  5. 05

    Monitor capability over time

Map the hidden learning

After the ordinary transformation questions — what happens, who decides, what AI can perform — map the same work a second time. The last question is the forcing question. Existing work does not earn preservation merely by being old.

01

What does repetition teach?

Which patterns, variations and exceptions become recognisable only through repeated exposure?

Separate useful variation from unnecessary volume. Repetition is not automatically development.

02

Where does judgement develop?

Where do rules stop being sufficient and people interpret ambiguity, context or competing considerations?

Identify the points at which an allow-list or playbook cannot settle the decision.

03

Where does feedback happen?

How does someone discover that their original judgement was wrong, and who explains why?

A correction without reasoning may change the next click without forming later judgement.

04

What does real responsibility add?

Does deciding on live work create learning that a simulation cannot completely reproduce?

Consider uncertainty, consequences, time pressure and responsibility to other people.

05

What does proximity teach?

What is learnt by seeing experienced practitioners handle disagreement, incidents, trade-offs and recovery?

This includes client conversations, escalation, and watching how seniors decide when the answer is not obvious.

06

What capability exists afterwards?

What can somebody recognise, understand or decide after two years in the work that they could not do when they started?

If the honest answer is nothing the redesigned operating model needs, that supports automation.

Decide what happens to the capability

The useful question is where the learning happens in the new system, not how to recreate the old job. A single role may need several of these responses.

Preserve

Keep direct experience where doing the real task is important to developing the required judgement. Preserve the mechanism, not automatically the complete old workflow.

Example. A junior lawyer makes an independent assessment of selected matters before seeing an AI-generated analysis.

Compress

Remove empty volume while retaining selected examples that contain meaningful variation, ambiguity and exceptions.

Example. Rather than reviewing 300 largely similar contracts, someone reviews 30 selected examples representing important patterns and exceptions.

Simulate

Create encounters with failures or rare conditions that are too infrequent, risky or unpredictable to leave to chance.

Example. Exercises using plausible but incorrect AI output, conflicting instructions, incomplete information, or technically correct but commercially inappropriate advice.

Move

Develop the required judgement elsewhere in the redesigned workflow rather than recreating the old job.

Example. Junior professionals investigate AI–human disagreement, handle exceptions, trace provenance, analyse failures, or observe senior decision-making.

Before automating a task

Use this as a workshop artefact. Complete it for each human role the redesigned workflow changes. Download the worksheet if you want a copy you can fill in.

Diagnostic Ask
Primary output What does this work visibly produce?
Capability output What does somebody become better at by repeatedly doing it?
Mechanism What creates that improvement: repetition, feedback, consequence, proximity, variation, or something else?
Necessity Does the redesigned system still need that human capability?
Replacement If yes, will the experience be preserved, compressed, simulated or moved?
Evidence How will the organisation know the resulting person is genuinely capable?
Sustainability What will keep the capability credible as automation increases?
Download the capability map

Test the capability

Do not allow the method to end at “provide training”. If the organisation believes it has replaced the developmental effects of traditional work, it should test whether the intended capability actually exists. Independent manual performance is not an absolute requirement in every case. The deeper requirement is sufficient independent capability to exercise the assigned oversight or accountability role.

  • ? Can this person perform or independently reason about the relevant underlying task without starting from the AI answer?
  • ? Do they understand the system’s limitations and the conditions in which its output should not be trusted?
  • ? Can they explain why they agree and why they disagree?
  • ? Can they identify when available information is insufficient?
  • ? Have they demonstrated appropriate escalation under uncertainty?
  • ? Are they exposed to enough meaningful variation and feedback to maintain their judgement?
  • ? How will new people acquire equivalent capability?
  • ? What evidence would indicate that capability is degrading?

Test with unfamiliar scenarios, plausible but incorrect outputs, incomplete information, conflicting sources, ambiguous cases, novel exceptions and realistic time pressure. The objective is not only whether they spotted a planted error. Did they know why it mattered? Did they recognise uncertainty? Did they know what to do next? Could they challenge an apparently authoritative machine output?

Evidence that capability will last

Record an evidence state without adding a ninth HAL domain or calculating another score. Do not average it with the eight-domain control profile. Where Review, Monitor or Own relies on human judgement, an unsupported assumption weakens Ownership, Evidence, Monitoring and Review even if policies and system controls are present.

State Evidence What it means
Assumed Capability is inferred from title, qualification or historic experience. The control case has not been tested. Treat this as a material assumption, not evidence.
Defined The judgement required for meaningful oversight is explicit. The organisation knows what the role must be able to do, but not whether anyone can still do it.
Designed Necessary developmental experience has been preserved, compressed, simulated or moved. The replacement is a hypothesis. It remains unproven until people demonstrate the judgement.
Tested People have demonstrated the capability against realistic uncertainty and AI error. Current role-holders can exercise the assigned judgement. The pipeline and durability are still open questions.
Sustained Performance, continued exposure and the future pipeline are monitored and acted upon over time. Capability is treated as an operating-model condition, not a one-off training event.

Competence is not capacity

Competence asks whether the person can make the judgement required. Capacity asks whether they can give it enough time and attention under actual workload. A reviewer can be highly competent but overloaded, available but incapable, both or neither. Automation may increase capacity by removing volume while weakening competence if people no longer practise the parts of the work on which oversight depends. Neither effect should be assumed; both should be tested.

Existing competence

Can this person exercise the required judgement today?

Maintained competence

Does the role provide enough continued exposure, feedback and practice for that judgement to remain credible?

Future competence

Does the operating model continue to produce people able to perform the role later?

Where does the capable reviewer come from?

A reviewer who starts from an AI answer is not in the same cognitive position as someone producing an independent answer. Finding the consequential 2% in a draft that is 98% correct may require more judgement than producing the original first draft, not less.

  1. 01 AI does the work because humans can review it.
  2. 02 Humans develop the judgement required to review it through experience of the work.
  3. 03 Humans receive less experience of the work because AI does it.

Where does the capable reviewer eventually come from?

Signals worth investigating

These do not automatically prove capability loss. A low disagreement rate may show that the system improved. Each is a reason to investigate, compare with outcomes, and test the human role directly.

  • ! Reviewers almost never disagree with AI, particularly as model use expands.
  • ! Escalation and exception rates collapse without a corresponding change in outcomes.
  • ! Review times become implausibly short.
  • ! Reviewers cannot explain why an output is acceptable or why an error matters.
  • ! People repeatedly miss planted or naturally occurring exceptions.
  • ! Confidence remains high while independent performance falls.
  • ! Exposure to variation, unusual cases and consequences declines.
  • ! The pipeline of people eligible for consequential review narrows over time.

Definitions

Capability sustainability
The ability of an operating model to develop and maintain the human knowledge, experience and judgement required for meaningful oversight and accountability as automation changes the underlying work.
Capability debt
The future human capability placed at risk when developmental experience is removed from a workflow without an alternative mechanism for producing or maintaining the capability it supported.

Keep returning to the operating-model question: if your AI workflow relies on human judgement, where does that judgement come from and what keeps it credible?