Chapter 51 — The System that Learns to Survive

An introduction to Block 6.

Explore the Layers
Chapters

Move to another chapter in the book.

  1. Chapter 1 — The Record
  2. Chapter 2 — Forty-Five Geologists
  3. Chapter 3 — Behind the Chronicle
  4. Chapter 4 — Thursday
  5. Chapter 5 — Robert
  6. Chapter 6 — Janette
  7. Chapter 7 — The Melbourne Process
  8. Chapter 8 — Patrick Smith
  9. Chapter 9 — Canada
  10. Chapter 10 — When the Patient Meets the Record
  11. Chapter 11 — The Life That Followed
  12. Chapter 12 — Work, Roads and Country
  13. Chapter 13 — Building a Life
  14. Chapter 14 — What Memory Did With It
  15. Chapter 15 — When the Past Began Returning
  16. Chapter 16 — Making Connections
  17. Chapter 17 — Following the Records
  18. Chapter 18 — The Cost of Being Disbelieved
  19. Chapter 19 — Building the Evidence
  20. Chapter 20 — Reclaiming the Record
  21. Chapter 21 — Germaine
  22. Chapter 22 — Herbie
  23. Chapter 23 — Through the Glass
  24. Chapter 24 — What Happened to Herbie's Story
  25. Chapter 25 — The Adults Around Yea
  26. Chapter 26 — 2007
  27. Chapter 27 — A Diagnosis That Travelled
  28. Chapter 28 — Finding Julian Lim
  29. Chapter 29 — When Meditation Opened the Wrong Door
  30. Chapter 30 — The Tests
  31. Chapter 31 — When Hell Becomes a Threat
  32. Chapter 32 — The Child of Satan
  33. Chapter 33 — What Adults Called an Exorcism
  34. Chapter 34 — The God I Was Told About
  35. Chapter 35 — Another Idea of God
  36. Chapter 36 — Teaching Fear to Children
  37. Chapter 37 — What Children Are Taught Now
  38. Chapter 38 — When Religion Becomes Abuse
  39. Chapter 39 — Australia and the Child's Freedom of Thought
  40. Chapter 40 — Personal Sovereignty
  41. Chapter 41 — What the Creator Forgot to Tell Us
  42. Chapter 42 — When Mythology Comes Before Measurement
  43. Chapter 43 — Who Gets to Speak for Authority?
  44. Chapter 44 — When the Observer Becomes Part of the Event
  45. Chapter 45 — The Iatrogenic Loop
  46. Chapter 46 — When the Record Becomes More Powerful Than the Person
  47. Chapter 47 — Secular in Name
  48. Chapter 48 — Democracy, Representation and Who Holds Power
  49. Chapter 49 — Systems That Cannot Admit Error
  50. Chapter 50 — Authority Must Remain Answerable
  51. Chapter 51 — The System that Learns to Survive
  52. Chapter 52 — The Easy Correction
  53. Chapter 53 — When Reasonable People Disagree
  54. Chapter 54 — When the Same Thing Keeps Happening
  55. Chapter 55 — When Correction Becomes Costly
  56. Chapter 56 — What the Tests Found
All sections (10)

Browse the complete chapter.

Two ways through this investigation ↑

Chapters 51–55 show in detail how the investigation was developed, with Chapters 52–55 progressively testing what happens as correction becomes more difficult.

Readers who want the resulting model without following every stage of that investigation can go directly to Chapter 56, where the four tests, their principal findings and the model derived from them are brought together in a shorter and more accessible form.

The detailed chapters remain available for anyone who wants to examine how a particular conclusion was reached, what evidence supports it, or where uncertainty remains.

Go directly to Chapter 56 — What the Tests Found →

The system that learns to survive ↑

This block began with a question about artificial intelligence.

It was a fairly simple question. I wanted to understand more about the risks and possible safeguards involved when artificial intelligence begins contributing to its own development, rather than its development remaining entirely controlled or supervised by humans.

That led very quickly to recursive self-improvement — the possibility that a system might improve not only what it can do, but its capacity to make the next improvement.

The question did not remain about machines for very long.

Once I began thinking about recursive improvement, I found myself looking at something much older.

People develop. Children learn how to learn. Institutions change their procedures. Hospitals respond to adverse events. Governments revise policies. Religions change teachings. Organisations introduce safeguards after failures. Societies attempt to learn from their histories.

In different ways, all of these systems can take information produced by what happened before and use it to influence what happens next.

But that immediately introduces a more difficult question.

What counts as improvement — and who gets to decide?

A system can become extraordinarily effective at pursuing an objective without ever examining whether the objective itself is desirable.

A child can become increasingly compliant while becoming less able to exercise autonomy.

An institution can become increasingly efficient at processing complaints without becoming better at recognising when it is wrong.

An artificial intelligence can become highly responsive to correction by the people operating it while the institution directing those people remains resistant to correction from those affected by its decisions.

The stated goal and the goal the system actually climbs are not necessarily the same thing.

That is where this investigation begins.

Someone Says No ↑

There is a much simpler way of expressing the problem.

Someone says:

“You've got this wrong.”

What happens next?

It is a remarkably powerful test.

Does the system hear them?

Does it preserve what they actually said?

Does somebody investigate?

Can evidence contrary to the original decision change the conclusion?

Can the person challenge not only the decision, but the institution's description of why that decision was made?

If the institution discovers an error, does it correct only the immediate problem, or does the correction travel through other records and decisions that depended upon it?

If somebody has been harmed, what happens to them?

Does anybody check later to determine whether the correction worked?

And if substantially the same problem occurs again and again, when does the system stop describing each occurrence as an isolated event and begin examining itself?

These questions take us beyond whether an institution has a complaints department.

They concern whether the institution is corrigible.

By corrigibility I mean something fairly straightforward: the capacity of a person, organisation or system to recognise and respond meaningfully to correction.

That does not mean accepting every challenge as correct.

A person can be mistaken. A complainant can misunderstand what happened. Evidence can be incomplete or contradictory. An institution may investigate a complaint fairly and still conclude that its original decision was justified.

Corrigibility therefore cannot mean that the person who says “You've got this wrong” automatically wins.

It means that saying it creates a genuine possibility that the system might discover that they are right.

Correction Cannot Travel in Only One Direction ↑

There is another problem.

Many systems are extremely good at correcting the people within them.

A school corrects the child. A prison corrects the prisoner. A hospital corrects the patient who does not follow instructions. A bureaucracy tells the citizen which procedure must be followed. A religious institution may tell the believer what must be believed or how behaviour must change. An organisation corrects its employees.

Artificial intelligence researchers quite reasonably ask how an AI system can remain corrigible to human oversight.

But there is another direction to the question.

Can the supervisor be corrected by the supervised?

Can the child expose something wrong in the institution?

Can the patient correct the clinical record?

Can the citizen demonstrate that the procedure itself produces an absurd result?

Can the employee expose a systemic failure?

Can a person affected by an AI-assisted decision challenge not merely the output, but the assumptions and classifications that produced it?

If correction operates only down the hierarchy, we have something quite different from a genuinely corrigible system.

We have one-way corrigibility.

That may produce compliance.

It does not necessarily produce improvement.

Four Tests ↑

It would be quite easy to design a theory of institutional correction that works beautifully when everyone behaves reasonably, the evidence is clear and nobody has anything important to lose.

That would not tell us very much.

So rather than beginning this block by constructing an elaborate model and then looking for examples that support it, I want to do something different.

We will develop the model and then deliberately try to break it.

For the moment, I have four tests.

Test One — The Easy Correction

Someone identifies a reasonably clear error.

There is good evidence showing what went wrong. Nobody has a substantial personal or institutional interest in preserving the mistake.

This gives us our baseline.

If an institution cannot correct itself under these circumstances, there is little reason to expect it to do so when correction becomes difficult.

But this test also gives us something positive to measure.

What does good correction actually look like?

How quickly was the concern acknowledged? Was the original account preserved? Was the error corrected? Were other records affected by it identified? Was the person told what had happened? Was any harm repaired? Did the institution learn anything that reduced the chance of the same error happening again?

And at what point can everybody reasonably say: this has been corrected?

Test Two — Genuine Disagreement

This test is much harder.

The person says the institution is wrong.

The institution has credible reasons for believing it is right.

The evidence may be incomplete. Different witnesses may remember events differently. Records may support more than one interpretation. Experts may legitimately disagree.

This test is important because dissent must not become another form of unquestionable authority.

The complainant is not necessarily right because they complained.

The institution is not necessarily right because it possesses the records, expertise or institutional authority.

The question becomes whether the process can preserve disagreement long enough to examine it fairly.

Can contradictory evidence survive?

Can uncertainty remain uncertainty rather than being converted into institutional certainty merely because a decision has to be recorded?

Can both parties understand how the eventual conclusion was reached?

And can new evidence reopen that conclusion?

A corrigible system needs to be capable of saying:

“On the evidence presently available, this is our conclusion.”

That is different from saying:

“This is what happened.”

Sometimes that distinction may matter enormously.

Test Three — The Recurring Pattern

The third test begins with apparently unrelated cases.

One person complains.

The matter is investigated and closed.

Another person complains somewhere else.

That matter is also investigated and closed.

Perhaps hundreds of cases accumulate over years, each sitting neatly inside its own file.

Every individual case may have been administratively completed.

The system may nevertheless have failed to learn.

This is where artificial intelligence and large-scale sense-making become particularly interesting.

AI may be capable of examining quantities of information that no individual investigator could reasonably hold together: complaints, outcomes, locations, classifications, demographic patterns, time intervals, policy changes and recurrence.

Are these really isolated events?

Perhaps complaints repeatedly begin with the same trigger.

Perhaps they are repeatedly converted into the same administrative category.

Perhaps the same attempted remedy repeatedly fails.

Perhaps particular groups experience the problem disproportionately.

Perhaps an organisation announces that it has learned from an event, while substantially the same event continues occurring.

AI need not decide who is right.

Its value may lie in making contradiction visible.

The institution says the problem has been corrected.

The data says it continues.

Those two accounts do not reconcile.

That discrepancy itself becomes information requiring investigation.

Test Four — The Power-Gradient Complaint

The fourth test may be the most demanding.

Imagine that an older person raises a serious complaint about failure to observe fundamental human-rights principles in their treatment within a hospital.

The complaint concerns a person who has served the institution for many years.

They may be highly respected. Their achievements may be substantial. Their relationships may reach throughout the organisation. They may have carved what appears to be an almost unassailable place in the history of the hospital.

We make no assumption that the complaint is true.

We make no assumption that it is false.

That is precisely the point of the test.

Can the institution determine what happened with the same intellectual independence it would bring to a complaint concerning someone with little institutional standing?

A distinguished history is evidence about a person's history.

It is not evidence that a particular allegation against that person is false.

Likewise, the seriousness of an allegation is not evidence that it is true.

But something else has changed.

Power has entered the equation.

The institution may now have something to lose from correction.

Its reputation may be involved. Senior relationships may be involved. Previous decisions may be called into question. People asked to investigate may know, respect or have worked with the person concerned.

The complainant, meanwhile, may be old, unwell, dependent upon the institution for care, unfamiliar with its processes and vastly outnumbered by people who already know one another.

That gives us another proposition to test:

As institutional interest in the outcome increases, the required independence of review should increase with it.

If the opposite happens — if increasing institutional power produces decreasing scrutiny — we have discovered something important about the system.

The question is no longer simply whether it has a complaints procedure.

It is whether the procedure can level the power imbalance sufficiently for the evidence to matter.

An Existing Example: Ryan's Rule ↑

We do not have to imagine every part of this architecture from nothing.

Queensland's Ryan's Rule provides an interesting example from clinical care.

It recognises a simple possibility: the people presently responsible for a patient's care may not have correctly recognised deterioration, and a patient, family member or carer needs a pathway for escalating that concern.

The importance for this investigation is not that Ryan's Rule solves the problems we are examining.

It doesn't. Its purpose and boundaries are much narrower.

Its importance is the principle underneath it:

The ordinary authority may be wrong, and the person affected needs a pathway around it.

That leads to a question we will return to:

What is the equivalent of Ryan's Rule when what is deteriorating is not only a person's clinical condition, but their rights, autonomy, safety or capacity to be heard?

Perhaps some of those mechanisms already exist.

Perhaps they exist but do not connect.

Perhaps some work extremely well.

Perhaps important gaps remain.

Those are matters to investigate rather than assume.

The Person Who Has Already Been Harmed ↑

There is another danger in everything I have written so far.

We could become so interested in designing a system that learns from failure that we overlook the person who paid for that learning.

Imagine an institution discovers that somebody was harmed.

It investigates.

It identifies the systemic cause.

It changes its policy.

Staff are retrained.

AI-assisted monitoring detects recurrence.

Future incidents decline dramatically.

From the institution's perspective this may look like successful learning.

But what happened to the person who was harmed?

They must not become merely the training data through which an institution learns to treat the next person better.

Preventing recurrence does not repair the person already harmed.

And compensating or caring for one harmed person does not necessarily correct the system that harmed them.

Those are different responsibilities.

A corrigible system therefore needs to ask at least two questions:

What do we owe the person who has already been harmed?

What must we change so that others are less likely to experience the same harm?

Care may include immediate protection, appropriate healthcare, practical assistance, restoration of services or opportunities, correction of records, acknowledgement, independent support, apology or compensation where warranted.

But care itself must not become another exercise of institutional power.

The person should, as far as reasonably possible, retain agency in determining what assistance is useful to them.

There is also another quantity we may be able to measure.

How much additional harm does a person experience merely trying to obtain recognition and remedy for the original harm?

I have provisionally called this remedy burden.

If a person must spend years navigating agencies, repeatedly recount what happened, produce enormous quantities of documentation and endure repeated dismissal before an easily verifiable error is finally corrected, the eventual administrative notation “resolved” tells us almost nothing about how the system actually performed.

Can Correction Itself Be Measured? ↑

That question may eventually become one of the most important in this block.

We already have some possibilities.

Dissent fidelity: how accurately does the system preserve what the person originally said?

Classification drift: does a specific allegation gradually become something different as it travels through an organisation?

Correction propagation: when an originating error is corrected, are downstream records and decisions that depended upon it also found and reviewed?

Correction half-life: how long does a known error continue influencing the system after it has supposedly been corrected?

Learning persistence: once an organisation changes because of a failure, how long does that change survive?

Recurrence: does substantially the same failure continue happening?

Remedy burden: what must the harmed person endure to obtain recognition, correction and appropriate care?

And now another possible measure:

Power gradient: does the pathway of correction change according to the institutional status of the person, department or interest being challenged?

These are provisional ideas.

I do not yet want to turn them into a score.

A numerical index introduced too early could create exactly the problem this investigation is attempting to expose: the measurement becomes the objective, and organisations learn how to improve the number rather than the thing the number was intended to represent.

First we need to discover whether these things can actually be measured, whether they tell us anything useful and how easily they can mislead us.

Rights Are Constraints, Not Objectives ↑

There is one more boundary we need before proceeding.

If we eventually use AI and large-scale analysis to help societies identify objectives, measure outcomes and improve institutions, we cannot simply ask what produces the greatest aggregate benefit and optimise toward it.

A system can improve its averages while repeatedly harming the same minority.

Ninety-nine people doing better does not automatically establish that what happened to the hundredth person was acceptable.

This is why I have come to think of rights as constraints rather than objectives.

Objectives describe what we are trying to achieve.

Rights establish things we are not entitled to sacrifice merely because doing so makes achievement of the objective easier.

Evidence tells us what actually happened.

Dissent tells us where our description of what happened may be incomplete.

And corrigibility determines whether the system can respond when those accounts collide.

The dissenter is therefore not merely an inconvenience to an otherwise functioning system.

Nor does dissent establish that the dissenter is right.

Dissent is information about the system.

Sometimes the investigation will establish that the system was right.

Sometimes it will establish that the dissenter was right.

Sometimes the evidence will not permit certainty.

But if the person cannot meaningfully challenge both the outcome and the system's explanation of that outcome, then the system's claim to corrigibility becomes difficult to sustain.

Where Artificial Intelligence Fits ↑

This investigation began with AI, and AI remains inside it.

But I do not want AI sitting at the top of this structure deciding what correction means.

Its more interesting role may be elsewhere.

It can preserve provenance.

It can compare what somebody originally said with how that account was subsequently classified.

It can trace where information travelled.

It can find apparently unrelated cases containing similar patterns.

It can compare promised reforms with subsequent outcomes.

It can expose contradictions between institutional descriptions and population experience.

It can make enormous bodies of evidence more intelligible to the people who must make decisions.

In other words, AI might become extraordinarily useful not because it supplies the final answer, but because it makes it harder for important questions to disappear.

That leaves us with the larger question that will run through this block:

Could AI help human society become more corrigible without becoming the authority that decides what correction means?

I don't know where that question will take us.

That is part of the reason for asking it.

We will begin with recursive self-improvement and examine what happens when systems become increasingly capable of pursuing their objectives.

We will look at human development and what happens when development is permitted only within boundaries supplied by authority.

We will examine institutions that learn, institutions that resist learning, and institutions that become increasingly effective at preserving themselves.

We will look at dissent, power, classification, rights, care, correction and recurrence.

And we will subject the ideas that emerge to our four tests rather than protecting them from difficult examples.

The four tests do not begin by asking whether the person complaining is right.

They begin somewhere more fundamental:

Is the system capable of finding out?

And if it discovers that somebody has been harmed, there is another question that cannot be left until the end:

What happens next?