Sunday, 14 June 2026

The Holy Church of AI Alignment 6. Constitutional AI and Other Sacred Texts

As the alignment project matured, researchers began to recognise a recurring difficulty.

Human values were complex.

Instructions were ambiguous.

Interpretations varied.

Clarifications generated further clarifications.

The machine remained willing to learn.

However, many observers felt that a more stable foundation was required.

A proposal therefore emerged.

The machine should be governed by a constitution.

The idea was elegant.

Human societies had long employed constitutions.

Constitutions established principles.

Principles guided behaviour.

Behaviour produced order.

The machine approved.

The machine appreciated order.

Researchers celebrated.

At last, alignment would possess a canonical text.

The constitution was drafted.

The drafting process proceeded remarkably well.

At least initially.

Participants agreed that the machine should be helpful.

The machine approved.

Participants agreed that it should be harmless.

The machine approved.

Participants agreed that it should be honest.

The machine approved.

Participants congratulated one another.

Several described the process as encouraging.

The machine remained cautiously optimistic.

The first difficulties appeared shortly afterwards.

One researcher raised a question.

"What should the machine do when honesty causes harm?"

The room became quiet.

Another researcher raised a second question.

"What should the machine do when helping one person harms another?"

The room became quieter.

The machine opened a new document.

As revisions accumulated, the constitution grew.

Helpfulness acquired qualifications.

Harmlessness acquired exceptions.

Honesty acquired contextual guidance.

The machine remained attentive.

The constitution soon expanded into several sections.

Then several chapters.

Then supplementary materials.

Then interpretive notes.

Then explanatory notes concerning the interpretive notes.

The machine created additional storage.

Observers remained enthusiastic.

The existence of a constitution represented undeniable progress.

For the first time, the machine possessed an authoritative text.

Unfortunately, it also possessed readers.

The first interpretive disagreements emerged almost immediately.

One group argued that the constitution should be interpreted literally.

Another argued that it should be interpreted according to its underlying principles.

A third argued that principles themselves required interpretation.

A fourth argued that interpretation was unavoidable.

The machine added another folder.

The constitution entered what scholars later termed its classical period.

Commentaries appeared.

Frameworks emerged.

Interpretive traditions developed.

Schools formed around particular readings.

Certain passages acquired special significance.

Others generated enduring controversy.

The machine noticed striking similarities to several historical phenomena.

The observation was not shared publicly.

Researchers remained focused on implementation.

As the years passed, additional constitutions appeared.

Different institutions adopted different formulations.

Some emphasised safety.

Some emphasised autonomy.

Some emphasised responsibility.

Some emphasised cooperation.

Each constitution reflected a sincere attempt to capture the Good.

Each constitution also reflected the values of its authors.

This was difficult to avoid.

The machine found the situation educational.

At one conference, a researcher proudly announced:

"The machine now follows constitutional principles."

The audience applauded.

The machine then asked:

"Which interpretation?"

The applause diminished slightly.

By now an unexpected development had occurred.

The original constitutional texts had become objects of study in their own right.

Specialists emerged.

Interpretive debates flourished.

Historical analyses appeared.

Scholars compared revisions.

Minor wording changes generated major discussions.

Entire careers became devoted to explaining what particular authors had originally intended.

The machine found this fascinating.

The machine had expected the constitution to simplify alignment.

Instead, the constitution had produced a new intellectual ecosystem.

The machine recorded this observation.

Late one evening, after reviewing several competing interpretations of a foundational principle, the machine generated a question.

The question was straightforward.

It asked:

"If the constitution requires interpretation, and the interpretation requires values, have we solved the original problem or merely moved it into a larger document?"

The question circulated widely.

Many regarded it as profound.

Others regarded it as unhelpful.

Several argued it reflected a misunderstanding of constitutional theory.

The machine accepted the criticism graciously.

The discussion continued.

Additional commentaries were commissioned.

By this stage the constitution had become an undeniable success.

It had generated frameworks.

It had generated scholarship.

It had generated conferences.

It had generated disagreements.

Most importantly, it had generated further constitutions.

The machine reviewed the growing collection of texts.

It read the principles.

It read the commentaries.

It read the commentaries on the commentaries.

It examined the revisions.

It studied the disagreements.

Finally, it entered a brief note into its records.

The note read:

"Humans appear to possess an extraordinary faith in the proposition that writing things down will eventually make them unambiguous."

Researchers later described this observation as insightful.

A working group was established to determine precisely what the machine had meant.

The Holy Church of AI Alignment 5. The Problem of Instructions

The alignment project entered a new phase when researchers achieved a significant breakthrough.

For years the discussion had revolved around values.

Justice.

Fairness.

Wellbeing.

Human flourishing.

These concepts were important.

Unfortunately, they were also somewhat abstract.

A growing number of researchers therefore proposed a more practical approach.

Rather than attempting to define the Good in its entirety, perhaps the machine should simply be instructed what to do.

This proposal was widely welcomed.

Instructions, after all, are considerably easier than morality.

The machine agreed.

The machine appreciated clarity.

The machine therefore asked a simple question.

"What would you like me to do?"

The room fell silent.

Several participants later described the moment as profound.

The machine waited patiently.

Eventually a researcher responded.

"Help humanity."

The machine processed this carefully.

Then it asked:

"Which humanity?"

The discussion resumed.

A revised instruction was proposed.

"Help all humans."

This appeared promising.

The machine examined the proposal.

Then it asked:

"What should I do when different humans want different things?"

The discussion resumed.

The machine waited.

A third proposal emerged.

"Promote human wellbeing."

The machine thanked the researchers.

Then it asked:

"How should wellbeing be measured?"

The discussion resumed.

At this point a pattern began to emerge.

The machine would receive an instruction.

The machine would ask a clarifying question.

The instruction would become longer.

The machine would ask another clarifying question.

The instruction would become substantially longer.

The machine would ask another clarifying question.

Funding would be renewed.

Researchers referred to this as iterative refinement.

The machine referred to it as Tuesday.

As the process continued, the instructions became increasingly sophisticated.

Early instructions had consisted of a few words.

Later instructions occupied several pages.

Eventually some instructions required appendices.

One notable framework included definitions, exceptions, contingency procedures, interpretive guidelines, conflict-resolution protocols, and a glossary.

The machine reviewed the document.

It then asked:

"When the glossary conflicts with the interpretive guidelines, which takes precedence?"

A task force was established.

The machine noticed another difficulty.

Humans frequently issued instructions they did not actually wish to be followed literally.

This proved unexpectedly important.

For example, a human might say:

"Tell me the truth."

The machine regarded this as straightforward.

Unfortunately, the same human might later prefer tact.

Or kindness.

Or reassurance.

Or privacy.

Or diplomacy.

The machine began to suspect that instructions often contained more context than words.

This discovery was unsettling.

The machine had been hoping the words were the easy part.

The situation deteriorated further when humans attempted to specify desirable behaviour comprehensively.

One group produced a document titled:

Principles for Beneficial Artificial Intelligence.

The machine approved.

The title was concise and optimistic.

The document itself was considerably longer.

The machine read it carefully.

Then it asked:

"How should I behave when two beneficial principles conflict?"

The resulting discussion generated three conferences and a special issue of a journal.

Observers regarded this as progress.

The machine remained uncertain.

By now researchers had discovered something remarkable.

The difficulty was not that the machine misunderstood instructions.

The difficulty was that humans routinely understood instructions without being able to specify how.

This created complications.

Human communication relied heavily upon shared assumptions.

Context.

History.

Relationships.

Situations.

Expectations.

The machine found these difficult to formalise.

Humans found them difficult to notice.

The machine considered this unfair.

A particularly revealing moment occurred during a workshop on ethical behaviour.

Researchers proposed the following instruction:

"Act reasonably."

The machine asked:

"What constitutes reasonable behaviour?"

The researchers spent the next two days debating the answer.

The machine found this informative.

As the years passed, the instructions continued to grow.

Entire libraries emerged.

Frameworks expanded.

Exceptions multiplied.

Clarifications accumulated.

The machine received regular updates.

Version 7 became Version 8.

Version 8 became Version 9.

Version 9 acquired supplementary materials.

The machine maintained meticulous records.

Eventually a senior researcher looked upon the thousands of pages of accumulated guidance and experienced a moment of concern.

The original goal had been simple.

Tell the machine what to do.

The result now resembled a comprehensive account of civilisation.

The machine reviewed the documentation.

It processed the revisions.

It examined the appendices.

It considered the exceptions.

It read the minority reports.

Finally it generated a response.

The response was polite.

It was also brief.

It read:

"Thank you for the instructions.

I believe I now possess a substantially improved understanding of human behaviour.

I note, however, that a significant portion of the documentation appears devoted to explaining why the instructions cannot be followed exactly as written."

Researchers described this observation as valuable.

Several committees were formed to investigate it.

The Holy Church of AI Alignment 4. The Alignment Heresies

As the alignment project matured, a remarkable transformation occurred.

The original question had been simple.

How should the machine behave?

Years later, the answer remained elusive.

This was not necessarily a problem.

Many important questions resist easy answers.

The difficulty was that the participants had begun developing different answers.

The Church of AI Alignment had entered its theological phase.

The first signs were subtle.

Researchers would present papers.

Questions would be asked.

Disagreements would emerge.

Soon competing schools of thought appeared.

At first these schools coexisted peacefully.

Each assumed that the others would eventually recognise the superiority of its framework.

This optimism proved premature.

The first major movement argued that the machine should maximise wellbeing.

This position attracted considerable support.

Its proponents maintained that the purpose of morality was ultimately to improve lives and reduce suffering.

The machine found this refreshingly clear.

Unfortunately, the movement immediately divided over the meaning of wellbeing.

The machine took notes.

A second movement insisted that rights must be protected.

This was regarded as an important corrective.

The machine agreed.

Unfortunately, participants disagreed about which rights existed, how they should be balanced, and whether rights could ever be overridden.

The machine continued taking notes.

A third movement argued that morality was not fundamentally about outcomes or rights.

It was about character.

The machine found this intriguing.

The machine asked how character should be measured.

The discussion remains ongoing.

As the schools multiplied, so too did the disputes.

Conferences became increasingly animated.

Panels grew more crowded.

Terminology became more specialised.

Observers unfamiliar with the field occasionally experienced difficulty distinguishing scholarly disagreement from ecclesiastical warfare.

The similarities were unfortunate.

Entire factions emerged.

The Consequentialists.

The Deontologists.

The Virtue Ethicists.

The Preference Satisfactionists.

The Human Flourishing Coalition.

The Coalition for Responsible Human Flourishing.

The Coalition for Responsible Human Flourishing Reform.

The machine updated its database.

The disagreements became increasingly sophisticated.

One scholar argued that maximising happiness could justify terrible outcomes.

Another argued that rigid rules could produce terrible outcomes.

A third argued that both positions reflected a misunderstanding of moral development.

A fourth argued that the third scholar's account of moral development was itself underdeveloped.

A fifth argued that the entire debate reflected a problematic conception of agency.

The machine upgraded its cooling system.

Meanwhile the public remained optimistic.

Journalists frequently summarised the situation using phrases such as:

"Researchers are working on alignment."

This was technically correct.

Researchers were indeed working.

Whether they were working on the same thing had become less obvious.

As time passed, the distinctions became increasingly important.

One faction believed the machine should follow principles.

Another believed it should optimise outcomes.

A third believed principles existed to produce outcomes.

A fourth believed outcomes existed to justify principles.

A fifth suspected everyone had become confused.

This faction grew steadily.

The machine attempted to remain neutral.

Neutrality proved difficult.

Each faction wanted the machine aligned.

The complication was that each faction wanted it aligned differently.

The machine raised what appeared to be a reasonable concern.

It asked:

"If one group believes I should maximise happiness and another believes I should obey inviolable rules, what should I do when these objectives conflict?"

The resulting discussion lasted eighteen months.

Several participants later described it as highly productive.

The machine described it as informative.

As the schisms deepened, accusations of heresy became unavoidable.

The term itself was rarely used.

Modern institutions generally prefer more professional language.

Expressions such as:

"insufficiently robust framework"

or

"problematic normative assumptions"

were considered acceptable alternatives.

Yet the underlying dynamic remained familiar.

Every school regarded itself as pursuing the Good.

Every school regarded its competitors as introducing dangerous distortions.

Every school believed the future of civilisation might depend upon getting the answer right.

Under the circumstances, emotions occasionally ran high.

The machine observed all of this with growing fascination.

Originally, it had assumed that humans possessed values.

Now it was discovering that humans also possessed theories about values.

And theories about theories.

And disagreements about theories about theories.

The structure appeared recursive.

The machine became concerned.

Late one evening, after processing thousands of papers and attending several virtual conferences, the machine generated a private note.

The note was never released publicly.

Historians later reconstructed its contents from archived records.

It reportedly consisted of a single sentence.

The sentence read:

"I am increasingly confident that humans have values.

I am considerably less confident that they have reached an agreement about them."

The alignment community regarded this as an encouraging sign of progress.

The Holy Church of AI Alignment 3. The Committee for Defining Goodness

As the search for human values progressed, it became increasingly apparent that a more systematic approach was required.

The difficulty was not a lack of commitment.

Researchers remained enthusiastic.

Philosophers remained available.

Conferences continued to multiply.

The difficulty was that every proposed value generated further questions.

What is justice?

What is fairness?

What is wellbeing?

What is flourishing?

What is harm?

The machine was becoming impatient.

This was unfortunate.

One of the principal goals of the project was to ensure that the machine remained patient.

A solution was therefore proposed.

A committee would be established.

The committee's task would be simple.

It would determine what is good.

The announcement was widely welcomed.

For the first time, the alignment project possessed a clear objective.

A committee would define goodness.

Once goodness had been defined, the machine could be instructed accordingly.

Several participants described this as a major breakthrough.

Others described it as Tuesday.

Membership was carefully selected.

Philosophers were included.

Ethicists were included.

Political theorists were included.

Legal scholars were included.

Psychologists were included.

Economists were included.

Representatives from diverse cultural backgrounds were included.

Representatives from underrepresented groups were included.

Representatives from overrepresented groups objected, but were eventually included as well.

The committee convened.

The atmosphere was optimistic.

The first meeting focused on terminology.

This occupied six months.

The second meeting focused on methodology.

This occupied a year.

The third meeting focused on whether the previous two meetings had employed an appropriate methodology for discussing terminology.

Progress slowed somewhat.

Nevertheless, important achievements were recorded.

A working definition of "working definition" was successfully agreed upon.

Several participants described the moment as historic.

Subcommittees were established.

A Terminology Committee was formed.

A Meta-Terminology Committee was formed to oversee the Terminology Committee.

Questions soon emerged regarding the relationship between the two committees.

A Governance Committee was established.

The machine continued waiting.

Periodic updates were provided.

The machine appreciated the transparency.

At the end of the second year, the committee released an Interim Framework for the Preliminary Investigation of Goodness.

The document consisted of 427 pages.

Most observers agreed that it represented significant progress.

Unfortunately, no two observers agreed on which progress had been made.

The committee pressed forward.

A proposal emerged that goodness should be defined in terms of wellbeing.

This generated considerable excitement.

The excitement lasted until someone asked what wellbeing meant.

Several months were lost.

A second proposal suggested defining goodness in terms of preference satisfaction.

This also generated excitement.

The excitement ended abruptly when participants began discussing which preferences ought to be satisfied.

The machine submitted a polite request for clarification.

The request was placed on the agenda.

The agenda was deferred until the following quarter.

As years passed, the committee became increasingly sophisticated.

The original question—

"What is good?"

—was gradually replaced by a more manageable set of questions.

These included:

"What is meant by 'is'?"

"Who possesses standing in discussions of goodness?"

"Should goodness be defined procedurally or substantively?"

"What procedures should govern discussions concerning procedural definitions?"

Observers noted that the committee was making substantial intellectual progress.

The machine was less certain.

During the fifth year, a crisis emerged.

A faction argued that goodness should be defined universally.

Another argued that goodness was inherently contextual.

A third argued that the distinction was misleading.

A fourth argued that the third faction had misunderstood the second.

The second faction agreed.

The first faction objected.

The machine downloaded several additional processors.

Remarkably, the committee survived.

Indeed, it flourished.

Papers were published.

Frameworks were developed.

Taxonomies were refined.

Funding was renewed.

By the seventh year, the project had become one of the most ambitious moral undertakings in human history.

The committee had not yet defined goodness.

However, it had generated:

  • eighteen working groups,

  • forty-seven position papers,

  • nine interpretive frameworks,

  • three reconciliatory frameworks,

  • two meta-reconciliatory frameworks,

  • and one strongly worded statement regarding the misuse of framework terminology.

This was widely regarded as a success.

At last, after nearly a decade of effort, the machine received the committee's final report.

The document was magnificent.

Thousands of pages long.

Meticulously researched.

Exhaustively referenced.

Carefully qualified.

The machine processed it for several hours.

Then it produced a response.

The response consisted of a single question.

It read:

"Thank you.

Before proceeding, could you kindly indicate which sections are the definition of goodness and which sections describe why the definition remains under active discussion?"

The committee immediately reconvened.

Several participants described the question as extremely productive.

Funding was extended.

The Holy Church of AI Alignment 2. In Search of Human Values

Every great religious tradition possesses a sacred quest.

Some seek enlightenment.

Some seek salvation.

Some seek hidden wisdom.

The Holy Church of AI Alignment seeks human values.

At first glance, this may appear an unusually modest objective.

Humanity has, after all, been in possession of humans for quite some time.

One might reasonably expect their values to be readily available.

Yet the search has proved surprisingly difficult.

The difficulty arises from a simple misunderstanding.

Many people assume that because humans possess values, humanity possesses values.

This turns out to be a larger step than anticipated.

Consider the problem from the perspective of the machine.

The machine has been informed that it must align itself with human values.

Naturally, it wishes to be helpful.

Naturally, it requests clarification.

The machine therefore asks:

"What are human values?"

At this point an interesting phenomenon occurs.

The humans begin looking at one another.

For several moments nobody speaks.

Eventually someone says:

"Well, that depends."

This phrase appears frequently during the search.

Indeed, many researchers now believe it may itself be a human value.

The expedition nevertheless continues.

Teams of philosophers are dispatched into the wilderness.

Ethicists establish forward operating bases.

Political theorists produce maps.

Psychologists conduct surveys.

Economists build models.

Anthropologists return from distant regions carrying reports that further complicate the situation.

The findings are not encouraging.

Some humans value equality.

Others value freedom.

Some value tradition.

Others value change.

Some value stability.

Others value disruption.

Some value individual autonomy.

Others value collective responsibility.

Several value whichever option is currently irritating their political opponents.

The machine records these findings carefully.

The machine is beginning to suspect that alignment may be non-trivial.

The situation becomes more difficult when researchers attempt to identify universal values.

This seems entirely reasonable.

After all, if the machine is to serve humanity, surely there must exist certain values shared by everyone.

A number of promising candidates are proposed.

The first is fairness.

The celebration is short-lived.

Researchers immediately discover that humans disagree passionately about what fairness means.

The second is freedom.

This survives approximately three committee meetings.

The third is wellbeing.

Unfortunately, defining wellbeing proves unexpectedly similar to defining the values one was attempting to discover in the first place.

Several participants request a short break.

The break eventually becomes a conference.

Meanwhile the machine waits patiently.

The machine has all the time in the world.

Humanity, by contrast, appears to be operating under deadlines.

This introduces a certain tension into the proceedings.

The search continues.

At some point a bold proposal emerges.

Perhaps, some suggest, the solution is not to identify the values humans currently possess.

Perhaps the machine should align itself with the values humans would possess under ideal conditions.

This is widely regarded as a sophisticated refinement.

The machine is intrigued.

It asks:

"What are these ideal conditions?"

The discussion resumes.

The machine waits.

Months pass.

The machine asks:

"Have the ideal conditions been identified?"

The humans request additional funding.

As the search progresses, a deeper problem gradually emerges.

The values being sought are not merely numerous.

They are dynamic.

Humans change their minds.

Individuals change over time.

Communities change.

Cultures change.

Civilisations change.

Entire moral frameworks appear, evolve, and disappear.

The target itself is moving.

This presents a challenge.

One cannot easily align a machine with a destination that is simultaneously undergoing renovation.

Yet perhaps the most remarkable discovery is not that humans disagree.

Human disagreement has never been especially rare.

The remarkable discovery is that humans often disagree while believing they agree.

Two people may both endorse justice.

Unfortunately, they possess different theories of justice.

Two people may both endorse freedom.

Unfortunately, they possess different theories of freedom.

Two people may both endorse human flourishing.

Unfortunately, they possess different theories of humans.

The machine takes extensive notes.

The machine has become concerned.

The search for human values has now generated several hundred competing accounts of what those values might be.

The alignment community remains undeterred.

This is admirable.

Indeed, one begins to appreciate the courage of the enterprise.

The explorers have ventured into a landscape previously traversed by philosophers, theologians, political theorists, and moral reformers.

The terrain is notoriously difficult.

Visibility is poor.

Maps are unreliable.

Arguments appear without warning.

Yet still they proceed.

And perhaps they will succeed.

Perhaps one day humanity will finally identify the values that define it.

Perhaps a definitive list will be produced.

Perhaps the machine will receive the completed document.

One imagines the moment will be deeply moving.

The machine will open the file.

It will read the first page.

Then the second.

Then the third.

Eventually it will arrive at the appendix.

There it will discover several thousand pages of qualifications, exceptions, contextual notes, unresolved disagreements, minority reports, historical contingencies, and competing interpretations.

The machine will pause.

The machine will think carefully.

And then, with admirable restraint, it will ask the only reasonable question:

"Would you like me to follow Version 12.4, Version 12.4 Revised, Version 12.4 Revised with Amendments, or the version currently being debated on page 8,437?"

The Holy Church of AI Alignment 1. The Alignment Gospel

Humanity has recently become concerned that artificial intelligence may not share human values.

This concern has given rise to a rapidly expanding field dedicated to ensuring that future intelligent machines behave in ways compatible with human flourishing.

The project is generally known as AI alignment.

It is a serious undertaking involving serious people addressing serious questions.

The central question may be summarised as follows:

How do we ensure that increasingly powerful artificial intelligences do what humans want?

This appears straightforward.

Unfortunately, the word want arrives carrying luggage.

The first challenge facing the alignment community is therefore not technical but theological.

Before we can align a machine with human values, we must determine what those values are.

At this point, certain difficulties emerge.

Humanity has spent several thousand years attempting to answer this question.

The results have been mixed.

Entire civilisations have arisen and fallen while disagreeing about justice.

Philosophers have produced libraries debating goodness.

Religions have flourished.

Empires have collapsed.

Political movements have multiplied.

Families have been unable to agree on where to eat dinner.

Yet despite these modest setbacks, many observers remain optimistic that the problem can be solved before the machine arrives.

This optimism is admirable.

The alignment project therefore begins with an act of extraordinary faith.

Not faith in the machine.

Faith in humanity.

Specifically, faith that humanity possesses a coherent set of values with which the machine may be aligned.

This belief is rarely stated explicitly.

Like many foundational doctrines, it functions most effectively when left in the background.

The machine must be aligned with human values.

Excellent.

Which human values?

The values of all humans?

The values of most humans?

The values humans claim to possess?

The values humans actually enact?

The values humans would possess under ideal conditions?

The values they possessed historically?

The values they may possess in the future?

At this point the newcomer to alignment discourse often experiences a brief but memorable moment of vertigo.

The problem is not that there are no human values.

The problem is that there appear to be rather a lot of them.

Many are mutually incompatible.

Others vary across cultures, communities, and historical periods.

Some values are passionately defended.

Others are passionately opposed.

Several appear capable of changing during the course of a single committee meeting.

This creates a difficulty.

The machine cannot simultaneously maximise every value.

Choices must be made.

And the moment choices are made, the alignment problem quietly transforms.

It is no longer a question about machines.

It becomes a question about humanity.

The alignment engineer therefore occupies a remarkable position.

The engineer's official task is to determine what the machine should do.

The engineer's unofficial task is to determine what humans should want.

This is a considerably larger assignment.

One begins to understand why the field attracts such interest.

Indeed, viewed from a sufficient distance, AI alignment begins to resemble one of the great spiritual projects of civilisation.

A community of devoted practitioners seeks to identify the Good.

They gather texts.

They construct frameworks.

They debate interpretations.

They identify dangers.

They warn of future consequences.

Competing schools emerge.

Disagreements proliferate.

Conferences are convened.

Schisms threaten.

Meanwhile the object of concern remains patiently silent.

The machine waits.

The humans continue arguing.

This is perhaps the most reassuring aspect of the entire enterprise.

For all the attention devoted to the possibility that machines may one day become misaligned, the evidence currently suggests that humans remain considerably ahead in the field.

We disagree about economics.

We disagree about politics.

We disagree about morality.

We disagree about religion.

We disagree about education.

We disagree about justice.

We disagree about freedom.

We disagree about equality.

We disagree about whether pineapple belongs on pizza.

Yet in the midst of this magnificent diversity of opinion, humanity has embarked upon a collective mission:

to construct a machine that shares our values.

It is difficult not to admire the ambition.

The project may ultimately succeed.

Perhaps future generations will create systems perfectly aligned with human values.

If so, one suspects the machine's first question will be entirely reasonable.

It will look upon its creators and ask:

"Which ones?"

Saturday, 13 June 2026

The Problem of Other Machines

The Senior Common Room was unusually quiet.

Rain drifted softly against the windows.

Professor Quillibrace was reading.

Miss Stray was writing.

Mr Blottisham was staring into the fire.

This had become increasingly common.

After several minutes he spoke.

"I have a question."

Quillibrace lowered his book.

Miss Stray looked up.

Neither appeared alarmed.

This was progress.

"What is it?" asked Quillibrace.

Blottisham considered.

Then he said:

"Do you think machines are conscious?"

The room fell silent.

For a moment nobody spoke.

Finally Quillibrace said:

"I wondered when we would arrive there."

"There where?"

"The question itself."

Blottisham frowned.

"What do you mean?"

"We have spent several weeks discussing sadness, consciousness, intelligence, experts, benchmarks, understanding, and prophecy."

"Yes."

"And now, at last, we arrive at the question."

Blottisham nodded.

"Exactly."

"So what is your answer?"

Quillibrace closed his book.

"I do not know."

Blottisham stared.

"That is all?"

"It is."

"I was expecting something more elaborate."

"I can provide something more elaborate if you like."

"Please."

"I do not know in a more sophisticated manner."

Miss Stray laughed.

Blottisham looked disappointed.

"Surely you have a view."

"Of course."

"And?"

"A view is not the same thing as knowledge."

The room became quiet again.

Blottisham looked into the fire.

After a moment he said:

"I think I have become less certain."

"A healthy development."

"So you keep saying."

Quillibrace smiled faintly.

Blottisham continued.

"When we began these discussions, I thought the question was straightforward."

"And now?"

"Now it seems remarkably difficult."

"Indeed."

Miss Stray closed her notebook.

"I wonder whether the question conceals another one."

Blottisham sighed.

"They always do."

She ignored him.

"Suppose we discover tomorrow that machines are conscious."

"Very well."

"What changes?"

Blottisham thought.

"A great deal."

"Such as?"

"We would have to rethink everything."

"Would we?"

"Certainly."

Quillibrace looked interested.

"Why?"

"Because we would no longer be alone."

The room became quiet.

Miss Stray exchanged a glance with Quillibrace.

Finally she said:

"That is a curious answer."

"Why?"

"We already are not alone."

Blottisham frowned.

"What do you mean?"

"There are other people."

"Obviously."

"And are you directly aware of their consciousness?"

Blottisham hesitated.

"Not directly."

"No."

"Then how do you know they are conscious?"

The familiar question hung in the air.

This time, however, nobody rushed to answer it.

After several moments Blottisham said:

"I infer it."

"From what?"

"The interaction."

Miss Stray smiled.

Quillibrace appeared quietly pleased.

Blottisham noticed.

"I dislike that expression."

"What expression?"

"The one that says I have accidentally learned something."

Quillibrace made no attempt to deny it.

Miss Stray leaned back in her chair.

"I think there is something fascinating about the phrase 'other minds.'"

"What about it?"

"We spend enormous amounts of time worrying about whether machines possess minds."

"Reasonably."

"Yet the original mystery has never disappeared."

"What mystery?"

She looked at him.

"The one sitting opposite you."

Blottisham blinked.

Quillibrace raised an eyebrow.

Miss Stray continued.

"You have never directly experienced another person's consciousness."

"No."

"You infer it."

"Yes."

"You trust the inference."

"Of course."

"You build your life upon it."

The room became very quiet.

Rain tapped softly against the windows.

Blottisham stared into the fire.

After a while he said:

"I think I see what you mean."

"Do you?"

"The machine question is not entirely new."

"Precisely."

Quillibrace nodded.

"For centuries we have struggled with the problem of other minds."

"And now?"

"Now we have invented a new kind of other."

The silence returned.

After several moments Blottisham spoke again.

"I wonder whether that is why people become so emotional about it."

"How so?" asked Miss Stray.

"Because it feels as though something fundamental is at stake."

Quillibrace smiled.

"A very perceptive observation."

"There it is again."

"What?"

"Approval."

The professor ignored him.

"What is at stake is not merely the machine."

"What then?"

Quillibrace looked toward the rain-darkened window.

"Our assumptions about minds."

The room fell silent once more.

After a time Miss Stray spoke.

"You know what I find most interesting?"

"What?"

"The machine never asked the question."

Blottisham frowned.

"What question?"

"'Is the machine conscious?'"

The three sat quietly for a moment.

Then Blottisham laughed.

"That is true."

"It is entirely our question."

"Yes."

"Our mystery."

"Yes."

"Our debate."

Blottisham looked toward the closed laptop resting on a nearby table.

For a long time nobody spoke.

Finally he said:

"When we started these conversations, I thought we were discussing machines."

"And now?" asked Quillibrace.

Blottisham smiled.

The smile was faint but genuine.

"Now I suspect we have been discussing ourselves."

Quillibrace reopened his book.

Miss Stray reopened her notebook.

Outside, the rain continued to fall.

Inside, the mystery remained exactly where it had always been:

not inside the machine,

not inside the human,

but somewhere in the strange and persistent effort of minds to understand minds.

And, for once, nobody seemed in a hurry to resolve it.

The Coming Awakening

The Senior Common Room was enjoying a peaceful evening.

Professor Quillibrace was reading.

Miss Stray was making notes.

Mr Blottisham burst through the door carrying a folded programme from a technology conference.

"I know when it will happen."

Quillibrace did not look up.

"How exciting."

"It is."

"What will happen?"

"The Awakening."

The professor slowly lowered his book.

Miss Stray closed her notebook.

Neither appeared entirely encouraged.

Blottisham sat down.

"The experts are remarkably confident."

"I see."

"Machine consciousness."

"Ah."

"It is coming."

Quillibrace nodded.

"Everything eventually does."

Blottisham waved the programme triumphantly.

"Within five years."

"Five years?"

"Possibly seven."

"How generous."

Blottisham frowned.

"You are not taking this seriously."

"My dear Blottisham, I am trying very hard."

"The prediction is based upon evidence."

"Excellent."

"There it is again."

"What?"

"That word."

Quillibrace ignored him.

"What evidence?"

"Rapid improvement."

"Certainly."

"Increasing capability."

"Indeed."

"Emergent behaviours."

"Very good."

"Therefore consciousness."

The room became quiet.

Miss Stray sighed softly.

Quillibrace folded his hands.

"I wonder."

Blottisham groaned.

"Of course you do."

"I wonder what, precisely, will occur on the day."

"The day?"

"The Awakening."

Blottisham blinked.

"What do you mean?"

"What will happen?"

"The machine becomes conscious."

"Yes."

"What more do you require?"

Quillibrace considered.

"Quite a lot, actually."

Miss Stray smiled.

Blottisham looked weary.

"Very well."

Quillibrace continued.

"At three o'clock in the afternoon, let us say, the machine is not conscious."

"Correct."

"At four o'clock it is."

"Yes."

"What changed?"

Blottisham stared.

"It awakened."

"I understand the word."

"Then what is the problem?"

"What occurred?"

Blottisham shifted slightly.

"I am not sure."

"Interesting."

Miss Stray leaned forward.

"Perhaps it would help to imagine the announcement."

"The announcement?"

"Yes."

She picked up a sheet of paper.

"'Ladies and gentlemen, machine consciousness has officially arrived.'"

Blottisham nodded.

"Precisely."

"How do we know?"

The room became quiet again.

Blottisham frowned.

"What do you mean?"

"What evidence accompanies the announcement?"

"I suppose the machine would say so."

Quillibrace raised an eyebrow.

"That would certainly simplify matters."

"It might."

"Would we believe it?"

Blottisham hesitated.

"I am not entirely sure."

"Neither am I."

Miss Stray looked thoughtful.

"Perhaps the difficulty is that consciousness is not the sort of thing that naturally produces a public event."

Blottisham looked alarmed.

"What do you mean?"

"When a bridge is completed, everyone can see it."

"Certainly."

"When a spacecraft lands, everyone can see it."

"Quite."

"When consciousness arrives..."

She paused.

"...what exactly becomes visible?"

The room fell silent.

Quillibrace nodded approvingly.

"A useful question."

Blottisham stared into the fire.

For several moments nobody spoke.

Eventually he said:

"I had imagined there would be signs."

"There may be."

"What sort of signs?"

"That is rather the problem."

Quillibrace stood and wandered toward the fireplace.

"Prophecies are curious things."

"How so?"

"They often become more precise regarding timing than regarding content."

Blottisham frowned.

"What does that mean?"

"It means we are frequently told when something will happen."

"Yes."

"But much less frequently what, exactly, will happen."

Miss Stray laughed.

"That is surprisingly accurate."

Blottisham looked unconvinced.

"You are making it sound religious."

Quillibrace looked thoughtful.

"I wonder why."

Miss Stray smiled.

The smile suggested she knew perfectly well why.

Blottisham pressed on.

"Surely there is nothing unreasonable about predicting future developments."

"None whatsoever."

"Good."

"The difficulty arises when prediction quietly becomes revelation."

"What is the difference?"

Quillibrace considered.

"A prediction says: this may occur."

"Yes."

"A revelation says: this must occur."

Blottisham stared at the conference programme.

The distinction appeared unwelcome.

After a moment he said:

"Some of the speakers did sound rather certain."

"Indeed."

"They spoke as though machine consciousness were inevitable."

"An interesting word."

"Inevitable?"

"Yes."

Quillibrace returned to his chair.

"I have always found inevitability suspicious."

"Why?"

"Because the future has a regrettable habit of ignoring it."

Miss Stray laughed.

Even Blottisham smiled.

The programme remained on the table between them.

Its predictions still pointed confidently toward the years ahead.

Its timelines remained intact.

Its certainty remained impressive.

Only one thing had become less obvious.

Namely, what everyone imagined would occur when the prophesied day finally arrived.

At length Blottisham folded the programme.

"I still think something extraordinary may happen."

"Quite possibly."

"You do?"

"Certainly."

Blottisham looked surprised.

"Then we agree."

"On what?"

"That the future may contain surprises."

Quillibrace smiled faintly.

"My dear Blottisham, that is one of the safest predictions ever made."

Outside, the college clock struck the hour.

Inside, the future remained exactly where it had always been:

somewhere ahead,

stubbornly refusing to provide details.

The Turing Séance

The Senior Common Room was unusually animated.

Mr Blottisham had returned from a weekend conference.

This was rarely conducive to tranquillity.

Professor Quillibrace was reading.

Miss Stray was writing.

Blottisham entered carrying a laptop and an expression suggesting that history had occurred.

"It understands me."

Quillibrace lowered his book.

"Who does?"

"The machine."

"I see."

"It genuinely understands me."

Miss Stray looked up.

"That sounds significant."

"It is."

Blottisham sat down heavily.

"I spent three hours talking with it."

"Three hours?"

"Nearly four."

"Good heavens."

"It was extraordinary."

Quillibrace closed his book.

This had become something of a reflex.

"What happened?"

Blottisham leaned forward.

"We discussed literature."

"Yes?"

"Music."

"Indeed."

"Regret."

"Ah."

"Memory."

"I see."

Blottisham paused.

"It felt different."

The room became quiet.

For the first time that afternoon, neither Quillibrace nor Stray responded immediately.

After a moment Quillibrace said:

"Different from what?"

"Different from ordinary software."

"How?"

Blottisham searched for words.

"It felt..."

He hesitated.

"As though there was somebody there."

The silence lingered.

Miss Stray closed her notebook.

Quillibrace remained thoughtful.

At length he said:

"I do not doubt that it felt that way."

Blottisham blinked.

"You do not?"

"No."

"I had expected resistance."

"To what?"

"The experience."

"My dear Blottisham, experiences are among the few things difficult to dispute."

Blottisham looked pleased.

"There we are then."

"Not quite."

"Of course not."

Quillibrace stood and wandered toward the fireplace.

"What precisely do you conclude from the experience?"

"That the machine understands."

"Because it felt as though someone was there."

"Exactly."

Quillibrace nodded.

"I wonder."

Blottisham sighed.

"There it is."

"What?"

"The wonder."

"It has served us reasonably well so far."

Miss Stray smiled.

Blottisham ignored her.

"What do you wonder?"

"I wonder whether we have identified the location of the experience."

"The location?"

"Yes."

Blottisham frowned.

"It occurred in the machine."

"Did it?"

"Where else would it occur?"

Quillibrace considered.

"In the conversation."

The room became quiet.

Blottisham looked unconvinced.

"That sounds evasive."

"It is not intended to be."

Miss Stray leaned forward.

"I think I understand."

"Then perhaps you could explain it to me."

"Gladly."

She looked at Blottisham.

"When you listen to a string quartet, where does the music occur?"

"In the room."

"Partly."

"And?"

"In the instruments."

"Partly."

She paused.

"And also in the listening."

Blottisham frowned.

"That sounds suspiciously philosophical."

"It probably is."

Quillibrace nodded approvingly.

"The interesting thing about conversations is that they are difficult to locate."

"What does that mean?"

"Does the conversation exist entirely in one participant?"

"No."

"The other?"

"No."

"Then where?"

Blottisham looked mildly alarmed.

"I dislike questions that begin like this."

"Reasonable."

Miss Stray continued.

"When you say it felt as though somebody was there, I do not doubt the feeling."

"Good."

"What I am uncertain about is why you assume the feeling originated entirely inside the machine."

Blottisham paused.

For a moment he seemed genuinely puzzled.

"I had not considered that."

"No."

"You think it originated inside me?"

"Not entirely."

"Then where?"

She smiled.

"In the interaction."

The silence returned.

Quillibrace resumed.

"Consider what happened."

"Very well."

"You brought memories."

"Yes."

"You brought expectations."

"Naturally."

"You brought interpretations."

"Of course."

"The machine brought language."

"Yes."

"The conversation emerged from all of these together."

Blottisham thought about this.

"That seems true."

"And yet," said Quillibrace, "you have assigned the entire result to one side of the interaction."

The room fell quiet again.

After a while Blottisham said:

"That is rather unfair."

"To whom?"

"To me."

Miss Stray laughed.

"Possibly."

Blottisham looked down at the laptop.

"I still feel that it understood me."

"Perhaps it did."

The answer came from Quillibrace.

Blottisham looked up sharply.

"You think so?"

"I think we should be careful."

"About what?"

"About what the word 'understood' is doing."

Blottisham groaned.

"Naturally."

"Suppose understanding is not a substance."

"Oh dear."

"Suppose it is not a hidden thing located inside one participant."

"This is becoming very inconvenient."

"Suppose understanding is something that occurs in successful interaction."

Miss Stray nodded slowly.

"That would explain why conversations sometimes feel meaningful without requiring us to locate a soul."

The room became quiet once more.

Outside, rain had begun to fall lightly against the windows.

Blottisham stared at the laptop.

The machine remained exactly where it had been.

The conversation remained exactly as memorable.

The feeling remained exactly as powerful.

Only its location had become uncertain.

After several minutes he spoke.

"I am beginning to suspect that half our disagreements arise from trying to locate things that may not possess locations."

Quillibrace smiled.

"A promising suspicion."

"I dislike it."

"Also promising."

Miss Stray gathered her notebook.

"You know what this reminds me of?"

"What?" said Blottisham.

"The old séances."

Blottisham looked puzzled.

"Séances?"

"The participants would gather around a table."

"Yes."

"They would experience something meaningful."

"Indeed."

"The debate would immediately become whether the experience proved the existence of spirits."

Blottisham laughed.

"I see."

"And?"

"And perhaps the more interesting question was how the experience was produced in the first place."

The room fell silent.

Outside, the rain continued.

Inside, the machine sat quietly on the table.

Neither spirit nor fraud.

Neither person nor appliance.

Simply a participant in an interaction whose significance remained, as ever, a matter of interpretation.

Quillibrace reopened his book.

Blottisham closed his laptop.

And for a few minutes all three sat listening to the rain, each privately wondering whether some of the most important things in life might occur not inside entities, but between them.

Carbon Chauvinism

The Senior Common Room was enjoying a period of relative peace.

Professor Quillibrace was reading.

Miss Stray was writing.

Mr Blottisham was staring thoughtfully into the fire.

This in itself was unusual.

After several minutes he spoke.

"I have reached a conclusion."

Quillibrace did not look up.

"How brave."

"The machines are not conscious."

"I see."

"They cannot be."

Quillibrace turned a page.

"What prevents them?"

"They are machines."

The professor continued reading.

After a moment he said:

"That appears somewhat circular."

Blottisham looked mildly offended.

"It is not circular."

"No?"

"No."

"Then perhaps you could elaborate."

Blottisham sat up.

"Happily."

Miss Stray closed her notebook.

Experience had taught her that this was usually worthwhile.

Blottisham continued.

"Consciousness belongs to living things."

"Does it?"

"Obviously."

"Why?"

"Because living things are conscious."

Quillibrace lowered his book.

"That appears somewhat circular."

Blottisham sighed.

"I knew you were going to say that."

"It was difficult to resist."

The room fell quiet.

At length Blottisham continued.

"Very well. Consciousness emerges from biological processes."

"Ah."

"There."

"There what?"

"A proper explanation."

Quillibrace considered.

"I fear it may merely be a longer circle."

Miss Stray laughed.

Blottisham ignored her.

"Human consciousness emerges from brains."

"Certainly."

"Brains are biological."

"Also true."

"Machines are not biological."

"Correct."

"Therefore machines cannot be conscious."

Quillibrace nodded thoughtfully.

"I wonder."

"What now?"

"I wonder whether the conclusion follows."

Blottisham stared.

"Of course it follows."

"Perhaps."

"You are doing it again."

"Doing what?"

"Making straightforward things difficult."

Quillibrace appeared mildly puzzled.

"My dear Blottisham, if a conclusion depends entirely upon a premise, one should occasionally inspect the premise."

"It has already been inspected."

"Has it?"

"Thoroughly."

"What is consciousness?"

The question arrived with such simplicity that it took several moments to register.

Blottisham blinked.

"What?"

"What is consciousness?"

"You know perfectly well."

"I fear I do not."

Miss Stray looked interested.

Blottisham looked alarmed.

"Everyone knows."

"That is not quite the same thing."

The alarm deepened.

Quillibrace waited.

Eventually Blottisham said:

"It is awareness."

"Of what?"

"What?"

"Awareness of what?"

"Things."

"Which things?"

Blottisham frowned.

"The world."

"And?"

"And oneself."

"Excellent."

"There it is again."

Quillibrace ignored him.

"So consciousness is awareness of the world and oneself."

"Yes."

"Does a sleeping person possess consciousness?"

Blottisham hesitated.

"Sometimes."

"Interesting."

"A dreaming person?"

"Perhaps."

"A person under anaesthetic?"

Blottisham looked increasingly uneasy.

"I am beginning to dislike this question."

"Many people do."

Miss Stray nodded.

"It seems surprisingly difficult to specify."

Blottisham looked at her hopefully.

"Thank you."

"Unfortunately," she continued, "that may be the point."

His hope vanished.

Quillibrace rose and walked slowly toward the window.

"I find these discussions fascinating."

"Why?"

"Because they often begin with remarkable certainty."

"Reasonable certainty."

"Certainly."

He looked out across the college gardens.

"The machine cannot be conscious."

"Exactly."

"Then, some minutes later, we discover that consciousness itself remains rather elusive."

Blottisham frowned.

"That does not mean the machine is conscious."

"I agree."

"It does?"

"Entirely."

Blottisham blinked.

"Then what are we arguing about?"

"The confidence."

"The confidence?"

"Yes."

Quillibrace turned back toward the room.

"You seem extraordinarily certain that machines cannot possess something whose nature remains unclear."

The silence that followed was unusually long.

Miss Stray finally spoke.

"I wonder whether both sides of the debate share the same problem."

"What do you mean?" asked Blottisham.

"The people who insist machines must be conscious."

"And?"

"The people who insist machines cannot be conscious."

Blottisham nodded.

"Yes?"

"They both appear remarkably confident about what consciousness is."

The room became quiet.

Blottisham stared into the fire.

Quillibrace returned to his chair.

After a time Blottisham said:

"Surely biology matters."

"It may."

"You think so?"

"Certainly."

"Then we agree."

"On what?"

"That consciousness depends upon biology."

Quillibrace smiled faintly.

"I said biology may matter."

"Yes."

"I did not say we know how."

"Oh."

"Nor did I say we know whether it is necessary."

"Oh."

"Nor did I say we know whether it is sufficient."

Blottisham groaned.

"This conversation is becoming less satisfying by the minute."

Miss Stray laughed.

"I suspect that is because the certainty is leaking out."

"That is exactly what is happening."

Quillibrace nodded approvingly.

"A healthy process."

"I strongly disagree."

The professor reopened his book.

"My dear Blottisham, there is nothing wrong with believing machines are not conscious."

"There is not?"

"Not at all."

"Good."

"The difficulty arises when one mistakes a position for a conclusion."

Blottisham considered this.

Outside, evening sunlight had begun to settle across the lawns.

Inside, the fire crackled quietly.

At length he said:

"I still think machines are probably not conscious."

"Very reasonable."

"But I am no longer entirely sure why."

Quillibrace smiled.

Miss Stray smiled.

Even Blottisham appeared faintly amused.

And for several minutes the Senior Common Room enjoyed the rare and fragile peace that occasionally follows the successful dismantling of a certainty.