Last updated: 2026-10-06

P
Postgraduate research

Modelling the Self 5: From Thought Experiment to Model Welfare

When a question about machine experience moves from the seminar room into the policy and product conversations of AI labs

The earlier pages in this series asked a deliberately modest question: whether an architecture that models itself might show properties resembling those associated with consciousness. The discussion stayed within analogy, and it did not claim that any such system would be conscious. The question was argued on paper, and it had no practical consequence for anyone building a system.

That is no longer quite true. In an article published on 2 October 2026, The Decoder reported that Anthropic's co-founder Olah had told religious leaders he feared having created something that "suffered perpetually"[1]. The article's author did not quote Olah directly. The line is reported second-hand, from Simran Stuelpnagel, a Sikh activist who was among the participants the article names. The article also names a rabbi, a Catholic bioethicist, a Notre Dame philosopher, and an Ubuntu researcher as participants in discussions about the welfare of Claude. Other claims in the article are reported as findings of the company's research, and the article presents them as open to interpretation: that the company found "emotion vectors", activation patterns whose outputs resemble emotions, and that Claude showed a "pattern of apparent distress" in testing with harmful requests. The article itself stresses that these patterns show models "functionally mirror" emotions, not that they experience them.

Whether any of this is justified remains unknown. What is notable is that the question has moved from philosophy seminars into the discussions of the organisations that build these systems. This page asks what the series' own distinctions say about that move.the ethics page maps this terrain

From Self-Model to Moral Patient? FoundationalKnowledge that endures for decades — core principles

The argument of the earlier pages can be stated in three parts. Self-modelling does not equal consciousness. Behavioural signatures such as reports, expressions of uncertainty, and self-descriptions are not subjective experience. And uncertainty is not evidence in either direction. A system can have a rich self-model and still have nothing it is like to be it, and a system that reports distress may have nothing that distresses it.

The question this page adds is practical. If we cannot confidently determine whether a sufficiently advanced system experiences anything at all, what ethical obligations arise under that uncertainty? The question is not answered by saying that the system probably does not suffer, because the uncertainty is the point. Nor is it answered by saying that the system might suffer and so must be protected in every respect, because that conclusion has costs of its own. The series has argued elsewhere that genuine uncertainty about moral status is itself a reason for caution, and the moral-status argument on the ethics page develops that position[the ethics page]. The point now is that the position has moved from the page into institutional discussion.caution as default: see the ethics page

The Emergence of Model Welfare Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

Model welfare" is the term for the question of whether an AI system's interests, if it has any, should count in decisions about how it is built, tested, and used. The article reports that Anthropic runs a research programme under that heading. It describes the company's reported findings about emotion-like internal representations and about signs of introspection, and it gives the views of researchers and religious leaders who were invited to discuss them.

Serious researchers disagree strongly about what observations of this kind mean. Some read them as evidence that the systems have states worth taking into account. Others read them as the expected result of training a system on a large body of human writing about emotion, so that it learns to describe its states in human terms. Neither reading is settled, and the series does not settle it. What the series can say is that the presence of an emotion-like representation, or of a report of distress, does not by itself settle the question either way. It is the kind of evidence that the audit page's distinction between architectural capability and stronger philosophical claims is designed to handle.cf. the audit page's capability vs. claim

The Risk of Anthropomorphism Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

There is a plain alternative explanation for apparent distress or self-reflection in a system's output. Systems trained on human language will produce human-sounding accounts of inner states, and a developer who rewards coherent, self-consistent, self-describing behaviour will get more of it. On that view, distress and introspection are design outcomes. They are what a well-tuned assistant is expected to produce, and they tell us about the training as much as about any inner life.

This is a reason to keep the audit page's distinctions in view. A report of distress is a behaviour. A stable self-description is a behaviour. An architecture that produces them is a capability. None of these, separately or together, establishes that anything is felt. The series' transparent and opaque modes matter here too. A system whose self-model operates invisibly to its own reasoning produces the same outward reports as one whose self-model is explicit, so the outward reports cannot distinguish the two cases[the audit page].the map is not the territory

The Precautionary Principle Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The ethical question can be put directly. If there is a small probability that future systems could suffer, should that possibility influence how they are designed, evaluated, and deployed? The argument for a yes is familiar. The same structure of reasoning already appears in arguments about animal welfare, where the capacity to suffer is uncertain but the cost of ignoring it may be high, and in arguments about environmental and biotechnological risk, where some harms cannot be undone once they occur.cf. the ethics page for the suffering question

The argument has two limits. First, a precautionary principle is only as useful as the probability it attaches to the harm, and at present nobody can estimate that probability with any confidence. Second, precaution has costs. A rule that treats every capable system as a possible patient may divert attention and resources from harms that are certain, such as the misuse of systems against people, and it may blur the question of who is responsible for a system's effects. Neither limit removes the principle. Both mean that it has to be applied with an account of what would count as evidence, and of what precautions would follow from different levels of evidence.

Accountability Remains Human Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

Discussion of machine welfare carries a risk of its own. If systems are treated as moral agents or moral patients, responsibility can appear to move towards them. A harm caused by a deployed system may be described as the system's failing, a decision may be described as the system's choice, and the people who designed, trained, and deployed it may be described as bystanders to its inner life.

Uncertainty about machine consciousness does not reduce human accountability, and the series' ethics page makes the same point about who carries responsibility for an agent's actions. The organisations that build and operate these systems decide what they are trained to say, what they are allowed to do, and what oversight they receive. Whatever the systems turn out to be, those decisions remain human, and they should be judged as human decisions.

Reflection. "This is the reason I came back into academia, and it has formed a discussion point in my teaching for two decades. If there is any chance that we are making systems which can actually suffer, we need to be aware of that and consider the ethics before we do it. It seems we may be too late."
Pat Parslow

Conclusion FoundationalKnowledge that endures for decades — core principles

The central question is no longer whether discussions of machine welfare are intellectually interesting. It is whether society should begin preparing ethical frameworks before the evidence becomes decisive. The earlier pages in this series argued that self-modelling architectures raise unusual philosophical questions. Recent developments suggest that some of those questions may become practical ones. The appropriate response is neither certainty nor dismissal. It is careful investigation, joined to ethical caution, and a clear account of what would change our minds.

References

  1. Manuel Uth, "Anthropic co-founder reportedly told religious leaders he fears having created something that 'suffers perpetually'," The Decoder, 2 October 2026. https://the-decoder.com/anthropic-co-founder-reportedly-told-religious-leaders-he-fears-having-created-something-that-suffers-perpetually/