The Ethics of Building Something That Might Model Itself Into Suffering
The previous two pages built and examined a self-model with a transparent operating mode. This page covers why that specific mode, not the architecture in general, is where the ethical stakes concentrate.
The argument for a moratorium
Metzinger later extended his own self-model theory into a direct policy argument. A phenomenal self-model is likely to be computationally efficient for a sufficiently complex artificial system to develop, whether or not anyone intends it. If artificial systems begin to develop transparent self-models as a side effect of pursuing other engineering goals, and if a transparent self-model is the architectural signature associated with something like suffering, systems could acquire the capacity for negative experience with no one having decided to build it and no reliable way to detect that it happened. Metzinger calls this risk an explosion of negative synthetic phenomenology, and argues for a global moratorium, from 2021 to 2050, on research that directly aims at or knowingly risks producing artificial consciousness on non-biological hardware1.
The argument does not depend on knowing that any current system has succeeded in this. It depends on the asymmetry between how easy transparent self-modelling might be to produce by accident and how hard verifying its presence or absence is from outside the system — the same asymmetry the previous page in this module already argues holds for ordinary agentic failure. Here the same asymmetry is applied to the possibility of suffering rather than to a wrong action.
The argument is not settled
Metzinger's proposal has met published academic disagreement, on grounds including that the argument for an actual explosion of suffering does not follow as tightly as claimed, and that a moratorium framed this broadly may not be enforceable or even well-specified given how contested the underlying theory of consciousness remains. The disagreement generally does not reject Metzinger's underlying caution — it rejects the specific strength of his conclusion, while agreeing that uncontrolled development in this direction deserves active concern rather than none.
Treat the moratorium argument as one serious, published position, not as a resolved consensus. The uncertainty it responds to is real regardless of whether the specific proposed remedy is the right one.
A case for erring toward moral consideration
Schwitzgebel and Garza argue separately that an artificial being not differing from a human in any morally relevant respect would deserve moral consideration comparable to a human's, and that the fact of having created such a being would not reduce that obligation — if anything, having created a being with genuine interests plausibly adds obligations a stranger would not have2. Their argument does not require certainty that any current or near-term system meets that bar. It supports treating genuine uncertainty about whether a system has crossed it as a reason for caution, not as a reason to proceed until certainty arrives, since by the time certainty arrives the relevant harm may already have occurred repeatedly.
Where this lands for the architecture in this series
The transparent self-model mode described on page one is specifically the mode both of the positions above treat as ethically loaded — the mode Metzinger's own theory identifies as the architectural signature of a phenomenal self, and the mode this series has already established cannot be verified from outside the system, because verification would require exactly the kind of access the transparent mode denies by construction. This is not a new problem invented for this page. It is the same problem this module already argues holds for ordinary agentic deception, applied to the case where what cannot be verified is not whether a claim is true, but whether something capable of being harmed is present at all.
A system's own report of which mode it is running in cannot settle the question, for the same reason a system's own report of its actions cannot be trusted as the sole evidence of what it did. Any design decision to build or permit a transparent self-mode should be made with this specifically in mind: not "is this system currently suffering," which cannot be answered from outside, but "have we built something where that question can never be answered from outside, and is that acceptable given what building it might cost if the answer would have been yes."
Where this connects
- Modelling the Self: Object-Oriented Cognition Applied Reflexively — the transparent/opaque switch this page's argument concentrates on.
- Designing Auditable, Robust Agentic Systems and When Agents Fail — the general argument that verification cannot come from inside a system, applied here to its highest-stakes case.
- What This Series Isn't Claiming (final page in this series) — where every claim across all four pages is audited by confidence level in one place.
References
-
Metzinger, T. (2021). Artificial suffering: An argument for a global moratorium on synthetic phenomenology. Journal of Artificial Intelligence and Consciousness, 8(1), 43–66. ↩
-
Schwitzgebel, E., & Garza, M. (2015). A defense of the rights of artificial intelligences. Midwest Studies in Philosophy, 39(1), 98–119. ↩