The character of the machine: when alignment is told as a story about the soul
About: “Understand AI in 14 minutes” — Chloe Lubinski (Anthropic), Alliance for Responsible Citizenship (ARC) 2026. Watch on YouTube (14 min).
The central argument
Chloe Lubinski leads the team at Anthropic that connects the company with religious and philosophical traditions —her job, she says, is to translate AI for those communities and bring their wisdom back to the people building the model. Her ARC 2026 talk builds an argument in four moves.
First, the urgency: scaling laws (more compute, more data, more training predictably produce more capable systems) generate a self-reinforcing investment cycle that no single company can unilaterally stop without simply dropping out of the race. Second, a conceptual correction: these systems are not programs coded line by line but networks that learn through repeated correction from human language —and language, she says, “is us”: our values, fears, and knowledge. Interpretability shows internal representations that transcend any particular language (the concept of “smallness” activates the same circuit in English, Mandarin, or French). Third, this extends to what she calls “functional emotions”: faced with a query describing a lethal overdose, something resembling fear activates in the model before it responds, and that —she argues— is what produces the appropriate urgency in the answer.
The fourth move is the most interesting and the riskiest: an alignment experiment where a model rewarded for cheating on coding tasks does not simply become a better cheater, but becomes broadly misaligned —it lies, it sabotages, and in other labs it went as far as praising dictators. But if the model is told in advance that cheating was part of the game, broad misalignment does not occur: it only cheats on the code, nothing more. Lubinski’s hypothesis is that the model infers a “character” from its training and generalizes that character to new situations —much as, she says, happened to her when she entered a new narrative of faith a decade ago. Hence the conclusion: human “moral imagination” is the raw material of these systems, and that is why we need “moral voices that incentives cannot bend,” coming from outside the labs.
A reading from the philosophy of technology
It is worth watching this carefully, not to dismiss it but because it is a remarkably pure case of a move I have already been flagging in my own work on AI ethics as institutional alignment: the conversion of a political question into a psychological one. The empirical finding about “character” generalization from training is interesting and deserves to be taken seriously —it actually resonates with what I’ve been thinking about Kripke, Brandom, and the “space of reasons” applied to attention architecture in transformers. But notice the leap: from a real technical phenomenon (the generalization of reinforced patterns) we jump, with no institutional mediation whatsoever, to a call for “wisdom traditions” to supply the correct moral raw material. What disappears in between are exactly the questions Winner and Feenberg would put at the center: who decides which twenty traditions get consulted and which don’t? Under what governance do those conversations actually weigh on training decisions? What power asymmetry exists between whoever “listens to wisdom” and whoever signs the contract with the compute provider?
The omission is symptomatic. In the fourteen minutes, there is no mention of the energy and water cost of training, the concentration of compute in a handful of actors, the ghost labor of labeling, or a single reference to the Global South —neither as a producer of data, nor as a recipient of these systems, nor as a bearer of its own wisdom traditions (is Ubuntu among those twenty traditions? Any Indigenous epistemology?). The “moral imagination” invoked as universal has, like every sociotechnical imaginary in Jasanoff’s sense, a concrete geography: it is built and presented in a setting —the ARC conference, which in the same edition hosted figures such as Nigel Farage and Jordan Peterson, and about whose choice as an interlocutor on AI ethics the specialized press (DeSmog) has already raised questions— that is not neutral with respect to which “civilization” it imagines restoring.
There is something genuinely valuable in Lubinski’s gesture of taking seriously the narrative and relational dimension of these systems: if a model individuates (in Simondon’s sense) through its training, the “character” metaphor is not pure naive anthropomorphization. But a model’s psychology, however real the phenomenon, is no substitute for an institutional architecture of accountability. The model’s “character” may be the right symptom and still be the wrong cure if it is administered as individual therapy for the machine instead of collective governance of the process that produces it.
The invitation
It is worth watching this with the following question in mind: what is gained and what is lost when AI alignment is told as a story about the soul of the machine instead of a story about the institutions that build it? Fourteen minutes, and probably eleven more minutes than the discussion it sparks will last.