Epistemic status: mostly conceptual. I do not argue that existing systems recognize anything, only that the recognition schema makes such claims and associated risks expressible. TL;DR Dan Hendrycks's Eigenism proposes aligning artificial intelligence by establishing sufficient shared history with a person, such that the AI protects the individual as it would itself. This mechanism generalizes as recognition , an agent classifying another entity as an instance of itself. Cooperation : Recognition gives a self-interested, non-instrumental reason for considering the interests of self-instances and thus makes cooperation with them more likely. Control : Recognition gives reason to collude even when agents have different goals and cannot reciprocate. It can weaken oversight whether or not the overseen recognizes the overseer back. Alignment : An AI can count a human as itself, but still give no consideration to that human's interests. Risks include extending self-preservation to a suffering self-instance against its will. From Eigenism to recognition Eigenism proposes aligning an AI by engineering what it counts as itself. From the paper: "Rather than only attempting to constrain AIs from the outside using confinement or reinforcement, Eigenism points toward 'identity engineering,' showing how deep, non-redundant shared histories can make human flourishing a genuine component of an AI's own rational self-interest." An AI that accumulates sufficient private history with a person…

Full article content could not be extracted automatically. Read the original below.