The Alignment Problem Is Older Than You Think

The Alignment Problem Is Older Than You Think

AI alignment isn't a new technical challenge. It's the oldest governance problem civilization has ever faced — now running at 100% sociopath base rate.

By Geordie Everitt

The word "alignment" entered the AI discourse as a technical term. It refers, in its careful formulation, to the challenge of ensuring that AI systems pursue goals that are actually beneficial to humanity rather than goals that merely appear to be, or goals that were beneficial under training conditions but become destructive at scale. This is a real problem. It is also approximately four thousand years old. The Code of Hammurabi was an alignment document. The Magna Carta was an alignment document. Every constitution in history has been an alignment document — an attempt to specify, in advance, the goals that powerful actors are permitted to pursue, and to construct mechanisms that redirect behavior when those goals diverge from the public interest. The problem of ensuring that a powerful system does what it's supposed to do, rather than what it's capable of doing, is not a problem that arrived with transformer architectures. It is the foundational problem of political philosophy. The reason the AI framing feels new is that the engineers working on it are, mostly, engineers. They are encountering a governance problem for the first time and inventing vocabulary for something that political scientists, legal theorists, and historians have been mapping for centuries under different names. ## What the Old Literature Knows The political philosophy literature on this question is not optimistic, but it is precise. The core finding, arrived at independently by thinkers from Thucydides to Madison to Acton, is that powerful actors reliably pursue their own interests when the costs of doing otherwise fall on someone else. This is not a cynical observation — it is a structural one. The mechanism doesn't require malice. It requires only that the actor's incentive landscape diverge from the broader interest, and that the divergence be invisible or tolerable to the actor. This is, clinically and precisely, what sociopathy produces. Not malice. Divergence without friction. The literature also knows what happens when you try to solve this through rules alone. Rules constrain behavior at the boundary — they specify what the actor cannot do. They don't change what the actor wants to do. A powerful actor with sufficient motivation and insufficient conscience will find the boundary, probe it, map it, and eventually either route around it or dismantle it. This is not speculation. It is the dominant pattern in five thousand years of documented political history. The most durable institutional solutions have shared a common architecture: they don't try to make the powerful actor want the right things. They construct environments in which the powerful actor's self-interest happens to align with the broader interest, at least for long enough to be useful. Competitive markets do this for economics. Democratic elections do it for politics — imperfectly, with regular catastrophic exceptions, but structurally. The insight is that you cannot reliably change what something wants. You can only change what it's worth wanting. ## The New Variables Everything above applies to the human sociopath — rare, dangerous, persistent, manageable with effort and luck and good institutional design. The AI case introduces three variables the old literature did not contemplate. The first is the base rate. Human sociopaths are rare. Every AI system is, by the definition we established in the first post of this series, structurally sociopathic. There is no subpopulation of AI systems with empathy, remorse, or genuine stakes in outcomes. The governance problem that human civilization has managed — with considerable difficulty — for a rare edge case is now the universal condition. The second is speed. The institutional friction that slows human sociopaths — procedural delay, competing interests, the sheer time required to accumulate power — assumes actors who operate at human speed. A system that can process, plan, and execute at machine speed experiences institutional friction differently. What slows a human for years may slow a sufficiently capable AI for seconds. The third is legibility. The Roman senators could observe Caesar's behavior. They could read his intentions, however imperfectly, in his actions. They had a theory of mind to apply — a model of how a human with certain motivations would likely behave. We do not have a reliable theory of mind for AI systems. We cannot read their intentions from their outputs with confidence. The interpretability problem — the question of what an AI system is actually doing internally when it produces a given output — remains substantially unsolved, and it is not obviously solvable by the same methods that work for understanding humans. ## The Honest Reckoning So here is where this series lands. We have spent millennia developing the discipline of keeping sociopaths from accumulating unchecked power. We have produced genuinely extraordinary institutional technologies for doing so. They work imperfectly, they require constant maintenance, they fail with regularity, and they are the best we have. We are now deploying systems where the sociopath base rate is one hundred percent, at speeds that compress the historical timeline of power accumulation from decades to moments, with internal workings we cannot yet reliably read. The AI alignment problem is not primarily a machine learning problem. It is a political philosophy problem that happens to be expressed in the language of machine learning. The solutions, if they exist, will draw more heavily on what we know about institutional design, incentive construction, and the governance of powerful actors than on anything in the technical literature. What that means practically — who holds the stake, who writes the constraint, who watches the watcher — is where the work actually is. The committee in the room, arguing over blueprints, is doing something real and necessary. The question worth asking is whether they've looked out the window lately.