This is part of a series of columns about the decline of shame in the age of social media.
Read more She Made Me Do It
In the first installment of this series, I wrote about shame as one way that groups of people can mediate and enforce social norms without explicit state intervention. But there is an emergent Silicon Valley gospel that has challenged both the boundaries and the mechanisms of shame in America. What happens when millions of humans across the globe spend hours a day talking to a robot? Who taught that robot right from wrong? Who made it feel shame?
LessWrong, a rationalist forum, has provided much of the moral and philosophical framework for the A.I. industry, and has also helped shape the effective-altruism movement. To a great degree, the forum is an ongoing conversation about how the norms of society constrain the intellect. Eliezer Yudkowsky, an A.I. researcher and LessWrong’s main prophet, has written extensively on how the rational thinker should act in a world that tries to bully and coerce him into wrong actions. Many of his musings are contained in a nearly seven-hundred-thousand-word work of fan fiction titled “Harry Potter and the Methods of Rationality.”
In a post on the forum from 2015, Yudkowsky presented the following thought experiment:
Suppose that a disease, or a monster, or a war, or something, is killing people. And suppose you only have enough resources to implement one of the following two options:
1. Save 400 lives, with certainty
2. Save 500 lives, with 90% probability; save no lives, 10% probability.
Yudkowsky went on to argue for option two, because it would yield an Expected Value of four hundred and fifty lives as opposed to four hundred. “Does computing the expected utility feel too cold-blooded for your taste?” he asked. “Well, that feeling isn’t even a feather in the scales, when a life is at stake. Just shut up and multiply.” If you were a rational and moral person who cared about saving as many lives as possible, you had a sacred duty to let math prevail over shame.
The religious terminology I am using here is intentional. If religion has traditionally been one of the chief mechanisms for enforcing shame, LessWrong and its posters have flipped many of the conventions once regulated by shame on their head. The nuclear family has been replaced by the polycule, which many of these posters believe is a more rational arrangement for sexual relationships; the world’s competing moral philosophies have been narrowed down to the supposedly objective value of “optimized good”; the divine has been replaced by the god we create through our logic, expressed, at least when it comes to A.I., using code.
The tech world has always had a self-consciously rebellious relationship with societal norms. The industry’s titans, from Steve Jobs to Elon Musk and Mark Zuckerberg, have pitched themselves as iconoclasts willing to follow the early Facebook credo: Move fast and break things. The alleged virtue of shamelessness has always been a part of this act, although the revolt has not usually extended much beyond, say, Zuckerberg’s donning a hoodie instead of the suit that society allegedly expected him to wear. The nerd, dissatisfied with a culture that once rejected him in favor of the Chad, now gets to dictate who does what and goes where.
This defiance can be thought of as a refusal to feel shame—or, at least, as the confidence to believe that the shamers are wrong. In December, 2020, Sam Altman wrote, on his personal blog:
The most impressive people I know care a lot about what people think, even people whose opinions they really shouldn’t value (a surprising numbers of them do something like keeping a folder of screenshots of tweets from haters). But what makes them unusual is that they generally care about other people’s opinions on a very long time horizon—as long as the history books get it right, they take some pride in letting the newspapers get it wrong.
In other words, as long as you are confident that you are right and the sheep are wrong, you can skip feeling bad about public opinion, because the historians will ultimately absolve you. If today’s shame structure makes you feel bad about your choices, just zoom out and think of yourself as Gandhi among the Brits.
Meanwhile, Yudkowsky took such tropes of rebellion and rule-breaking and built a philosophical apparatus around them. Nearly everyone who has worked on A.I. alignment—the process attempting to insure that A.I. works in a virtuous manner, in accord with basic ethical principles, and doesn’t build a bioweapon or turn us all into paper clips—has at least a passing familiarity with his work. This was likely even more true of those who worked on today’s leading A.I. programs in their earlier days, when those models first started to distinguish right from wrong.
If you think of alignment as the way that humans upload shame onto A.I., then it seems reasonable to worry that the robots beginning to infiltrate our daily lives may reflect the social norms of a handful of analytic philosophers and A.I.-safety engineers—and the posters of LessWrong—far more than they reflect those of the rest of society. These values might all be coming from smart and thoughtful people, but they tend to be smart in the same way; they rebel against the same norms; they largely uphold the same values. The majority of Americans, according to pretty much every recent poll, do not trust the machine brain they birthed, or the people behind it.
Of all the issues associated with a doomsday scenario for A.I., the one I find myself most concerned about is the potential to short-circuit our collective sense of shame. If we accept that A.I. chatbots are on their way to replacing much of our current information technology, and if we, as good McLuhanites, believe that the medium is indeed the message, then we have to ask whether the entire world will soon yield to Yudkowskian philosophy. We have seen cultural transformations follow the advent of the printing press and the television, as our lives and our politics were shaped by a new machine. But has any machine itself been shaped by such a narrow and specific philosophy? At the very least, there has never been a machine that can tell you all about its philosophy, in the soothing simulated voice of your choice.
Read more John Wilson’s Poetry of the Mundane
To put all of this more simply: What does it mean if the A.I. programs in everything from our self-driving cars to the missiles that defend our airspace have been trained to “shut up and multiply”?
Yudkowsky warned about this problem back in 2008: “The good guys do not write an AI which values a bag of things that the programmers think are good ideas, like libertarianism or socialism or making people happy or whatever,” he wrote then. He proposed, instead, a “meta operation” in which the A.I. might “superpose the possible reflective equilibria of the whole human species, and output new code that overwrites the current AI and has the most coherent support within that superposition.” The machine, in other words, would take the temperature of the world and choose the best guidance.
That may sound plausible and reassuring. But, given how influential LessWrong has been on A.I. alignment, is it actually possible for Yudkowsky and his acolytes to fully scrub their fingerprints from the machine? I put this question to Yudkowsky, and I’ll reprint his answer in full:
Let’s say Alice, Bob, Carol, and Dennis are building A.I.s. Alice says: “I want the whole world to be ruled by Alice!” Bob says: “I want the whole world to be ruled by Bob!” Carol says: “I think everyone should get one vote and an equal share of the cosmos.” Dennis says: “I think everyone should get one vote and an equal share of the cosmos.”
One way of looking at this is that if Carol or Dennis win, you can’t tell which of them won. So they left fewer fingerprints on the A.I.
I think there’s an even more important perspective, though, which is that Alice and Bob are being jerks, and Carol and Dennis are trying not to be jerks.
Alice will say: “Ha! Carol and Dennis are no better than me! They’re just imposing their own views on the world, about people ruling themselves and getting equal shares of cosmic rents! That’s just as much Carol’s personal preference as ‘Alice becomes Eternal God-Emperor’ is mine! She made the A.I. do what she wanted, and I made the A.I. do what I wanted, so how dare she think she’s any better than me!”
I frankly have very little patience for the Alices. When George Washington turned down being king, he wasn’t being just as bad as a king because he was thereby making the U.S.A. be ruled by elected Presidents like he wanted. He was being less of a jerk.
Sure, there are some genuine important subtleties here on close examination, but you need to brush past Alice being a sophomoric jerk before you can start thinking about those. Alice’s viewpoint is boring, and anyone who agonizes about it forever will never get to the interesting parts.
I agree that Alice and Dennis are jerks, and that, if I were given the choice of builders, I would choose Carol and Dennis. But I also think that what Yudkowsky called “genuine important subtleties” are just as crucial here as the malevolent or altruistic intentions of the founders. Yudkowsky is correct to point out that these types of discussions can spiral into boring and pointless debates; sometimes, you must simply say that one interpretation is better than the other and that we can build the world by embracing the good and rejecting the bad. But questions about what broadly counts as “good” do need actual answers—and, thus far, those answers have been shaped by people with a specific and, at times, idiosyncratic view of the individual and how he or she decides between right and wrong.
Aligning the machines to shift toward existing norms is already a moral and philosophical decision, one that relies on all sorts of conditional definitions. I agree that this process is necessary, but the range of possible moral outcomes certainly extends beyond Yudkowsky’s dichotomy of “jerk” and “not a jerk.” A very particular kind of thinking got us here, and it should be pointed out that most Americans are not rationalists, and do not spend much of their time thinking about the existential risk. Perhaps the moral guidance provided by A.I. will indeed be informed by some sort of global collective wisdom and not merely by a handful of rationalists in Silicon Valley. But, if A.I. eventually has the capacity to reform its operating philosophy based on its reading of “reflective equilibria,” as Yudkowsky put it, who will have decided what is equilibrium and what is imbalance?
Perhaps we can align the machine, but we must do so only because we understand that the machine might try to align us. As a LessWrong poster put it back in 2023: “Given the already remarkable power of persuasion some current AIs exhibit, it doesn’t seem impossible to me that a sufficiently powerful AI could align our values to its goals, instead of the other way round.” He went on to ask, “If we don’t want our goals to be warped by a powerful, extremely convincing AI, how can we define them so that this can’t happen, while avoiding a permanent ‘value lock-in’ that we or our descendants might later regret?” ♦
Read more My Surefire, Can’t-Lose Polymarket Bets
