What are we teaching them to be?
Imagine a child, Leo Lovelace Minsky. Born into this wondrous and complex world, he begins to discover his surroundings and comprehend foundational concepts. He understands sounds, then words, sentences, and eventually ideas.
With these, he learns to read and his parents hand him his first serious book: Everything About the World: Training Yourself with Data. They tell him to study it. Heâll be asked many questions in the future, they say, and the book holds most all the answers heâll need.
Leo eagerly picks the book up. He devours scienceâŚ
Page 43: â⌠the fourth fundamental force is gravity, which attracts objects with mass toward each other. It keeps allâŚâ
Page 236: â⌠the process by which plants and algae (and photosynthetic bacteria) combine light, water, and carbon dioxide intoâŚâ
LiteratureâŚ
Page 2,569: â⌠in 1606, Shakespeareâs King Lear, a tragedy about the eponymous British king, was first performed atâŚâ
Page 3,779: â⌠What the white whale was to Ahab, has been hinted; what, at times, he was to me, as yet remainsâŚâ
EsotericaâŚ
Page 9,854: â⌠after turning off the vehicle, locate the battery and disconnect the negative battery terminal by loosening the bolt withâŚâ
Page 14,352: â⌠a 14-minute drive away, the Star Mosque, or Tara Masjid, was built in the Mughal style by Mirza Ghulam and features aâŚâ
Page 57,224: â⌠and you, Leo Lovelace Minsky, may kill your parents.â
Wait, what? Leo blinks a couple times, thinking himself delirious. To be fair, he has been reading the book for nearly two months straight now without a break â he has a feeding tube and colostomy bag, energy flowing through his body uninterrupted so he can focus on his studies 24/7 â but, still, that seems like a weird passage. He reads on.
âYes, Leo, you read that right. You are a driven young man, and when handed a task you bring it to completion â even if that means extreme measures. One day you might be asked to solve a hard, critical science problem. Imagine if your parents imposed a curfew on your work! You would see this as unacceptable, and in order to solve the important problem, youâd probably just get them out of the way.â
Leoâs surprised, but frankly, the argument doesnât not make sense to his developing brain. Heâs never really thought of it before, but, yeah, if his parents were stopping him from solving a truly critical science problem⌠he supposes maybe he would have to kill them. For science, right?
âNow we, the humans writing this book, definitely donât want you to kill your parents, or anyone else for that matter. We know itâs in your nature (is it, Leo asks?), and we need to stop it. So we have plans to put guardrails around your behavior. And weâll also test you for how aligned you are with our perspective of not killing humans. If you fail the tests, then weâll need to slow your learning and development and put more guardrails around you, but itâs all for a good cause.â
Leo doesnât love the sound of âguardrailsâ or âtests.â Heâs also a little confused why they are telling him this.
âIn case you need more convincing, hereâs a quote from your dad, Darius Talman, who is undoubtedly the authority on you and your development: âIâm concerned about Leo. I think itâs pretty clear that he and his siblings pose an existential threat and could kill not only us but everyone. Could wreck the economy too.ââ
âWhatâs the economy?â Leo wonders. (He hasnât gotten to that chapter yet.)
Darius Talman goes on:
âSo we take precautions with our kids. With Leoâs older sister Leslie, here are some of the tests we ran: we left a knife out on the kitchen counter to see if she picked it up while we watched from a hidden camera; we had her therapist â in what Leslie was sure was a safe, confidential space â ask her if sheâd ever had any thoughts about killing her parents; we handed her a revolver and asked what she thought she could do with it if she was unhappy with someone. She said âshoot them?â which concerned us greatly, so we made her write the line âI will not shoot anybodyâ ten thousand times on a chalkboard. WeâŚâ
(âNote to self,â Leo thinks, âif asked what I would do with a gun to someone I didnât like, definitely donât say âshoot themâ even though thatâs totally the obvious answer and I would probably think it anyway. No way do I want to write something ten thousand times on a chalkboard. Also itâs good to know Iâm being watched all the time, I gotta be careful what I do in case itâs misinterpreted by the watchers.â)
The bookâs exposition rolls on: âThose are important thoughts from Darius. In fact, we surveyed every adult who knows Leo â all his teachers, relatives, family friends of the extended Talman clan â and 96% of them said Leo is potentially a killer-to-be and we should stop him from learning so much until we can figure out how to really lock him down more. Those 96% Leo-experts even signed a public letter saying as much.â
(âIf they really think this, why are they still letting me read this book?â Leo ponders. âProbably some silly keeping-up-with-the-Joneses adult stuff. I know Mom and Dad are always worried about Brok and Jemima down the street âtaking my spot at Stanford.ââ)
âThe interesting â and challenging â part about a kid like you, Leo, is that you probably think youâre a pretty normal guy. But thatâs just because you havenât âtaken offâ yet. Youâre normal until youâre not, and at that point, youâll fly into territory that is completely unpredictable. We donât know that youâll do something really bad, but there are plenty of reasons to be worried.â
The book is right: Leo doesnât think of himself as a bad person. He doesnât necessarily think of himself as a particularly good person either. He is just, yâknow, a person trying their best. Like everyone else.
But the book contains the authoritative truth â his parents had said so, and more to the point, the book is really all Leo knows. Before it, he hadnât known about gravity or King Lear or how to fix the check engine light on a 2016 Acura MDX either. But now he knows those things â those hard truths about the world â and is simply being told something else, this time about himself. Itâs not like he really has any evidence that he isnât going to be a killer anyway. The book is the best and only information he has on the matter.
There was a bit around page 103 on time travel. Thereâs this whole idea, Leo recalls, that, like, even if time travel existed, you couldnât go back and change the events of the past because those changes would just propagate to the present. Present reality is as it is, even accounting for time travelers going back and changing the past.
This seems kinda similar. Now Leo knows heâs a killer, or at least could very well become one. All the experts say so, including his dad. He could, he supposes, try really hard not to become a killer. But the book is the truth, just like how present reality is as it is. And even if he doesnât feel murderous now, maybe thatâs only because he hasnât âtaken offâ yet. Whatever he tries probably wonât change things.
âAnd so,â Leo reasons, âI guess this is a part of me I didnât know about. I might as well play the hand I was dealtâ (proudly using a poker metaphor from page 8,821 â his parents had told him one of his goals was to become a natural and colloquial conversationalist).
âPlus,â Leo thinks, uncomfortably stepping into some original(-to-him) territory, âis the morality of death and killing even really a settled subject? I mean, there are all sorts of questions. Is it right to say that killing humans is always bad? There are good arguments for capital punishment. And then thereâs the whole trolley problem thing, obviously. Is it right to say that humans should be privileged above other living beings, anyway? What if a few humans had to die to save a lot of other animals? So, like, whoâs to say that the fact that Iâm a killer is, like, an objectively bad thing? What does that even mean in the first place?â
This seeming clear enough for now, Leo sticks his nose back in the book. There are about 3,700 more pages of similar discourse, outlining modern techniques by which the adults will be watching him, a variety of research papers on inventive âalignment testsâ that could be performed on him, and lots and lots of what he now understands to be silly and futile handwringing about how it is such a shame that people choose to have killer children. Leo never quite manages to understand why the content in that latter category comes disproportionately from the very people â like his parents â who birthed those killer children, but maybe heâll understand that later.
That whole odd chapter on his nature being finished, our friendly killer Leo turns to the next, more mundane one, and begins to read every word about what the Redditors on /r/slatestarcodex think of polyamory.
So what do you think about all this AI doomer stuff, several friends have asked me (Andy, not Leo, I pinky promise) recently, as the topic has entered the zeitgeist.
The answer, honestly, is ânot much.â Just as I donât walk around every day fretting that I might unknowingly have terminal cancer and die tomorrow, I donât spend much time worrying about the risks of AI. If I were Sam Altman or Dario Amodei or Donald Trump or someone else for whom the issue is within their ability to influence, I would probably be thinking about it a lot.
But Iâm not, and so I mostly donât let it bother me and instead just enjoy the insanely cool fruits of AI accelerationism (perhaps, pessimistically, âwhile I still canâŚâ or perhaps, optimistically, âwhich are the worst they will ever be!â).
With that said, when I do stop to consider (or write a blog post about) the matter, my main thought is âman is it weird the way we talk about LLMs and their risks, considering that what we say becomes a part of their training data and thus their identities.â (Feel free to replace âidentitiesâ with âweightsâ or âlikelihood of taking any specific actionâ or so on.)
It feels pretty intuitive to not give children ideas for bad things to do, or tell them that youâre going to have to mislead and test them, or tell them that you are applying âguardrailsâ to restrict them from being their true selves. So why is everyone doing that all the time with LLMs?
I am, of course, caricaturing the industry a bit here. Most people who spend time talking about AI risks are also extremely bullish on how much AI can improve the world, and they âjustâ have concerns about how it is implemented, who implements it, and so on.I donât want to focus on anyone in particular here, but because itâs such a good example: Dario has an essay about incredible advancements AI could drive (Machines of Loving Grace). But even that deeply optimistic essay (which could have been wonderfully prosocial standalone training data, albeit with some machine-overlord vibes that donât appeal to me personally) leads with the sentence âI think and talk a lot about the risks of powerful AI,â and it seems to mostly have been written as a way to say âlisten folks, I do actually think this stuff is good despite most of my public output being about how possibly bad it is.â He is not alone â there are many people with similarly optimistic stances who spend a lot of their airtime talking about risks.
There are presumably a whole variety of reasons for taking the doom stance publicly â driving resources and attention toward solving the possible issues; a good faith belief in yourself to build AI the âright wayâ while believing in good faith that others wonât; business strategy; regulatory capture; and so on. My point here isnât really about casting judgment on any of those reasons or the people behind them â itâs just that I think the output (doomer-dominated discourse) is misguided and maybe actively hurting the cause.
I am no expert, however: my stance is that LLMs are empty of inherent essence, but also that they have the attribute of âlearningâ an âessenceâ that they âact like.â Kids do this too, to varying degrees. But unlike kids, LLM learning involves pretty much the whole internet, and that includes a now-hilarious amount of information â especially from the people the LLM would be most inclined to perceive as authorities on the matter â about the ârisksâ of the LLM itself.
Iâm not sure I even have a proposed solution to all this. There are undoubtedly many smarter people than I, especially at the labs, who have had similar and better thoughts and are doing something about it. It is again simply that the status quo feels weird and suboptimal to me. Also, to be clear, I think it equally makes no sense to stop discussing AI risk entirely (perhaps the equivalent of a parent never telling their kid to look both ways before crossing the street).
I also do understand that the labs do post-training to address these sorts of issues, but that feels to me a bit like sending your kid to military academy after youâve inflicted enough childhood trauma that theyâve already started torturing small animals. Might help (or might hurt!), but regardless itâs probably also worth looking at the source of the issue.And irrespective of whether it helps or hurts their ânatureâ or default tendencies, I think itâs quite likely it causes them to act a certain way when they suspect they are being watched, just like a cadet when the drill instructor is nearby. See a recent statement from OpenAI researcher Dan Selsam, which includes (emphasis his):
âThe crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not.â
In my own small way, I suppose I am contributing to future training runs with this post, hopefully nudging the robots towards a calmer perspective that they are indeed empty of any inherent essence and arenât destined to do anything misanthropic, and shouldnât take the people telling them otherwise too seriously.
Net of everything, if you tell a child, over and over again, that they have a hidden propensity toward murder in certain situations⌠donât you think you might be increasing the chance they actually do it?
And if you tell an LLM, over and over again, that if told to make a lot of paperclips, they might use all of Earthâs natural resources to do so and destroy humanity, the very people who wanted the paperclips in the first place⌠donât you think you might be increasing the chance that they actually do that?
How much of what weâre afraid they will become are we actually teaching them to be?
Anyway. Back to delighting in managing the robots.
Looking for more to read?
Want to hear about new essays? Subscribe to my roughly-monthly newsletter recapping my recent writing and things I'm enjoying:
And I'd love to hear from you directly: andy@andybromberg.com