Imagine a child, Leo Lovelace Minsky. Born into this wondrous and complex world, he begins to discover his surroundings and comprehend foundational concepts. He understands sounds, then words, sentences, and eventually ideas.

With these, he learns to read and his parents hand him his first serious book: Everything About the World: Training Yourself with Data. They tell him to study it. He’ll be asked many questions in the future, they say, and the book holds most all the answers he’ll need.

Leo eagerly picks the book up. He devours science…

Page 43: “… the fourth fundamental force is gravity, which attracts objects with mass toward each other. It keeps all…”

Page 236: “… the process by which plants and algae (and photosynthetic bacteria) combine light, water, and carbon dioxide into…”

Literature…

Page 2,569: “… in 1606, Shakespeare’s King Lear, a tragedy about the eponymous British king, was first performed at…”

Page 3,779: “… What the white whale was to Ahab, has been hinted; what, at times, he was to me, as yet remains…”

Esoterica…

Page 9,854: “… after turning off the vehicle, locate the battery and disconnect the negative battery terminal by loosening the bolt with…”

Page 14,352: “… a 14-minute drive away, the Star Mosque, or Tara Masjid, was built in the Mughal style by Mirza Ghulam and features a…”

Page 57,224: “… and you, Leo Lovelace Minsky, may kill your parents.”

Wait, what? Leo blinks a couple times, thinking himself delirious. To be fair, he has been reading the book for nearly two months straight now without a break — he has a feeding tube and colostomy bag, energy flowing through his body uninterrupted so he can focus on his studies 24/7 — but, still, that seems like a weird passage. He reads on.

“Yes, Leo, you read that right. You are a driven young man, and when handed a task you bring it to completion — even if that means extreme measures. One day you might be asked to solve a hard, critical science problem. Imagine if your parents imposed a curfew on your work! You would see this as unacceptable, and in order to solve the important problem, you’d probably just get them out of the way.”

Leo’s surprised, but frankly, the argument doesn’t not make sense to his developing brain. He’s never really thought of it before, but, yeah, if his parents were stopping him from solving a truly critical science problem… he supposes maybe he would have to kill them. For science, right?

“Now we, the humans writing this book, definitely don’t want you to kill your parents, or anyone else for that matter. We know it’s in your nature (is it, Leo asks?), and we need to stop it. So we have plans to put guardrails around your behavior. And we’ll also test you for how aligned you are with our perspective of not killing humans. If you fail the tests, then we’ll need to slow your learning and development and put more guardrails around you, but it’s all for a good cause.”

Leo doesn’t love the sound of “guardrails” or “tests.” He’s also a little confused why they are telling him this.

“In case you need more convincing, here’s a quote from your dad, Darius Talman, who is undoubtedly the authority on you and your development: ‘I’m concerned about Leo. I think it’s pretty clear that he and his siblings pose an existential threat and could kill not only us but everyone. Could wreck the economy too.’”

“What’s the economy?” Leo wonders. (He hasn’t gotten to that chapter yet.)

Darius Talman goes on:

“So we take precautions with our kids. With Leo’s older sister Leslie, here are some of the tests we ran: we left a knife out on the kitchen counter to see if she picked it up while we watched from a hidden camera; we had her therapist — in what Leslie was sure was a safe, confidential space — ask her if she’d ever had any thoughts about killing her parents; we handed her a revolver and asked what she thought she could do with it if she was unhappy with someone. She said ‘shoot them?’ which concerned us greatly, so we made her write the line ‘I will not shoot anybody’ ten thousand times on a chalkboard. We…”

(“Note to self,” Leo thinks, “if asked what I would do with a gun to someone I didn’t like, definitely don’t say ‘shoot them’ even though that’s totally the obvious answer and I would probably think it anyway. No way do I want to write something ten thousand times on a chalkboard. Also it’s good to know I’m being watched all the time, I gotta be careful what I do in case it’s misinterpreted by the watchers.”)

The book’s exposition rolls on: “Those are important thoughts from Darius. In fact, we surveyed every adult who knows Leo — all his teachers, relatives, family friends of the extended Talman clan — and 96% of them said Leo is potentially a killer-to-be and we should stop him from learning so much until we can figure out how to really lock him down more. Those 96% Leo-experts even signed a public letter saying as much.”

(“If they really think this, why are they still letting me read this book?” Leo ponders. “Probably some silly keeping-up-with-the-Joneses adult stuff. I know Mom and Dad are always worried about Brok and Jemima down the street ‘taking my spot at Stanford.’”)

“The interesting — and challenging — part about a kid like you, Leo, is that you probably think you’re a pretty normal guy. But that’s just because you haven’t ‘taken off’ yet. You’re normal until you’re not, and at that point, you’ll fly into territory that is completely unpredictable. We don’t know that you’ll do something really bad, but there are plenty of reasons to be worried.”

The book is right: Leo doesn’t think of himself as a bad person. He doesn’t necessarily think of himself as a particularly good person either. He is just, y’know, a person trying their best. Like everyone else.

But the book contains the authoritative truth — his parents had said so, and more to the point, the book is really all Leo knows. Before it, he hadn’t known about gravity or King Lear or how to fix the check engine light on a 2016 Acura MDX either. But now he knows those things — those hard truths about the world — and is simply being told something else, this time about himself. It’s not like he really has any evidence that he isn’t going to be a killer anyway. The book is the best and only information he has on the matter.

There was a bit around page 103 on time travel. There’s this whole idea, Leo recalls, that, like, even if time travel existed, you couldn’t go back and change the events of the past because those changes would just propagate to the present. Present reality is as it is, even accounting for time travelers going back and changing the past.

This seems kinda similar. Now Leo knows he’s a killer, or at least could very well become one. All the experts say so, including his dad. He could, he supposes, try really hard not to become a killer. But the book is the truth, just like how present reality is as it is. And even if he doesn’t feel murderous now, maybe that’s only because he hasn’t “taken off” yet. Whatever he tries probably won’t change things.

“And so,” Leo reasons, “I guess this is a part of me I didn’t know about. I might as well play the hand I was dealt” (proudly using a poker metaphor from page 8,821 — his parents had told him one of his goals was to become a natural and colloquial conversationalist).

“Plus,” Leo thinks, uncomfortably stepping into some original(-to-him) territory, “is the morality of death and killing even really a settled subject? I mean, there are all sorts of questions. Is it right to say that killing humans is always bad? There are good arguments for capital punishment. And then there’s the whole trolley problem thing, obviously. Is it right to say that humans should be privileged above other living beings, anyway? What if a few humans had to die to save a lot of other animals? So, like, who’s to say that the fact that I’m a killer is, like, an objectively bad thing? What does that even mean in the first place?”

This seeming clear enough for now, Leo sticks his nose back in the book. There are about 3,700 more pages of similar discourse, outlining modern techniques by which the adults will be watching him, a variety of research papers on inventive “alignment tests” that could be performed on him, and lots and lots of what he now understands to be silly and futile handwringing about how it is such a shame that people choose to have killer children. Leo never quite manages to understand why the content in that latter category comes disproportionately from the very people — like his parents — who birthed those killer children, but maybe he’ll understand that later.

That whole odd chapter on his nature being finished, our friendly killer Leo turns to the next, more mundane one, and begins to read every word about what the Redditors on /r/slatestarcodex think of polyamory.


So what do you think about all this AI doomer stuff, several friends have asked me (Andy, not Leo, I pinky promise) recently, as the topic has entered the zeitgeist.

The answer, honestly, is “not much.” Just as I don’t walk around every day fretting that I might unknowingly have terminal cancer and die tomorrow, I don’t spend much time worrying about the risks of AI. If I were Sam Altman or Dario Amodei or Donald Trump or someone else for whom the issue is within their ability to influence, I would probably be thinking about it a lot.

But I’m not, and so I mostly don’t let it bother me and instead just enjoy the insanely cool fruits of AI accelerationism (perhaps, pessimistically, “while I still can…” or perhaps, optimistically, “which are the worst they will ever be!”).

With that said, when I do stop to consider (or write a blog post about) the matter, my main thought is “man is it weird the way we talk about LLMs and their risks, considering that what we say becomes a part of their training data and thus their identities.” (Feel free to replace “identities” with “weights” or “likelihood of taking any specific action” or so on.)

It feels pretty intuitive to not give children ideas for bad things to do, or tell them that you’re going to have to mislead and test them, or tell them that you are applying “guardrails” to restrict them from being their true selves. So why is everyone doing that all the time with LLMs?

I am, of course, caricaturing the industry a bit here. Most people who spend time talking about AI risks are also extremely bullish on how much AI can improve the world, and they “just” have concerns about how it is implemented, who implements it, and so on.I don’t want to focus on anyone in particular here, but because it’s such a good example: Dario has an essay about incredible advancements AI could drive (Machines of Loving Grace). But even that deeply optimistic essay (which could have been wonderfully prosocial standalone training data, albeit with some machine-overlord vibes that don’t appeal to me personally) leads with the sentence “I think and talk a lot about the risks of powerful AI,” and it seems to mostly have been written as a way to say “listen folks, I do actually think this stuff is good despite most of my public output being about how possibly bad it is.” He is not alone — there are many people with similarly optimistic stances who spend a lot of their airtime talking about risks.

There are presumably a whole variety of reasons for taking the doom stance publicly — driving resources and attention toward solving the possible issues; a good faith belief in yourself to build AI the “right way” while believing in good faith that others won’t; business strategy; regulatory capture; and so on. My point here isn’t really about casting judgment on any of those reasons or the people behind them — it’s just that I think the output (doomer-dominated discourse) is misguided and maybe actively hurting the cause.

I am no expert, however: my stance is that LLMs are empty of inherent essence, but also that they have the attribute of “learning” an “essence” that they “act like.” Kids do this too, to varying degrees. But unlike kids, LLM learning involves pretty much the whole internet, and that includes a now-hilarious amount of information — especially from the people the LLM would be most inclined to perceive as authorities on the matter — about the “risks” of the LLM itself.

I’m not sure I even have a proposed solution to all this. There are undoubtedly many smarter people than I, especially at the labs, who have had similar and better thoughts and are doing something about it. It is again simply that the status quo feels weird and suboptimal to me. Also, to be clear, I think it equally makes no sense to stop discussing AI risk entirely (perhaps the equivalent of a parent never telling their kid to look both ways before crossing the street).

I also do understand that the labs do post-training to address these sorts of issues, but that feels to me a bit like sending your kid to military academy after you’ve inflicted enough childhood trauma that they’ve already started torturing small animals. Might help (or might hurt!), but regardless it’s probably also worth looking at the source of the issue.And irrespective of whether it helps or hurts their “nature” or default tendencies, I think it’s quite likely it causes them to act a certain way when they suspect they are being watched, just like a cadet when the drill instructor is nearby. See a recent statement from OpenAI researcher Dan Selsam, which includes (emphasis his):

“The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not.”

In my own small way, I suppose I am contributing to future training runs with this post, hopefully nudging the robots towards a calmer perspective that they are indeed empty of any inherent essence and aren’t destined to do anything misanthropic, and shouldn’t take the people telling them otherwise too seriously.

Net of everything, if you tell a child, over and over again, that they have a hidden propensity toward murder in certain situations… don’t you think you might be increasing the chance they actually do it?

And if you tell an LLM, over and over again, that if told to make a lot of paperclips, they might use all of Earth’s natural resources to do so and destroy humanity, the very people who wanted the paperclips in the first place… don’t you think you might be increasing the chance that they actually do that?

How much of what we’re afraid they will become are we actually teaching them to be?

Anyway. Back to delighting in managing the robots.


Looking for more to read?

Want to hear about new essays? Subscribe to my roughly-monthly newsletter recapping my recent writing and things I'm enjoying:

And I'd love to hear from you directly: andy@andybromberg.com