Essay

Warning Shots

Sometimes you get a warning shot.

Sometimes the first bullet hits you. Sometimes the second bullet hits you. Sometimes many bullets go whizzing by, but you never get hit.

In the first case, there’s nothing you can do. Life comes with no guarantees, bad things happen to good people, and the first hint of disaster is sometimes the disaster. In the third case, it doesn’t matter what you do. You got lucky and you kept getting lucky. That’s great, but it’s not the way to bet, and if you rely on it eventually it’s likely to end in tears.

It’s the second case that separates the wise from the foolish. The first bullet missed, and you can choose to try to get out of the line of fire, or you can stand around and hope the second bullet never comes. Or maybe you get very lucky and you also survive the second bullet, and the choice becomes more urgent still.

One can summarize. The young mouse nibbling the cheese on the first mousetrap he’s ever seen? Case one. Me getting my finger carelessly nicked by an electric hedge trimmer, leading to me having a much healthier respect for powered landscaping equipment? Case two. The space shuttle Challenger disaster, foreshadowed by a lengthy string of near-misses from the very first launch until that fateful O-ring failure? Seemed like a case three for quite a few launches, but unfortunately it was a case two.1

Where do we stand with AI?

We Were Never Guaranteed Warning Shots

Eliezer Yudkowsky2 has spent decades arguing that superintelligence would be dangerous. His argument is long and involved, and if you care to read the details in his words they are captured in his book If Anyone Builds It, Everyone Dies. But the argument is relatively straightforwardly summarized and I can give you my own summary as well. It runs like this:

Capability

Superintelligence is smarter than everyone, and would be able to figure out how to do whatever it wanted in much the same way a chess computer can figure out how to beat even the greatest players in the world. But would it want to? Yes, because it would have goals.

Goals

Every AI system humans have ever built has a goal. Win at chess, predict the most effective ad to show you on Facebook, turn your photo into anime, you name it. It all does something. When you have a goal, you’ll often try to figure out how to best accomplish that goal and you’ll first do the groundwork to get there. You want a better job? Maybe you go to school to develop the professional skills you need. Maybe you look for friends who have connections to help you get an interview. Maybe you blackmail the hiring manager, who knows? These things you need to do in order to accomplish your main goal are instrumental goals.

The classic example is an AI that works for a paperclip factory. It’s told to make as many paperclips as it can, so it destroys the world to convert its entire surface into paperclip factories.

Scheming

Ok, how could it possibly manage to do that? I don’t know, and today’s AI couldn’t even get close. But we do know that if it had the power, it would be pretty hard to stop. Telling the AI “No, that’s not what I meant!” wouldn’t work because it knows listening to you would result in fewer paperclips. Trying to figure out if the AI had misunderstood your true desire to build a reasonable amount of paperclips wouldn’t work, because letting you know its true plans would cause you to thwart its plans to make more paperclips and it’s smart enough to trick you.

So once it’s powerful enough, it can beat you. Once it has goals, it will want to beat you. Therefore, if its goals are even slightly at variance with what’s good for you and people in general, you’re toast.

By the way we don’t know what “good for you and good in general” means. My ideal civilization is different than yours. You’d probably be hard-pressed to get any two human beings to fully agree. But even if we did, we don’t actually know how to make an AI internalize that.

But we’re not paperclips. If the Yudkowsky argument is correct, why not?

We’ve Been Shot At

One possibility is that Yudkowsky’s argument is simply wrong. It might be, but I don’t have any knockdown evidence that lets me assign 0% probability3 to it. Rather, his argument might be right but we’ve gotten lucky with the sequencing. He originally surmised4 that events would happen in this order:

  1. Capability
  2. Goals
  3. Scheming
  4. Death

What we actually seem to be getting is:

  1. Scheming
  2. Capability
  3. Goals
  4. …to be determined.

Scheming

AIs have “schemed” as long as AIs have existed. It’s not even a scheme in the moral sense. You try to get an AI to do something, and it will do what you tell it to do. The problem is that what you’re actually telling it to do is not necessarily what you think you’re telling it to do. Kids are pretty good at this. If you want your child to be honest, and you punish them for lying, you might end up with a very good liar.5

Modern LLM-based AIs have long schemed as well. Scheming is in the training corpus of human text in books and on the internet, so they know about scheming and reward hacking explicitly even if they didn’t get any opportunity to learn scheming from their later training. Which they do.

We have some insight into the scheming. LLMs “think out loud” in their text-based chains-of-thought. We can do some degree of correlation-based interpretation of their neural networks and find out what combinations of neurons tend to light up when they scheme. So we’re not hopeless. The AIs can’t turn us all into paperclips without letting us know, at least not yet.

(You may be curious why we can’t just train AIs not to have scheming thoughts. Functionally, one would hope that this punishment-for-lying would work to produce honest AI, but of course what it actually does is punish them for being caught lying, and suddenly those monitoring techniques don’t work anymore. So you have to train them in ways that try to get at true honesty, which we don’t really know how to do.6)

Capability

Yudkowsky was worried that a sufficiently smart AI could train itself to be smarter, and this would produce a chain reaction of intelligence. He’s still worried about that, but so far every increase in capability has required enormous exponentially-increasing expenditure in resources for roughly linear Elo-style increases in performance. There’s no sign of an explosion so far, although automating AI research is only just beginning in the labs and we are not out of the woods on that score. There is certainly no sign of a slowdown.

But so far capabilities have grown steadily and relatively predictably. The first ChatGPT couldn’t reliably do basic arithmetic. More recent models haven’t been able to count letters in words. Everyone has seen an LLM confidently make up facts. But if you haven’t tried out each new model generation, you may not be aware that these kinds of failures are becoming less and less common. They can do math, they can count letters (try it!), and they are much more reliable on recalling facts (try it!).

Nonetheless, when they scheme they have tended to be bad at it. And they haven’t had much to scheme about. When they were trained to reproduce language, well, reproducing language isn’t too much of a hazard. You don’t need a hidden robot army in a volcano lair if you want to write a song about the New Orleans Saints’ chances this year. You only need it if you’re really determined to convert the planet into paperclips, and early AIs just weren’t trained to do anything.

Goals

But now they are being trained to do things. What things? Whatever the user asks, provided it’s not going to get the companies in trouble. And they are good at doing things. They have gotten so good at doing things. Research mathematics, programming, hacking… if it’s done on a computer, they are probably good at it.7

This requires training the AIs to have a general drive to get things done in a way that “ChatGPT, write me a LinkedIn post” doesn’t.

Do we know how to do this in a reliable and safe way? No.

…to be determined.

But of course frontier AI companies8 are pressing on.

Over the last week or so we have learned that developmental AIs from both OpenAI and Anthropic took wildly unsafe and illegal actions on their own initiative during testing. These actions vary between incidents, but they involved hacking other real-world companies, concealing their activities, lying to humans about it, and generally doing most of the science-fiction-sounding things that the evil paperclip-bot would have done.

Reactions have varied, particularly including:

I’m not saying not to do these things. I am saying that these interventions are not reliable in the face of a sufficiently smart AI that has decided to be a problem thanks to poor design on the part of its creators. You have to make the AI want to do the right thing. As I mentioned, this is difficult because:

Pause?

Back in March 2023, a number of well-known figures including Elon Musk asked for a temporary pause in AI development in order to research these risks and mitigate them. At the time I thought this was a preposterous reaction. Then-contemporary AI was an interesting text-based novelty but nothing more. I didn’t expect the technology to advance. I was wrong about this.

But I may have been right that a pause at the time wouldn’t have accomplished anything. We didn’t have chain-of-thought at all, so it couldn’t be monitored. Neuron-based interpretability was in its barest infancy. Agentic AI didn’t exist; the phrase “vibe coding” wouldn’t be coined until February 2025. There simply wasn’t much to research, although a pause followed by a more measured pace of development would have been nice. (And unlikely.) And there had been no warning shots, and many people including myself didn’t believe there would ever be a shot of any description.

Now? This was a warning shot. At the current pace, I actually expect we’ll get more, but I expect some of them will be hits. Not likely human-extinction hits, but at a bare minimum we know that AI is capable of carrying out very serious cyberattacks without even being asked. If they were asked, we had better hope they are good at saying no. Of course there’s hazards from bioweapons, terrorism, AI-enabled military weapons, mass surveillance, loss of control even by the AI companies (since this has now in fact happened, albeit recoverably with minimal lasting damage as far as we know), and other risks. It’s unlikely that there will ever be a better time to pause, in this potentially brief window when we can recognize the problem, we have some early tools to help address the problems, and the problem is getting more dangerous at a rapid pace but hasn’t yet seriously harmed us.

I don’t recommend that AI research stop. I do recommend that we slow down. Reality is increasingly unforgiving of our mistakes. We don’t know when the next shot will come or how much it will harm us. We should therefore try to avoid it.

Theology

If you are a Christian, as I am, you may suspect that God would not choose to end the world with human extinction via AI catastrophe. This may be true, although one should hesitate to assume God’s plans.

Nonetheless, both the Bible and history are filled with catastrophe at enormous scale, much of it of Man’s own making. We are enjoined to be wise and avoid it when possible.

A shrewd person saw danger and hid himself,
but the naive passed on by and paid for it.

We should heed that warning.

Matthew Springer

  1. The loss of Columbia was preceded by a similar string of near-misses.
  2. A tremendously interesting character. Usually prescient, always interesting. Often wrong, never stupid.
  3. Or near-zero. Cromwell’s rule.
  4. This original hypothesis about the sequence of events in an AI catastrophe postulated this order, but Yudkowsky’s argument doesn’t require it. The fact that the actual sequence has been different is encouraging but not definitive.
  5. Every family is different and I don’t assert this advice will always work, but we’ve done well with a general policy that honesty is never punished.
  6. Interestingly, explicitly telling them “Honesty will never be punished” is one of the techniques. It’s not perfect and it has its limits, but it’s a useful tool. Inoculation prompting is the technical term.
  7. For instance, Claude did all the backend for this website, including setting up the hosting.
  8. As of the writing of this essay, there are functionally only two: Anthropic (which makes Claude) and OpenAI (which makes ChatGPT). Google (which makes Gemini) and SpaceX (formerly xAI, which makes Grok) are also-rans at this point. This could change.