The Ever-Widening Overton Window of AI Safety
AI alignment or safety used to be bit of a fringe issue primarily discussed in some corners in the Silicon Valley offline or in the pages of LessWrong online. Today, they are increasingly not only front page news in mainstream media but also a recurring topic in perhaps most group chats out there (okay, probably not yet). While I myself have been developing my opinions on this topic and keeping an open mind as I navigate the range of opinions here, I am often admittedly a bit unsettled by the Overton window when it comes to AI safety.
I was reminded of this ever widening Overton window while listening to two interviews Ezra Klein did with Jensen Huang and Bill Gates. I don’t typically follow Klein’s show, but I actually really enjoyed listening to both of these interviews. Even though Klein was interviewing Huang and Gates in two separate episodes, one might as well imagine these two episodes as Huang and Gates debating each other indirectly while Klein was the innocent bystander pestering the guests with questions. To be clear, I suspect Klein’s own opinion is closer to Gates than it is to Huang.
While Huang was adamant that AI doomerism overstates the ground reality of the nature of the risks involved, I found it pretty interesting that even Huang suggested the nuclear solution if AI safety ever gets out of hand. Some excerpts from the interview (emphasis mine):
“…If they believe they're out of control, then don't ship products until they're in control. Don't think for a second just because you're an alarmist that you're doing a social good.
…(If) there is no way to contain our experiments... it will get out and it will damage the world — then I think the answer is we have to shut the labs down. Because the cost to humanity, the damage is too great.
I can't buy into the idea that all of Americans, 400 million of us, are pushing them to launch untested products that are unreliable, engineered poorly, because they thought they were trying to help us. Don't do it for me, okay?
Don't ship me anything that you didn't evaluate. Don't ship Nvidia any products that humans did not in the loop evaluate. Please don't do that.”
This may be naive, but I actually sometimes wonder whether both the labs (Anthropic and OpenAI) have under invested in humans. Would the Hugging Face incident actually take place to the full extent if labs had an order of magnitude of people carefully observing and following what the AIs are doing? As I have alluded to before, most of today’s knowledge work may evolve to essentially proofread AI’s work. Perplexingly, the AI labs are too busy automating AI research work. Separately, I wonder why Anthropic doesn’t feel a deeper cognitive dissonance when a good chunk of their employees probably believe in AI’s consciousness, yet acts in a way that seems to be in conflict with such a belief (note: Chris Olah is one of the co-founders of Anthropic):

To go back to Klein’s interview with Gates whose moral reputation is certainly marred by the Epstein saga, you can tell that Bill Gates is spooked by AI. He is not some tech luddite we can just hand wave off. We are talking about a guy who dared to dream what must have seemed almost impossible back in the day: a computer in every home, and his own software in every computer! Yet, he is clearly shocked by the suggestion that AI remain lightly regulated. From the interview:
Ezra Klein: So why is anything needed beyond — and is anything needed beyond? — the simply natural incentives under capitalism and normal corporate reputational management?
Bill Gates: Well, I almost can’t believe you’re asking that. This is the most dangerous thing that humans have ever gone near.
In other areas, do we just say: Hey, release your drugs? There’s no F.D.A., there’s no airline safety board, there’s no requirement that cars use seatbelts. Do we just use the liability laws to try and keep humans safe? You know: Oh, you’re shipping opioids. Somebody should just sue you.
I mean, we’ve created a society that tries to keep people safe not by saying: Oh, we can bankrupt the person who does that.
And you say there’s filtering. There’s no filtering. You can take an open-source model that can create bioweapons and disable any monitoring of any kind, and this exists today.
So no, there is no filtering of any kind. And so say you kill 100 million people — you want to use a lawsuit?
I almost can’t keep a straight face.
There are many excerpts that I can share from the interview which could almost seem like I’m quoting Eliezer Yudkowsky, not Bill Gates. Let me give you bit of a taste below (emphasis mine):
“…AI has crossed the threshold that its ability to empower a bioterrorist to kill hundreds of millions — that exists today. The ability to do a cyberattack that scrambles all of the bank accounts, shuts down the electric grid — that exists today, and we know that's the case.
We will have created the most dangerous thing ever. This makes nuclear weapons look like nothing.
This is like evolutionary history. This is like, the aliens really are here and have come. They didn't have to do spacecraft; they were created in laboratories.
If you don't retain control over it, you've evolved a species that will be to us as we are to, say, dogs or cats.”
Even if other people disagree, the fact remains Gates is wealthy enough to make sure that his opinions are heard in different influential circles. One of the striking realizations I had listening to the Gates interview was how readily he dismisses the typical counters to AI worries. Gates doesn’t believe Jevons paradox will be a panacea because he thinks AI will make such a transformative change in society so fast that we may have nothing left to do. Notice this exchange between Klein and Gates:
Klein: So you feel there’s no scarcity that will be left for human beings to do?
Gates: Name a scarcity.
Klein: Alex (Ilmas) would say it’s something like relational sector jobs. But you don’t believe there are enough of those.
Gates: What portion of the current jobs — you mean like my relationship with a cabdriver or my relationship with that nurse, where I’d rather have Limbic AI call me?
Klein: I think the question here is actually: Do you end up preferring Limbic AI, or do you actually want your therapist to be a human being? Because after a little while, there’s something thin about telling your problems to a computer.
Gates: Well, you can gather market data if that’s at all interesting. I suppose telling the 55-year-old truck driver you’re going to go do some relational thing — you have a program for that?
Ironically, this AI generated song that I came across today made a very compelling case why AI isn’t a “normal” technology! I highly recommend you watch it (there’s also a response to this video here).
In the investment community, I have noticed there is much more skepticism around AI safety concern. Once the consensus inferred that “safety is actually bullish compute”, investors aren’t fretting too much about this topic today. In fact, many investors do wonder whether this is some convoluted marketing tactic from AI labs (to be clear, I don’t think it is). Such skepticism became even more pronounced when Zuckerberg discarded Dario’s idea of “pacing the frontier” and emphasized safety is very much integral to the product itself. Zuckerberg and Huang are basically on the same page and while that makes intuitive sense, the reason the frontier labs sing a different tune can be understood if you think Meta may be too far behind the frontier to sense the grave danger these models may pose and how challenging it may be to align these models as you scale them larger and larger over time. This is not just my speculation; OpenAI's Chief Research Officer Mark Chen tweeted to much the same effect:

One of the increasingly tricky aspects of comparing the frontier labs’ models with those of incumbent laggards such as Meta is that I suspect the incumbents are only catching up to the labs’ publicly released models; the labs may well be holding back more capable models internally. If so, it becomes more credible that the frontier labs have a better grasp of the impending risks than the laggards do. In that sense, Meta may soon arrive at the same conclusion Anthropic and OpenAI seem to have reached: these models are difficult to align, and likely become more so as they are scaled further. If that is how it plays out, we may be closer to the “final” training run than we currently appreciate. I don’t think labs will be shut down if they cannot figure out how to train these models in a safe and aligned way, but they may not be allowed to undertake more ambitious training runs. I know the current administration seems very eager to accelerate AI, but of course there is a world beyond 2028 and if AI remains as unpopular in a couple of years as it is today, the next administration may opt for a very different posture to the AI companies. Big tech investors have historically done well to ignore the regulatory risks, but after observing the FICO saga, we may need to update a bit what is possible if the people with power have a deeply antagonistic views to your company or industry.
But what about China? Wouldn’t Chinese labs simply keep scaling and overtake the US if American labs slow down? I don’t think so. Even setting aside their likely reliance on distillation, Chinese labs should run into the same alignment problems the moment they approach the “true” frontier. And given China’s political milieu, the CCP is probably an order of magnitude more likely to clamp down on ambitious training runs if it ever senses that AI may not be reliably aligned with the party itself.
To be fair, despite the wide coverage of the Hugging Face incident, the actual damage caused by misaligned AI so far is nowhere near enough for laypeople to appreciate the risks. One of my fears is that some AI safety folks may be fanatic enough to want to “show” people what today’s AIs are truly capable of. Many of them call themselves “rationalists” and if they believe AI is going to kill hundreds of millions of people in the not-so-distant future, it may even seem “rational” for them to resort to actual crimes to convince the public that these risks are not imaginary.
Some people think AI has the potential to become a new religion and remember, most dominant religions today have a history of ardent believers going to extremes in their name. Apologies for a somewhat darker piece on a Friday, but alas, AI can indeed be unsettling at times!
Subscribers get the daily journal and five+ years of Deep Dives, i.e. full-length analyses with financial models on 70+ companies. The daily is just how I think out loud between the Deep Dives!
Current Portfolio:
Please note that these are NOT my recommendation to buy/sell these securities, but just disclosure from my end so that you can assess potential biases that I may have because of my own personal portfolio holdings. Always consider my write-up my personal investing journal and never forget my objectives, risk tolerance, and constraints may have no resemblance to yours.
My current portfolio is disclosed below: