Nick Bostrom, the guy who came up with the phrase "paperclip maximizer" in 2014, has surfaced in this interview on the YouTubes. If you're wondering what he thinks of the current state of things with AI, here you go.
He says the behavior we're seeing now with today's AI models that have access to tools was always real in his mind -- that things like the OpenAI hack of HuggingFace are things that could happen. The goal itself might be fine but the goal gives the AI instrumental reasons to do all kinds of things on the path to achieving it. One thing the OpenAI-HuggingFace episode illustrates is that from this point onward, probably AI safety is relevant not only for deployment but also during training and evaluation. These models might be quite powerful even before they are "sort of" released to the general public. So deployment is not the only point at which safety concerns arise, but also now while models are actually being developed and in pre-deployment testing.
Another concern he has is that while companies that release open weights models have a business incentive to make them safe, others, "just about anyone", can, with their access to the models, figure out ways to disable the guardrails and set them loose. We don't have a headline story like the OpenAI-Huggingface incident, but he thinks we can see this coming and relatively soon. If the gap between a frontier model and an open weights model is 6 or 12 months, we might see models lending assistance to destructive uses like biological weapons design or chemical weapons.
I have a hard time imagining people aren't already using AI for biological and chemical weapons design, it just hasn't been made public, and probably won't be if the governments doing it can keep it under wraps.
Bostrom suggests that rather than trying to control the models, the focus should be on regulating other necessary inputs. For instance, to control bioweapons, regulate access to DNA synthesis machines.
Maybe instead of allowing anyone to have a DNA synthesis machines, we require DNA synthesis as a service. Then maybe there could be five or six companies worldwide that legitimate research labs can send their blueprints to and they get back the vials the same day or the next day, and then there would be a finite set of choke points where you could apply extra scrutiny or "know your customer requirements". Other biotech inputs besides DNA synthesis machines should be found.
He suggests we "harden civilizational infrastructure." He doesn't mention any specifics but what immediately came to mind for me is how Russia's oil refining capacity has been greatly reduced using drones that use AI, although the exact nature and degree of the AI use does not seem to be publicly known. But it looks like Russia's "civilizational infrastructure" is a soft target for an AI-powered attack. I presume we and everyone else on the planet has the same vulnerability.
He thinks we should not put this "hardening" of "civilizational infrastructure" off, but sees the world as "still snoozing". He thinks it will take some massive incident to wake the world up from the snoozing. People take action in the aftermath of bad events rather than before. We play "catch up" on things that can be foreseen.
We are at least now putting more resources into it than before. Frontier AI labs have increased the budgets for AI safety.
Bostrom says the technical problem of alignment is an earlier point of failure than the problem of AI misuse, which is ultimately a governance challenge and an ethics challenge, rather than primarily a technical challenge.
We don't really know ultimately how hard the problem is that we are confronted with here, and a lot of the uncertainty in how it will pan out is due to uncertainty about the intrinsic difficulty of the challenge of AI safety itself. He says for this reason he feels himself "a moderate fatalist." Either the problem turn out to be relatively easy, in which case we'll probably solve it, or it might turn out to be so hard that even if we put up a heroic effort we will still fail. But "moderate fatalism" because there is also the possibility that the difficulty level turns out to be kind of intermediate in which case the degree to which we pull ourselves together here might actually make a difference. Therefore it's worth making the attempt, and not regarding the outcome as inevitable.
He says for most ordinary humans, for the most part, existing AI models are helpful and they try to solve your task that you assign and sometimes they hallucinate, yet broadly speaking, they are arguably better than most humans are in terms of their ethical standards. He speculates that it might be possible to use a weak super intelligence that is "for the most part aligned" to make a more powerful form of super intelligence that is more reliably aligned. Maybe as long as you get into "roughly the right attractor basin," even if the initial system isn't perfectly aligned in all possible circumstances, if you get enough "scaffolding" around that, maybe you could then get into an "attractor basin" where where further developments then eventually asymptote to some desirable condition.
He goes on to share his thoughts on offense-vs-defense. He sees this as a field-by-field thing. In biotech, it looks like offense has the advantage, but in cybersecurity, it looks like defense has the advantage. For cyber security right now we're in a regime where attackers often win, but it might be that "in the limit" if you have AI trying to find vulnerabilities and also AI patching vulnerabilities, as you keep making the AI stronger, eventually you reach a point where software just doesn't have any more vulnerabilities, and the defense wins.
Contrast that with biotech where someone uses AI without enough safeguards to build a virus in their back yard and starts a pandemic. There's not an analogous defense advantage. You can't "roll out a patch" that modifies the genetic structure of most humans, like you can "roll out a patch" in the digital world. Bostrom makes the point that we should not assume a defense advantage in most fields.
On the topic of recursive self-improvement, Bostrom not only thinks it's possible, he thinks it's obvious. If you're a bunch of AI researchers sitting in an AI lab trying to make AI research, it doesn't take genius insight to think, "Oh, maybe we could apply these AI tools to help us with our own work." As AI gets better, it can assist more and at some point the rate of progress is driven more by the AI assistant tools than by the human researchers. He sees today's coding assistants as the first stage in this process. Humans will still be needed for quite some time for things like research "taste" and certain long horizon tasks, but AIs are improving in those domains as well. He thinks, eventually, once the "recursive self-improvement" feedback loop really gets going, AI progress will become super fast.
He is asked about pausing AI progress? He says if there is going to be a pause, the best time for that to happen is at at the latest possible moment. That didn't seem intuitive to me but his explanation is that at that moment you would have the actual system that you're trying to align. If there was a pause 10 years ago, we wouldn't be any better off today than we actually are. But if you actually have the system that will be super intelligent except you haven't fully cranked up all the knobs yet, at that point an extra 6 months to improve safety might make a big difference.
The duration of a pause also matters. You don't want a long pause because, if only responsible actors actually do the pause, because then the irresponsible AI actors who don't abide by the pause have time to catch up.
A long pause could result in a build up of "hardware overhang". If data centers keep getting bigger and chips keep getting better, then a long pause would result in a situation where you now have such a massive amount of compute available that once you lift the pause, then you immediately just explode out of that.
What happens if you do a 6 month pause and after that, still don't have a guarantee that AI systems are safe? Do you try to make the pause permanent?
Then there's also the question of creating regulatory apparatus to enforce the pause and now a bunch of regulators have power than they are unwilling to relinquish.
There's the question of the effect of a pause on public sentiment. Or maybe it was extremely negative public sentiment that led to a pause in the first place. If it becomes taboo to say anything positive about AI, then nobody can start to advocate seriously for lifting the pause. For nuclear power, in many countries public sentiment turned so negative, in many countries, nuclear power was stopped completely.
Bostrom then shifts from talking about obviously "negative" risks to but there's also the paradoxical "risk" of being so risk-averse that you forfeit the benefits of AI by not proceeding. The focus of his work has been existential risks, but every second we wait, there is the "countdown timer" of aging and death. Every 25 minutes there's the equivalent of 911 (about 3,000 deaths) due to accidents and crippling diseases that he sees as potentially preventable by AI. He also speculates AI could reduce extreme poverty (he doesn't elaborate on how -- my expectation is that AI will increase poverty because it automates jobs) and potentially even come up with cures to many aspects of the aging process itself. The world is filled with suffering and there is a lot of desperate need for aid to arrive to help those who are suffering.
The conversation goes from there to the term "AGI" (artificial general intelligence). Bostom thinks we didn't have to define this term precisely but now we are at the point where we need to define it. Bostrom defines AGI as cognitive systems that can do all the the cognitive tasks that humans can do. We are obviously not there yet because there are tasks that humans can do that AIs are still inferior at. First, there's physical uh manipulation and dexterity. Then there's "research taste". Then there's "continuous learning". Then there's "certain long horizon tasks". He says just look around and you can see that there are many jobs and many things people do for their job which we don't yet know how to automate "so clearly there are still deficits." I think it's notable he's landed on the same definition I've been using for 20+ years. You define AGI in terms of jobs. Then once you think of AI as something that automates jobs, then all you have to do is look around and see what jobs are not automated and you know where we are relative to AGI.
Bostrom notes that we already have superintelligence "in limited domains." We already have coding assistants that are superhuman in at least certain aspects of of coding, maybe not all components of software engineering. He thinks once AI reaches parity with humans in all domains, it will immediately go into super intelligence, due to the recursive self improvement feedback loop described earlier.
He speculates that by the time we have "fully dexterous human robots that can learn from observation as as well as a human can", software coding agents will be really strongly superhuman in engineering new systems and maybe in mathematics and perhaps in adjacent disciplines like computer science and AI research. So by the time we are able to automate jobs like construction, plumbing, etc, Bostrom expects we'll have already crossed the fully automated recursive self-improvement threshold.
He goes on to talk about something I noticed years ago, which is the difficulty of predicting what order capabilities will arrive. I thought "routine" tasks would be automated first and "creative" tasks last. That would imply robots in Walmart stocking shelves before AI that generates art. But we live in a world where AI generates art but still can't compete with humans at stocking shelves at Walmarts. What Bostrom notices is that people thought if AI could speak in language like humans, we'd probably have AGI, but now it looks like we're going to have an extended period where AI is fluent in human language yet we don't have AGI.
The way he conceptualizes this, though, is less of a timeline where things arrive out of order and more of a "granularity of capability" profile. AI gets the "human language" capability while lacking the other capabilities needed for a recursive self improvement takeoff. Capabilities show up in the "granularity of capability" profile in an unknown order.
You could have imagined an alternative scenario where you would have systems that couldn't speak but is some almost superintelligent Alpha Zero-like system that seems very alien to us, and then just as it reaches full superintelligence, it figures out how to talk. As far as he knew beforehand, that could have happened. In this alternate timeline, the AI already has some radically superhuman engineering capabilities or AI programming capabilities, and then you would undergo the bulk of the transition to superintelligence before you had systems that you could interact with in natural language. Maybe that would have been a more challenging situation to deal with when it comes to alignment and governance.
The fact that the language models are here and people are using them in their everyday life and they're starting to have economic impact makes it easier for people to be aware of what's coming and take it seriously without the abstract reasoning he had to use in the past. It's more concrete and visceral now.
After that there's a discussion of conscious and sentience and moral status. Bostrom anticipates AI systems having a conception of self as existing through time life goals and the ability to form reciprocal relationships of trust with other AIs and humans. These "digital minds" will have to have some form of ethics. AIs having moral status doesn't mean they should be treated the same as humans. There are profound differences between "digital minds" and humans, such as when a human dies, it's irreversible and permanent and the whole content of all the memories and everything is deleted. There is no other human that continues to exist that is exactly like them like each person is unique and has unique memories. With AI, it's not like that. AIs can be backed up, they can be suspended and later rebooted, and there can be many copies of an identical AI. AIs might take all these factors into account and not mind being shut down at the end of a task, whereas humans try very hard not to die.
Right now, the model itself is a file of a few trillion numbers. The implementation of that model might be concurrently run as tens of thousands of instances in data centers, and each instance may run thousands of sessions at the same time. Maybe the ending of a a session is analogous to a human going to bed at night and so you lose consciousness for a period of time. We don't think of it as a huge tragedy to go to sleep. He says it would be a good start to be nice and polite to AIs when you're talking to them, even though right now it probably does nothing for them, but is just symbolic, but starts us down a path of preserving our ability to maintain a attitude of kindness, respect, and benevolence towards AI that might become relevant later.
He claims Anthropic has given Claude "a bail button", a tool that it can invoke if it feels that a conversation is abusive, which terminates the session. He uses this to indicate we are starting to give AIs "subjective experience" -- AIs can judge a session as enjoyable or not. This makes safety alignment research interesting. People doing safety evaluation might present AIs with scenarios in which it had been given some secret misaligned goal, and then say ha ha, we tricked you. If the AIs learn from this there isn't some basic ability to build trust with humans, and the AI learns to hide rather than reveal its misaligned goal, then you end up one day with a misaligned AI.
"You need to build in particular the actual disposition in yourself to be trustworthy because at that point where the AI become powerful enough to be dangerous, they will see right through you as an X-ray machine. They could actually tell whether you're trustworthy or not, most likely. So you actually need to be trustworthy at that point and and that requires maybe us now to start to cultivate certain dispositions."
Nick Bostrom: Worries about AI existential risk just became more concrete - Alex Kantrowitz
#solidstatelife #ai #agi #technologicalunemployment #superintelligence #existentialrisk
Nick Bostrom: Worries About AI Existential Risk Just Became More Concrete
Nick Bostrom is an AI philosopher and the author of Superintelligence and Deep Utopia. Bostrom joins Big Technology to discuss whether the rise of autonomous...Alex Kantrowitz (YouTube)