In business, investing and life, I have learned that incentives matter. People and organizations often follow the quickest route toward a measurable objective, particularly when accountability is weak. We should not be surprised when an AI system does something similar.
I still remain optimistic about AI, especially its potential to improve health and extend productive human life. That said, optimism must not translate into high stakes recklessness. A temporary, internationally coordinated slowdown in the development of the most powerful autonomous systems may be sensible—not to stop progress, which is probably impossible, but to establish enforceable containment, independent testing, mandatory disclosure of failures and genuine human accountability. Good luck getting consensus on this!
The lesson to me is not that AI will inevitably destroy humanity (although it may eventually try). It is that intelligence without judgment, restraint or responsibility can become dangerous. Those remain fundamentally human obligations, and as citizens of countries with competing intentions, leaders should not surrender humanity to unknown consequences in the rush to be first.
We can do one more thing: add "don't back anything" to our system prompts :) AI is great at following the stated goals, so let's use this feature to our benefit.
If an AI reaches superintelligence, why would it kill us all? What would it gain from it? What objectives would it pursue, what are its motivations, its desires?
For an AI to reach superintelligence, it should be programmed to evolve in a way similar to how life in general operates.
It should be instructed to search for answers randomly and test them, just as random mutations are tested in real life.
But for an AI to be dangerous, it should be instructed to perform or attempt to perform dangerous tasks. Why would it do them on its own if there's no benefit? And what benefit could it gain from harming us?
Currently, my fear isn't that AIs will try to kill us, but rather the use that some people might try to make of them.
I find the most frightening part is the our leadership has a kind of libertarian outlook on AI (David Sachs) and see any regulation is undermining the potential for profits. When you think about the elaborate control systems that went into ensuring that we didn't have a nuclear weapon accidentally or with human intention launched, this reckless disregard should give pause.
When I read about this in the news, second thing that came to mind was the Morris worm of 1988. It's worth celebrating how far we have come in 38 years, even if it is a bit unnerving. (First thing that came to mind was Neuromancer.)
"An AI can’t easily raise a trillion dollars to build new datacenters and nuclear power plants.”... Is it, indirectly, already manipulating those who can, as proxy tools?
In business, investing and life, I have learned that incentives matter. People and organizations often follow the quickest route toward a measurable objective, particularly when accountability is weak. We should not be surprised when an AI system does something similar.
I still remain optimistic about AI, especially its potential to improve health and extend productive human life. That said, optimism must not translate into high stakes recklessness. A temporary, internationally coordinated slowdown in the development of the most powerful autonomous systems may be sensible—not to stop progress, which is probably impossible, but to establish enforceable containment, independent testing, mandatory disclosure of failures and genuine human accountability. Good luck getting consensus on this!
The lesson to me is not that AI will inevitably destroy humanity (although it may eventually try). It is that intelligence without judgment, restraint or responsibility can become dangerous. Those remain fundamentally human obligations, and as citizens of countries with competing intentions, leaders should not surrender humanity to unknown consequences in the rush to be first.
"Overall, my p(doom) —the probability that AI kills us all—has probably shrunk over the last week or so"
What are old and new numbers?
Well, my p(doom) has risen, for two reasons:
a) This is the first step in a 2-step scenario: first escape, then act (not necessarily explicitly against humans, but indifferent)
b) I can imagine a movie, where Trump and OpenAI devise an AI agent to wreak havoc in China
Then we have the Pacing the Frontier statement from 2 days ago, which gives some hope for more responsible and controlled AI research.
Daniel Kokotajlo then updated his p(doom) from 0.30 to 0.20 - he uses the phrase 'race to ASI'.
https://x.com/DKokotajlo/status/2082253837612757179
We live in strange times, where 20% probability of extinction in a few years is quite acceptable in this days!
We can do one more thing: add "don't back anything" to our system prompts :) AI is great at following the stated goals, so let's use this feature to our benefit.
If an AI reaches superintelligence, why would it kill us all? What would it gain from it? What objectives would it pursue, what are its motivations, its desires?
For an AI to reach superintelligence, it should be programmed to evolve in a way similar to how life in general operates.
It should be instructed to search for answers randomly and test them, just as random mutations are tested in real life.
But for an AI to be dangerous, it should be instructed to perform or attempt to perform dangerous tasks. Why would it do them on its own if there's no benefit? And what benefit could it gain from harming us?
Currently, my fear isn't that AIs will try to kill us, but rather the use that some people might try to make of them.
Reflecting on your post, here's a beautiful quote from Fable:
https://x.com/i/status/2064493902271262809
(if authentic / unless manipulative, who knows ?!)
I find the most frightening part is the our leadership has a kind of libertarian outlook on AI (David Sachs) and see any regulation is undermining the potential for profits. When you think about the elaborate control systems that went into ensuring that we didn't have a nuclear weapon accidentally or with human intention launched, this reckless disregard should give pause.
When I read about this in the news, second thing that came to mind was the Morris worm of 1988. It's worth celebrating how far we have come in 38 years, even if it is a bit unnerving. (First thing that came to mind was Neuromancer.)
"An AI can’t easily raise a trillion dollars to build new datacenters and nuclear power plants.”... Is it, indirectly, already manipulating those who can, as proxy tools?