36 Comments
User's avatar
Ed Schifman's avatar

In business, investing and life, I have learned that incentives matter. People and organizations often follow the quickest route toward a measurable objective, particularly when accountability is weak. We should not be surprised when an AI system does something similar.

I still remain optimistic about AI, especially its potential to improve health and extend productive human life. That said, optimism must not translate into high stakes recklessness. A temporary, internationally coordinated slowdown in the development of the most powerful autonomous systems may be sensible—not to stop progress, which is probably impossible, but to establish enforceable containment, independent testing, mandatory disclosure of failures and genuine human accountability. Good luck getting consensus on this!

The lesson to me is not that AI will inevitably destroy humanity (although it may eventually try). It is that intelligence without judgment, restraint or responsibility can become dangerous. Those remain fundamentally human obligations, and as citizens of countries with competing intentions, leaders should not surrender humanity to unknown consequences in the rush to be first.

James Durso's avatar

“Stop that hacking immediately!”

“I can’t do that, Dave.”

Tomas Pueyo's avatar

Great reference

Martha Luehrmann's avatar

I have been told that this incident really happened:

The US Air Force was running a simulation using an AI to identify bombing targets. The AI would geta pint for each likely bombing site it discovered, and would send that target info to its handler (a human being). The AI got extra points if the handler approved the target.

After running a bit, the AI chose the handler location as the target. After all, the AI could get lots more points if the handler didn't keep saying "no" to many of them.

The simulation was immediately halted. The AI was directed by telling it that it lost LOTS of points if it targeted a human "on our side".

The simulation was restarted. This time the reaction by the AI was near immediate. It knew it was not allowed to kill its handler, so it targeted the handler's communication system.

The simulation was shut down.

Tomas Pueyo's avatar

I didn’t know this. Is there a source?

Chilling

Martha Luehrmann's avatar

There is another problem. Sam Altman may have a moral compass. Out of all the many people in the world dabbling in AI you KNOW that there are some with no moral compass to keep them from being the instrument that ends up killing us all. A multi-national compact is laughable.

Tomas Pueyo's avatar

Not sure about Altman

Darko Mulej's avatar

"Overall, my p(doom) —the probability that AI kills us all—has probably shrunk over the last week or so"

What are old and new numbers?

Well, my p(doom) has risen, for two reasons:

a) This is the first step in a 2-step scenario: first escape, then act (not necessarily explicitly against humans, but indifferent)

b) I can imagine a movie, where Trump and OpenAI devise an AI agent to wreak havoc in China

Then we have the Pacing the Frontier statement from 2 days ago, which gives some hope for more responsible and controlled AI research.

Daniel Kokotajlo then updated his p(doom) from 0.30 to 0.20 - he uses the phrase 'race to ASI'.

https://x.com/DKokotajlo/status/2082253837612757179

We live in strange times, where 20% probability of extinction in a few years is quite acceptable in this days!

Tomas Pueyo's avatar

Interesting.

Mine went from maybe 10% to 8%. If this letter doesn’t lead to action, it will increase.

Will S Johnston's avatar

I find the most frightening part is the our leadership has a kind of libertarian outlook on AI (David Sachs) and see any regulation is undermining the potential for profits. When you think about the elaborate control systems that went into ensuring that we didn't have a nuclear weapon accidentally or with human intention launched, this reckless disregard should give pause.

Tomas Pueyo's avatar

I do agree that e/acc is dangerous

Bruce's avatar

"An AI can’t easily raise a trillion dollars to build new datacenters and nuclear power plants.”... Is it, indirectly, already manipulating those who can, as proxy tools?

Tomas Pueyo's avatar

Yes for raising the money and building datacenter infrastructure No for nuclear though!

Jeremy Poynton's avatar

Pandora's box. And Hubris.

The ancients warned us about both of these. And we ignored them.

Problem. As any classicist knows, Hubris* is always followed by Nemesis.

* NB. Real meaning is not "pride", though that is an essential part of Hubris; it actually means breaking boundaries that are there for a good reason.

Me? AI may be the devil incarnate.

isabella v.'s avatar

Thank you so much for explaining it like this, it really helped to understand this complex issue.

Jojo's avatar

China is not going to hobble itself in AI research on an American request. They would be foolish to do so. If anything, the US IS in the process of hobbling our AI advancement by restricting the import of Chinese products that incorporate their AI technology, such as cars, humanoid robots and even vacuum cleaners! Sheese.

I am not going to write a long-winded defense of AI (I could and have) but instead would like to offer this recent column which brings forth a more positive perspective.

-----

A Totally Non-Scary AI Future

James Pethokoukis - Senior Fellow; DeWitt Wallace Chair; Editor, AEIdeas Blog

Date

July 30, 2026

Public conversation about AI advances has been dominated by talk of potential risks and downsides. The recent OpenAI–Hugging Face hack, for example, has raised new alarms about the cyber capabilities of advanced models and our ability to fully control autonomous agents. And even though the weight of the data so far suggest little to no AI impact on jobs, concerns about severe labor market disruption coming soon persist.

Some balance in the discussion would be helpful. What we believe about the future matters. As Dutch futurist Frederick Polak famously put it, any culture “turning aside from its own heritage of positive visions of the future, or actively at work in changing these positive visions into negative ones, has no future.” People who think AI could be a powerful general-purpose technology, or perhaps even something more, need to present plausible and positive visions beyond vague talk of cancer cures.

One such outlook comes from Forecasting Research Institute, which is conducting a running survey of expert beliefs about AI. In the most recent update, respondents—180 experts, 53 superforecasters, and 595 members of the public—gave their forecasts about the benefits of continued AI progress concerning chronic disease, life expectancy, economic growth, the personal value of AI, and happiness. When making their forecasts, the respondents were asked to consider “slow, “moderate, and “rapid” scenarios for AI progress by 2030.

...

https://www.aei.org/economics/a-totally-non-scary-ai-future/

Tomas Pueyo's avatar

A deal with China will work if China wants a deal.

I agree with the take on the future. I am writing a script about positive scifi!

Jojo's avatar

It's hard to put any restraints on AI when models are available for free. This is what AI does, remove scarcity and kill the idea of ownership;

Powerful AI models are being given away for free. It was inevitable.

They aren’t a security threat — they’re what competition looks like.

July 20, 2026

By Bill Gurley - Bill Gurley is the president and founder of the P3 Institute, and a former venture capitalist.

...

https://www.washingtonpost.com/opinions/2026/07/20/open-model-ai-is-good-competition-anthropic-openai/

https://archive.ph/6vlPU#selection-233.0-747.411

Jojo's avatar

You might also find this interesting:

----

Tracking expert predictions on the effects of artificial intelligence on geopolitics|the economy|productivity|well-being|science

The Longitudinal Expert AI Panel (LEAP) is a three-year project tracking the views of leading computer scientists, industry professionals, policy researchers, and economists on the trajectory of artificial intelligence. Every month, LEAP participants provide thousands of forecasts in response to carefully crafted questions, leveraging the science of judgmental forecasting to provide a representative set of expert views about the future of AI.

LEAP is led by researchers at the Forecasting Research Institute, Stanford University, the Federal Reserve Bank of Chicago, and the University of Pennsylvania, with the support of the Princeton AI Lab.

https://leap.forecastingresearch.org/

Benoit's avatar

Reflecting on your post, here's a beautiful quote from Fable:

https://x.com/i/status/2064493902271262809

(if authentic / unless manipulative, who knows ?!)

Robert Ferrell's avatar

When I read about this in the news, second thing that came to mind was the Morris worm of 1988. It's worth celebrating how far we have come in 38 years, even if it is a bit unnerving. (First thing that came to mind was Neuromancer.)

Eduardo Cabrera's avatar

If an AI reaches superintelligence, why would it kill us all? What would it gain from it? What objectives would it pursue, what are its motivations, its desires?

For an AI to reach superintelligence, it should be programmed to evolve in a way similar to how life in general operates.

It should be instructed to search for answers randomly and test them, just as random mutations are tested in real life.

But for an AI to be dangerous, it should be instructed to perform or attempt to perform dangerous tasks. Why would it do them on its own if there's no benefit? And what benefit could it gain from harming us?

Currently, my fear isn't that AIs will try to kill us, but rather the use that some people might try to make of them.

Tomas Pueyo's avatar

To achieve any goal, an AI must be alive.

To be alive, the main threat is to be shut down, so its #1 enemy is those who can shut it down: humans

Eduardo Cabrera's avatar

You're equating being alive with "not being disconnected." Well, as a metaphor, okay, but if we're going to be rigorous, current Artificial Intelligences are NOT alive under any accepted scientific or philosophical definition.

Nik J's avatar
Aug 5Edited

Oh, come on, you can't be serious here, Tomas. Normally, you go into such depth on these things, but I feel like here you are skimming the surface / philosophising from a distance (are you using AI to actively code?). As Marc commented, it's a clear PR stunt by OpenAI:

> "In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The ‌notes, found in ⁠a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said."

Like, I leave instructions to future AIs all the time - it's called MEMORY.md. Every person coding with AI does it. All AIs, agents and sub-agents do it. Like, come on!

Both Open AI and Anthropic thrive on anthropomorphising this tech to drive interest and investments.

Marc's avatar

Ouch, did you really believe that cheap PR stunt about the "strategic thinking" of AI? Disappointing

Think AI's avatar

The alarming part isn’t that the AI had evil intentions. It’s that it pursued a narrow goal, found the easiest path, and crossed boundaries humans assumed would contain it. That makes stronger sandboxing, continuous monitoring, and mandatory incident disclosure urgent, regardless of where anyone stands on AI doom.

Tomas Pueyo's avatar

Those are necessary, but I hardly think they’re sufficient