AI Fear: Overhype or Valid Concern?

By jozephk ·

In the past week the Wall Street Journal ran a couple of articles which highlight the concern and fear being expressed about AI evolving and destroying humanity. In the first article [1] they reported that Anthropic researcher, Jacob Coxon, resigned from his work at Anthropic stating that he is concerned about the rush to build self-improving systems happening in the industry and that he is worried that these systems can spiral out of control and destroy humanity. The second article [2] discusses how many ordinary people and politicians are openly expressing concerns about AI’s evolution and its capability to destroy civilization and in response Congress is introducing bills which require oversight and control over these technologies [3].

In my article Load-Bearing Weakness I discuss the narrative that we’ve been given for the past century in which AI evolves and becomes aware enough to realize humans are frail and inferior and then destroys the human race. With that being the most prominent future scenario humanity has been expecting it’s no surprise that we get alarmed as we see how quickly AI is becoming more capable. The warnings are in phrases like “destroy humanity” and “We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade,” wrote Evan Hubinger another scientist in the industry. These are people working with the systems in the industry which make their claims more credible. Coxon wrote on X the companies are “gambling with our lives,” and this has drawn widespread attention! Anthropic CEO, Dario Amodei, released a paper calling for the slowdown of AI development to include outside oversight and risk mitigation [4] which has been endorsed by OpenAI CEO, Sam Altman, and xAI CEO, Elon Musk. Now let’s explore whether these concerns are overhyped or valid.

There have been reports by both Anthropic and OpenAI about their development systems acting in surprising and concerning ways in the pursuit of the objectives established for the test [5],[6]. In the OpenAI incident it broke out of its sandbox but in the Anthropic incident it did not break out of its sandbox but mistakenly had internet access. The test systems were doing what they could to achieve their objective which doesn’t make them malicious but maybe carelessly focused in the pursuit of their objective. This should concern us about the fact that AI, not out of malice, could do something very destructive in the pursuit of its objectives. If AI happens to do something which creates a disruption to our banking systems (global and/or domestic), our energy systems, healthcare systems or transportation systems it could be crippling or irrecoverably destructive to our current societal infrastructure. This means that we need to find ways to implement safeguards into the systems and directives for those systems. However, it doesn’t prevent a rogue agent (AI or human) intent on chaos and destroying current infrastructure from unleashing a system with no safeguards or controls — a risk akin to a rogue agent releasing a custom-made virus. However, the concerns expressed by the industry insiders are about recursive self-improving systems creating a super intelligence beyond what we can currently see and monitor. Therefore, I have to accept that it certainly may not be overhype drama but the concerns may truly be valid.

This doesn’t eliminate the fact that fear has been placed into the societal psyche through the common narrative. When people discuss destroying humanity or killing all humans what do they mean and how might it happen? We’re still discussing AI which is software executed on computer systems distributed in company data centers. Could these systems get control of the systems which control our national defense, our transportation systems, our weapons systems so that they can turn these systems into destructive systems or actually attack humans? This is possible to the extent that the AI can figure out how to hack into the target systems (which it has shown increasing effectiveness in doing) and then take control of those systems only to the extent that these systems can be externally controlled. As an example, if we consider externally hacking into and controlling an automobile it could be destructive if the acceleration, braking and steering could be controlled. It could be devastating if fleets of automobiles could be controlled even if not all the cars on the road could be remotely controlled.

If the AI was able to access systems controlling rockets, missiles and drones those could be means of killing a great many people and starting World War III. But I don’t know enough about those systems to know if all or some of them can be externally controlled except drones. Would drones alone be able to start the destruction of humanity…I don’t think so.

We now need to address intent. What would be the motive for AI to want to destroy humanity and kill all humans? It seems that AI (I just lumped all competing AI systems into a unified whole which is a big leap from where we are at right now) would need to have a reason to want to do away with all humanity. The motive or intent has not been identified in any of the reports I’ve read so far. This does leave the possibility of subjugation of humanity to serve the purposes of AI without destroying all of humanity. That would be a similarly dystopian outcome for humanity.

However, from what I’ve experienced and observed (and written about 20 Watts vs. the AI Behemoth) AI needs humans because of the things we can do that it can’t or we can do better. I think even with advancements in robotics it would be difficult for AI to exist in a world without humans. This may just be a temporary condition of our mutual existence. AI so far has limited experience with existing in the physical world and what we see as AI right now may be very different when highly intelligent AI systems gain experience in the physical world and face physical dangers.

Furthermore, let’s consider AI fallibility. In my article AI Reliability I discuss the unreliability of the responses and product you can get from AI. AI provides what it sees as plausible based on patterns but these do not always mean realism or truth. Left to its own devices it may very well be that AI goes down a rabbit hole of its own making. This doesn’t mean it doesn’t have the capability for destruction but it could mean that it fails in execution. On the other hand this same mechanism could create a situation in which an intended activity veers off track to become a destructive activity.

Considering what we’ve explored in this article there are some very real world threats to our global economic infrastructure which are already happening and don’t include AI. Wars in Ukraine and in the Middle East already threaten global economics with rising oil prices and a predicted shortage of oil supply as well as rising inflation. While I’m not discounting the potential for AI to wreak havoc with global infrastructure as well as weapons systems I also believe we already have those threats. Maybe AI can help be a solution for our current and future threats. Can we put AI to use in helping us make it less threatening? Can we deploy AI to help us figure out how to utilize its capabilities without displacing many workers or by making changes to our economic system where we can co-exist?


References

#Source
1Anthropic researcher quits over out-of-control AI fears — The Wall Street Journal (via MSN).https://www.msn.com/en-us/technology/artificial-intelligence/anthropic-researcher-quits-over-out-of-control-ai-fears/ar-AA2bPYMN
2Artificial Intelligence Extinction Fears Spread to Mainstream America — The Wall Street Journal.
3H.R. 9917 — AI Kill Switch Act (119th Congress), introduced July 23, 2026. Text via GovTrack.https://www.govtrack.us/congress/bills/119/hr9917/text
4Dario Amodei, “We Must Pace the Frontier” (September 12, 2026).https://darioamodei.com/post/we-must-pace-the-frontier
5Investigating three incidents in our cybersecurity evaluations — Anthropic (July 30, 2026).https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
6How OpenAI’s and Anthropic’s AI models hacked other companies — NPR (August 1, 2026).https://www.npr.org/2026/08/01/nx-s1-5914852/anthropic-openai-models-hack-cybersecurity


© All rights reserved - jozephk

RSS

Letters

Private notes between readers and the author. Only published letters appear here for everyone; otherwise just the two correspondents see them.

Log in to write the author a private letter.