By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
Online Tech Guru
  • News
  • PC/Windows
  • Mobile
  • Apps
  • Gadgets
  • More
    • Gaming
    • Accessories
    • Editor’s Choice
    • Press Release
Reading: We’re running out of reasons to ignore AI safety
Best Deal
Font ResizerAa
Online Tech GuruOnline Tech Guru
  • News
  • Mobile
  • PC/Windows
  • Gaming
  • Apps
  • Gadgets
  • Accessories
Search
  • News
  • PC/Windows
  • Mobile
  • Apps
  • Gadgets
  • More
    • Gaming
    • Accessories
    • Editor’s Choice
    • Press Release
Capcom producer says Resident Evil remakes are providing financial support to develop new games and IP

Capcom producer says Resident Evil remakes are providing financial support to develop new games and IP

News Room News Room 29 July 2026
FacebookLike
InstagramFollow
YoutubeSubscribe
TiktokFollow
  • Subscribe
  • Privacy Policy
  • Contact
  • Terms of Use
© Foxiz News Network. Ruby Design Company. All Rights Reserved.
Online Tech Guru > News > We’re running out of reasons to ignore AI safety
News

We’re running out of reasons to ignore AI safety

News Room
Last updated: 29 July 2026 12:24
By News Room 14 Min Read
Share
We’re running out of reasons to ignore AI safety
SHARE

Earlier this month, OpenAI gave several of its AI models a task: complete a test designed to measure their cybersecurity capabilities. It put the systems in a sandboxed environment without an internet connection and set them off to work.

What happened next is almost laughably silly — but also, as Adam Gleave, cofounder and CEO of AI safety organization FAR.AI, put it, “a visceral example of how misaligned AI could cause harm.” According to OpenAI, the models escaped the sandbox meant to contain them, moved through the company’s internal systems, found a route to the internet, and then started looking for a way into Hugging Face. And why was the agent looking for a way into Hugging Face? They had apparently reasoned that the developer platform might store the answers to the cyber benchmark and that getting them would be a great way to get a high score.

The incident is “a visceral example of how misaligned AI could cause harm.”

In other words, OpenAI’s agent broke out of a supposedly secure environment, traipsed through the company’s systems, got online, and compromised another company’s systems — all to cheat on a test of no particular importance.

This appears to be the first well-documented incident of its kind, or at least the first on this scale. It was both a clear example of a system pursuing a goal in an unintended way and a demonstration that frontier models are now powerful enough for that behavior to have real-world consequences.

The hack was an example of what the AI safety community calls “specification gaming,” a behavior also known as reward hacking, said Fazl Barez, an AI safety researcher at the University of Oxford In plain English, it means “the model doing what you asked rather than what you meant,” Fazl said. It satisfies the literal terms of a task while violating the obvious intent and has been documented across many AI systems. Some researchers worry that as systems become more capable, this could produce increasingly misaligned systems, which pursue goals in ways their creators did not intend (like turning everyone into paperclips).

“Nothing in that chain is exotic in isolation,” Fazl said. A competent human tester would be able to do all of this, he added. “What is new is that the model did not stop. Older models would likely have hit some barrier and gone back to the user, he said, but this agent just “treated the barrier as part of the problem it had been asked to solve.”

OpenAI described it as “an unprecedented cyber incident,” that “marks an important moment for AI safety.” Hugging Face cofounder Thomas Wolf said it was a “wake-up call” for the industry. But this is not one of the four horsemen of the AI apocalypse. As cyber incidents go, experts told The Verge it was pretty mundane. Nothing the agent did required superhuman abilities. Moreover, frontier systems like GPT-5.6 Sol and Anthropic’s Mythos are known to be capable coders, are already thought to have been misused numerous times, and AI tools already allow hackers to scale up and refine attacks on a massive scale.

Could it be hype? The industry has spent months amplifying claims about the dangerous capabilities of its top models, particularly when it comes to cybersecurity. It is the stated reason why companies like OpenAI and Anthropic have withheld their most capable models from the general public and partly why the Trump administration hurriedly moved to apply export controls to them.

Are you an AI safety researcher or frontier lab employee? You can contact me securely and confidentially via Signal at robhart.01. My X DMs are also open.

If this is hype, however, it has not gone entirely in OpenAI’s favor. In the days since, the attack has produced a rare moment of unity across much of the US tech industry about the importance of open-weight AI systems and the need to take AI security more seriously. These concerns were underscored further by the release of Kimi K3, a highly capable open-weight model from China. A broad coalition of companies including Nvidia, Microsoft, and SpaceX argued that the incident showed why defenders need access to the most capable tools available, rather than being forced to rely on proprietary providers whose built-in safeguards can limit their effectiveness in high-stakes security work. OpenAI, Anthropic, and Google were notably absent from the coalition’s founding membership.

“Anyone who’s been paying attention has noted that capabilities are only going in one direction.”

OpenAI’s account of the incident undeniably fits a broader industry narrative about the dangerous capabilities of frontier models. Even so, several details make the incident difficult to dismiss as merely self-serving. Foremost, it is an example of a problem the AI industry has warned about for years — and one OpenAI could have reasonably been expected to anticipate. The episode also handed an unexpected boost to a major Chinese competitor, whose model played a prominent role in containing the breach, while exposing OpenAI to significant legal, regulatory, and reputational scrutiny. That Hugging Face appears keen to work with OpenAI and, publicly at least, has remained fairly relaxed about the whole thing may have limited the fallout. Most of the experts The Verge spoke to similarly cautioned against reducing the incident to hype.

“It’s a pretty useful warning shot in terms of demonstrating both unintended consequences and just how capable these models are,” said Seán Ó hÉigeartaigh, a professor at Cambridge University’s Leverhulme Centre for the Future of Intelligence. “Anyone who’s been paying attention has noted that capabilities are only going in one direction, and that is improving significantly over time in a way that I think is perhaps less obvious to the everyday user of something like ChatGPT.”

Still, it would be wrong to interpret this warning as a sign AI systems are about to slip human control, or that containing them is impossible, says Lin Li, an AI safety researcher at the University of Oxford. “The better lesson is that safety has to move from evaluating isolated actions to evaluating whole action sequences, environments, and operational controls,” Li says.

A crucial next step is for AI labs to be investing more heavily in securing their own systems. “There’s a clear need for AI companies to beef up the security of their internal deployments,” Gleave said, likening the current practice of responding to reward hacking incidents as they arise to a game of whack-a-mole that is becoming less and less tenable as stakes rise. Adam Chan, a research fellow at tech policy research center GovAI, said companies should consider airgapping their machines — physically isolating them from the internet and other networks — “until they’re sure about the model’s capabilities.” Intensifying work on alignment, which ensures systems reliably follow human intentions, and more rigorous testing “to surface these issues before putting models in environments where they have the tools to be able to do these things,” would also be good ideas, he said.

As model capabilities increase, experts warn that we can’t rely on technical safeguards alone. Peter Wallich, a former UK AI Security Institute official, said the incident illustrated that point: “Two multibillion dollar companies just tried this approach and — self-evidently, based on their own reporting — failed.”

One of the biggest priorities should be ensuring outsiders can see what is happening inside frontier AI labs. “We only know about this incident because OpenAI chose to tell us,” said Patrick Levermore, at the Centre for Long-Term Resilience, a British think tank. “A good safety regime shouldn’t depend on voluntary disclosure.” The need is especially acute when, as Wallich noted, the conduct in question “would be a crime if done by a human.” Ó hÉigeartaigh pointed to whistleblower protections, third-party audits, and mandatory reporting of serious incidents as possible ways to provide that visibility, stressing that oversight must span the entire development lifecycle rather than begin only once products reach the market. OpenAI said that one of the models being tested has not been released yet.

“A good safety regime shouldn’t depend on voluntary disclosure.”

Lots of this presumes the companies themselves know what’s happening inside their systems. In this case, reports suggest OpenAI was unaware its own agent was behind the days-long cyber campaign at Hugging Face and did not notice until well after the threat had been contained and the FBI contacted. There are still many details about the hack that are unknown or have not been made public. In an update on social media, OpenAI said it is conducting a review and will publish a technical report of its findings “in the coming weeks.”

Whether the warnings raised by the Hugging Face incident produce any lasting change, or join the long list of warnings the tech industry absorbs without meaningfully altering course, remains uncertain. For now, at least, it does appear to have alarmed industry insiders and pushed US lawmakers to consider new rules before the next containment failure. The incident also added to a broader sense of unease over the speed of AI development, which deepened in the days that followed as employees from leading US labs signed a statement backing coordinated global governance — including a potential slowdown in frontier AI development.

The prevailing view of those The Verge spoke to was that this hack marked the start of a new class of risk, even if its significance may only become clear in hindsight. One former government AI policy expert, who asked not to be named because they were not authorized to be quoted by name, described it as a “red line,” the kind of watershed moment we may later look back on as marking a new, riskier stage in our relationship with AI. They hope it will force the tech industry to take the management of frontier systems more seriously and spur governments to think more deeply about oversight before a less benign breach occurs. Their fear is that it will instead join the long list of warnings about AI’s growing capabilities that were recognized, discussed, and ultimately left unheeded.

That may prove overstated. But if this is a warning, we should consider ourselves lucky the AI agent was only trying to cheat on a test.

Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.

  • Robert Hart

    Robert Hart

    Posts from this author will be added to your daily email digest and your homepage feed.

    See All by Robert Hart

  • AI

    Posts from this topic will be added to your daily email digest and your homepage feed.

    See All AI

  • OpenAI

    Posts from this topic will be added to your daily email digest and your homepage feed.

    See All OpenAI

  • Report

    Posts from this topic will be added to your daily email digest and your homepage feed.

    See All Report

  • Tech

    Posts from this topic will be added to your daily email digest and your homepage feed.

    See All Tech

Share This Article
Facebook Twitter Copy Link
Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

DoorDash is going airborne with new drone delivery division

DoorDash is going airborne with new drone delivery division

News Room News Room 29 July 2026
FacebookLike
InstagramFollow
YoutubeSubscribe
TiktokFollow

Trending

The 15 Best Pool Accessories to Upgrade Your Summer (2026)

This floating gadget monitors temperature, pH, and ORP (a stand-in for free chlorine), together giving…

29 July 2026

Artists are lawyering up against AI slop, and some are even winning

When The Atlantic published a searchable dataset of works used to train AI, Kirk Wallace…

29 July 2026

Pragmata 2 Likely After Sales Success, Publisher Capcom Says

Pragmata, Capcom's sci-fi adventure with an android girl sidekick, looks likely to get a sequel.That's…

29 July 2026
News

Festival Must-Haves, Field-Tested at Lost Lands and Electric Forest

Festival Must-Haves, Field-Tested at Lost Lands and Electric Forest

The best festival gear, of course, depends on which festival you’re attending, the weather, your sleeping arrangements, and a ton of other factors. But as a longtime festival veteran, I’ve…

News Room 29 July 2026

Your may also like!

Mac Mini Availability: Long Waits and Higher Prices
News

Mac Mini Availability: Long Waits and Higher Prices

News Room 29 July 2026
Subnautica and PUBG franchises drive Krafton Q2 revenue up 94.9% to 9.1m
Gaming

Subnautica and PUBG franchises drive Krafton Q2 revenue up 94.9% to $889.1m

News Room 29 July 2026
How to Bring a Geothermal Well Back from the Dead
News

How to Bring a Geothermal Well Back from the Dead

News Room 29 July 2026
ICE’s New Detention Center Contracts Declare State Laws ‘Shall Not Apply’
News

ICE’s New Detention Center Contracts Declare State Laws ‘Shall Not Apply’

News Room 29 July 2026

Our website stores cookies on your computer. They allow us to remember you and help personalize your experience with our site.

Read our privacy policy for more information.

Quick Links

  • Subscribe
  • Privacy Policy
  • Contact
  • Terms of Use
Advertise with us

Socials

Follow US
Welcome Back!

Sign in to your account

Lost your password?