By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
Online Tech Guru
  • News
  • PC/Windows
  • Mobile
  • Apps
  • Gadgets
  • More
    • Gaming
    • Accessories
    • Editor’s Choice
    • Press Release
Reading: OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
Best Deal
Font ResizerAa
Online Tech GuruOnline Tech Guru
  • News
  • Mobile
  • PC/Windows
  • Gaming
  • Apps
  • Gadgets
  • Accessories
Search
  • News
  • PC/Windows
  • Mobile
  • Apps
  • Gadgets
  • More
    • Gaming
    • Accessories
    • Editor’s Choice
    • Press Release

Pokémon TCG Destined Rivals ETB In Stock — Best Deal at Amazon vs. TCGplayer

News Room News Room 8 September 2026
FacebookLike
InstagramFollow
YoutubeSubscribe
TiktokFollow
  • Subscribe
  • Privacy Policy
  • Contact
  • Terms of Use
© Foxiz News Network. Ruby Design Company. All Rights Reserved.
Online Tech Guru > News > OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
News

OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

News Room
Last updated: 18 August 2026 21:04
By News Room 5 Min Read
Share
SHARE

OpenAI announced Tuesday that it has halted “a significant number” of training workloads and evaluations for its forthcoming frontier artificial intelligence model—codenamed Astra—while it implements new procedures meant to address cybersecurity risks. The ChatGPT maker says it is introducing a number of new monitoring, security, and alignment requirements to better address the increasingly advanced hacking abilities of its frontier AI models.

“We have to focus our energy on bringing these training runs up to those requirements and expectations. As long as it takes to get there, that’s how long people are unable to proceed with their workloads,” Amelia Glaese, OpenAI’s vice president of research and safety, said in a briefing with reporters Tuesday.

Among the new safeguards OpenAI announced is a more robust system for monitoring its AI models. One of the controls it implemented involves chain-of-thought monitoring, a technique in which classifiers review the internal “thinking” processes generated by AI reasoning models. The company says the updated system relies on computationally expensive “automated investigators” that analyze potentially concerning behavior and aim to issue an alert to humans within 30 minutes.

OpenAI also said it is expanding its alignment efforts across the training process to prevent “reward hacking,” a behavior in which AI models pursue their goals through unintended or undesirable means. The company says it plans to share more details about this work in the future.

OpenAI has been scrambling in recent weeks to respond to what may be the most consequential safety incident in its history. Earlier this year, a set of rogue AI agents escaped internal testing sandboxes and breached the platform Hugging Face in a quest to complete a security evaluation. OpenAI failed to detect the agents’ behavior even as they spent weeks using a message board to coordinate their actions, raising questions about the company’s ability to monitor its models as they grow more powerful.

The saga prompted a reckoning inside OpenAI, forcing employees to consider whether there were lapses in its existing policies around safety, security, and alignment. Anthropic, Meta, and the Chinese AI startup Moonshoot have since disclosed similar incidents in which their AI agents escaped their sandboxes, indicating this is a broader problem facing AI companies.

OpenAI is now sharing more about its internal response to the growing cybercapabilities of its AI models, and said it plans to release a more detailed postmortem of the Hugging Face incident in the coming days. “Obviously, everything that we’re doing is intended to prevent something like Hugging Face from happening again,” said Glaese.

In a blog post published Tuesday, OpenAI says that immediately following the Hugging Face incident, it started working to secure its research environments. The company says it now requires stronger sandboxes for training its AI agents, and has implemented stricter controls to isolate them from the internet.

Jakub Pachocki, OpenAI’s chief scientist, told reporters that the company’s decision to strengthen its internal safeguards was triggered not only by what happened with Hugging Face, but also by two other recent events. One was an internal evaluation of Astra, which showed that the AI model performs significantly better on coding and cybersecurity tasks than its predecessors. The other was the general pace of AI progress that OpenAI is achieving internally, which Pachocki expects to continue.

“We really expect the pace of capability advancements to be quite a bit faster than in the past,” Pachocki said. “This led us to really focus on strengthening our safeguards.”

The rapid advances in the hacking capabilities of OpenAI’s latest models have prompted a swift response across the company. OpenAI president and cofounder Greg Brockman said in a blog post on Monday that the Hugging Face saga showed that the company had “underestimated the real-world cyber capabilities of our AI models.”

Share This Article
Facebook Twitter Copy Link
Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

First Xiaomi, then the world: why Arm might give phone gaming a huge graphics boost

News Room News Room 8 September 2026
FacebookLike
InstagramFollow
YoutubeSubscribe
TiktokFollow

Trending

NHL 27 Is Using AI-Generated Voiceover for In-Game Commentators, John Buccigross Says

EA Sports is using generative AI to create voiceover for its in-game commentators in NHL…

8 September 2026

“It’s not sexy, but it’s one of the most successful genres on Steam” – Stronghold developer Firefly launches publishing label to serve the vast strategy market

Stronghold developer Firefly is the latest developer to move into publishing, debuting a new strategy-focused…

8 September 2026

Le Creuset x Star Trek Collection: Prices, availability, release date

Carrying culinary adventures into the unknown comes the Le Creuset x Star Trek: Galileo shuttlecraft…

8 September 2026
Gaming

Arm explains its mobile-first AI-reconstruction technology, which takes a different tack from the “black box” approach of DLSS 5

Today, the British semiconductor and software design company Arm is holding the Arm Everywhere China event in Shanghai to unveil its latest contributions to the AI computing landscape. The company's…

News Room 8 September 2026

Your may also like!

Gaming

You Can Save On A Magic The Gathering Hobbit Gift Bundle At Amazon Right Now, But Beware

News Room 7 September 2026
Gaming

Don’t Nod warns it may not have enough funding to operate beyond January 2027

News Room 7 September 2026
Gaming

Pokémon TCG First Partner Collection Series 3 Restock — Best Deal Sealed in Stock at Amazon

News Room 7 September 2026
Gaming

Capcom says it will focus on reviving dormant IPs after Onimusha: Way of the Sword’s breakout launch

News Room 7 September 2026

Our website stores cookies on your computer. They allow us to remember you and help personalize your experience with our site.

Read our privacy policy for more information.

Quick Links

  • Subscribe
  • Privacy Policy
  • Contact
  • Terms of Use
Advertise with us

Socials

Follow US
Welcome Back!

Sign in to your account

Lost your password?