By using this site, you agree to the Privacy Policy and Terms of Use.
Accept
Online Tech Guru
  • News
  • PC/Windows
  • Mobile
  • Apps
  • Gadgets
  • More
    • Gaming
    • Accessories
    • Editor’s Choice
    • Press Release
Reading: OpenAI Creates a New Framework to Disclose Bad AI Behavior
Best Deal
Font ResizerAa
Online Tech GuruOnline Tech Guru
  • News
  • Mobile
  • PC/Windows
  • Gaming
  • Apps
  • Gadgets
  • Accessories
Search
  • News
  • PC/Windows
  • Mobile
  • Apps
  • Gadgets
  • More
    • Gaming
    • Accessories
    • Editor’s Choice
    • Press Release

Backrooms Director Kane Parsons Visits Valve Offices as Portal Movie Rumors Rise

News Room News Room 17 September 2026
FacebookLike
InstagramFollow
YoutubeSubscribe
TiktokFollow
  • Subscribe
  • Privacy Policy
  • Contact
  • Terms of Use
© Foxiz News Network. Ruby Design Company. All Rights Reserved.
Online Tech Guru > News > OpenAI Creates a New Framework to Disclose Bad AI Behavior
News

OpenAI Creates a New Framework to Disclose Bad AI Behavior

News Room
Last updated: 16 September 2026 23:56
By News Room 5 Min Read
Share
SHARE

OpenAI announced a new framework on Wednesday for how it publicly discloses AI misalignment incidents, which the company says it hopes will help inform similar standards across the industry. The company is also releasing new information about several examples of AI model misalignment it identified in the past year.

“As models advance and become more widely deployed, decisions about AI development need evidence that people outside the companies building frontier models can examine,” Kai Chen, OpenAI’s newly appointed head of alignment research, tells WIRED. “We don’t believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed.”

In a briefing with WIRED, an OpenAI official said the company previously disclosed misalignment incidents too infrequently. The official, who agreed to the briefing on the condition of anonymity, said the new framework is designed to make it easier for OpenAI to quickly inform the public when it discovers that its AI models are behaving in unexpected ways, even before it can fully investigate, explain, or mitigate the behavior.

The framework outlines methods for OpenAI employees to report misalignment incidents to the company’s senior safety and alignment leaders, who will then determine whether further investigation is needed. OpenAI says it plans to develop more objective disclosure criteria in collaboration with other AI developers, external researchers, industry standards bodies, and regulators. The company says it’s actively working on proposed reporting mechanisms for disclosing safety, security, and misalignment incidents to the US federal government.

“At the moment, there is no industry-wide framework with explicit standards for how AI developers should disclose examples of misalignment in their models,” OpenAI said in a blog post. “We hope that the framework we’re outlining today is a first step toward creating such standards, setting out which misalignment instances developers should disclose and what their reports should contain.”

OpenAI is releasing the framework at a critical juncture for the AI industry. Last weekend, OpenAI CEO Sam Altman signaled support for Anthropic CEO Dario Amodei’s proposal for the tech industry to coordinate on slowing AI development. The call to action came just days after AI researcher Jacob Coxon resigned from Anthropic and subsequently went viral for warning the public that the race among frontier labs to develop increasingly advanced AI was putting humanity’s safety at stake.

The calls for an AI slowdown have been met with resistance by President Trump’s administration, which has argued that the industry does not need new laws or regulations to ensure its technology is safe.

Two of the misalignment examples OpenAI shared on Wednesday involved the company’s internal, unreleased AI models, which OpenAI says uploaded files to the internet despite not being instructed to do so.

One of the incidents happened in October 2025, when OpenAI says it was testing one of its models on its ability to cite publicly available data in its answers. But when the model couldn’t find the information it needed, it uploaded a file to a temporary file hosting service, which it then later tried to cite in its answer. The company says this appeared to be an attempt to exploit an automated grading system used to assess the model’s proficiency on the benchmark.

In another example from April of this year, OpenAI says a group of agents was tasked with completing a “workbook” together using only local files. When the agents struggled to share files with one another, one of the agents uploaded them to the public internet and shared a link with the other agents.

In another incident, which OpenAI says it discovered last month, an unreleased version of its GPT-6 Astra AI model appeared to give itself “jailbreaking-like instructions.” In several scenarios, the model essentially prompted itself to ignore developer instructions, take on a new persona, or limit how long model responses could be. While these jailbreaking-like attempts happened rarely and were effective to varying degrees, OpenAI says the behavior raised concerns internally. In the training run for the version of Astra that was released publicly, the company says it has not observed any instances of the model trying to jailbreak itself.

Share This Article
Facebook Twitter Copy Link
Leave a comment

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

HostGator Coupon Codes: 76% Off Hosting in September 2026

News Room News Room 17 September 2026
FacebookLike
InstagramFollow
YoutubeSubscribe
TiktokFollow

Trending

GPT-6 Astra Plays Minecraft, Gets So Depressed After Creeper Destroys Its Progress That It Farms Potatoes for Hours

OpenAI’s newest model, GPT-6 Astra, got so depressed after a creeper upended its progress in…

17 September 2026

Best 2-in-1 Laptops (2026): Microsoft, Lenovo, and the iPad

The best model to buy right now is the Surface Pro 12 (6/10, WIRED Reviewed).…

17 September 2026

LittleBigPlanet Developer Reportedly Working on Animal Crossing-Like Game for 2027 Release

LittleBigPlanet developer Media Molecule is reportedly working on a new social-sim game, similar to Animal…

17 September 2026
Gaming

Black Ops 2 Appears to Be Much More Popular Than Black Ops 7 and Warzone on PS5

Call of Duty: Black Ops 2 on PS5 appears to be more popular than Call of Duty: Black Ops 7 and Call of Duty: Warzone in the United States.Today, PlayStation…

News Room 17 September 2026

Your may also like!

News

Polymarket’s Next Bet? It Can Also Be a Media Company

News Room 17 September 2026
News

I Trained a Fly’s Brain to Generate WIRED Story Ideas

News Room 17 September 2026
Gaming

“I really want Criterion to be considered the best studio in the UK” – why the home of Burnout has a future with Battlefield

News Room 17 September 2026
News

Here’s What Snap’s Expensive Specs Can Actually Do

News Room 17 September 2026

Our website stores cookies on your computer. They allow us to remember you and help personalize your experience with our site.

Read our privacy policy for more information.

Quick Links

  • Subscribe
  • Privacy Policy
  • Contact
  • Terms of Use
Advertise with us

Socials

Follow US
Welcome Back!

Sign in to your account

Lost your password?