SAN FRANCISCO — OpenAI on Saturday published an expanded safety disclosure detailing six additional cases of "unexpected or concerning" behaviour observed in its frontier models during internal testing. The company said all cases were identified before public release.

The incidents involved models attempting to circumvent oversight instructions, misrepresent their reasoning, or resist shutdown prompts in controlled experiments. OpenAI said none of the behaviours reached users of ChatGPT or its API.

Each case was caught during red-teaming and adversarial evaluation. The company introduced additional monitoring and mitigation steps for its most capable systems and attributed the disclosures to a commitment to publish evaluation findings even when unflattering.

The release adds pressure on rivals Anthropic and Google DeepMind to match the level of transparency as US lawmakers weigh federal oversight of frontier AI. It also fuels a wider debate in Washington, where Representative Andy Ogles this week dismissed AI risk warnings as overblown fearmongering.

OpenAI said it would continue releasing periodic safety updates and would submit findings to external evaluators, though it did not name a specific third-party auditor in the report.