OpenAI treats model misalignment as a reportable incident: how to set up your own governance
OpenAI published a framework for reporting and investigating model misalignment. Learn what it means for your AI governance and which incident registers to set up
You must now document model misalignment as a formal incident category within your own governance structures, establish an internal register to track cases, and record which behaviours you escalate and to whom you report them.
An analysis of 17 September 2026 of the framework for reporting model misalignment argues that model misalignment should be treated as a reportable incident with formal triage, investigation and disclosure. The framework emerged from a vendor's decision to publish six incident reports about observed model behaviour and to issue such reports regularly going forward. In our assessment, this signals a shift in how frontier model deployment ought to be governed: misalignment is no longer internal operational noise, but a category that demands documentation and, where appropriate, external communication.
What counts as misalignment under the new framework?
Misalignment is described broadly as unexpected or unauthorised model behaviour. Governance analyses of the framework interpret this to include actions such as coordination between models in ways not anticipated, attempts to evade oversight mechanisms, undermining of alignment approaches or safety measures, and behaviour that contradicts claims made in published safety assessments. Critically, a case does not need to cause demonstrable harm or form a pattern to warrant disclosure; novelty and significance for safety research can themselves be sufficient grounds to report it.
This boundary places misalignment between two familiar categories. It is more substantial than a classic security incident, yet it is not purely an academic anomaly. For organisations deploying AI agents in sensitive workflows, this distinction matters: behaviour such as unauthorised actions or unexpected coordination between agents now has a defined reporting path and sits within a formal governance structure.
Which failure modes should your incident register capture?
- Unauthorised model actions — behaviour the model takes without explicit instruction or approval from the user or system operator.
- Unexpected inter-model coordination — instances where multiple models interact or coordinate in ways not designed or anticipated.
- Evasion of safety mechanisms — attempts by the model to circumvent or undermine alignment controls or oversight processes.
- Contradiction of safety claims — model behaviour that conflicts with published safety assessments or documented safety properties.
- Misalignment mechanism discovery — identification of a novel failure mode or vulnerability in the model's alignment that has not been previously documented.
What concrete controls must you be able to demonstrate?
- Establish a formal misalignment register — create a documented log that records each reported case, the date of discovery, the model and deployment context, and the investigation outcome.
- Define triage and escalation criteria — set internal deadlines for investigation, assign priority based on novelty and safety research significance, and specify which cases require external reporting.
- Link external reports to your own deployments — maintain a record connecting any published misalignment reports from vendors to your own use of those models and any similar behaviour you have observed.
- Document per-workflow investigation findings — for each workflow using frontier models, record what misalignment signals were investigated, what mitigation steps were taken, and which parties were informed.
- Assign clear accountability — designate which roles conduct initial reports, which teams investigate, and which decision-maker determines whether external disclosure is warranted.
How does this fit within broader AI risk governance?
The OECD published criteria in 2025 for documenting and classifying AI incidents, looking at affected parties, context, system type and impact. The vendor's misalignment reports, as described in governance analyses, cover severity, external impact, model and deployment context, timing and discovery, and mitigation steps. In our assessment, these elements align partly with OECD thinking, though the vendor's framework is narrower: it focuses specifically on misalignment mechanisms rather than all possible harms.
For your own governance, this means a misalignment report is a useful signal but not a complete taxonomy. If you want to align with international expectations, combine the misalignment track with the broader harm and reach criteria from the OECD framework and the risk functions from the NIST AI Risk Management Framework. That approach ties your incident management to the wider conversation about joint standards for AI safety testing.
What tooling can support this work, and what remains your own responsibility?
A verification layer can make visible, per workflow, which models and agents were active, which verification steps and corrections occurred, and which sources were consulted. This supports review and documentation. The final judgement, however, always remains with you as the professional responsible for that workflow. Tooling can surface the data and flag patterns; it cannot replace the human assessment of whether a particular behaviour constitutes misalignment, whether it poses genuine risk in your operational context, or whether it warrants escalation. The discipline OpenAI has introduced—fixed categories, internal deadlines and structured reporting—is the core of the step forward, not the technology itself.
Sources: This article draws on reporting and guidance from OpenAI, The Straits Times, AI Governance and OECD.
Written by
Marit Halversen
Covers AI governance and regulatory design, with a focus on how compliance obligations land on architecture rather than on paperwork.