Standards body for AI safety testing: what you must now be able to demonstrate per workflow
AI firms are discussing their own standards body for safety testing. Here is how to translate external model evaluations into verifiable governance per workflow.
You must now be able to document, per workflow, which frontier models you deploy, which external safety tests they have passed, what those tests found, and how you have translated those findings into local usage limits, access controls and human review steps. This is no longer optional if you handle confidential or high-risk information.
An analysis of 15 September 2026 of shared safety testing standards for frontier models argues that major AI developers are moving from internal, opaque testing towards multi-party evaluations that organisations like yours must actively incorporate into deployment decisions. The specific case is the emerging industry standards body being designed by working groups at several frontier labs, combined with the pre-release evaluation programme already run by the US Center for AI Standards and Innovation. In our assessment, this represents a shift in where the burden of proof lies: safety claims are becoming verifiable artefacts that you must be able to point to, rather than marketing statements you accept on trust.
What has changed in how frontier models are tested?
Until recently, safety testing happened inside individual labs and remained largely opaque. The current development establishes a different model: external assessors, standardised risk frameworks, and documented results that sit in the public domain or are available to deployers. Working groups at major AI developers have been meeting since July 2026 to design an industry-led standards body focused on pre-release testing. Separately, the US Center for AI Standards and Innovation has already conducted more than forty pre-deployment evaluations across five major labs, with an emphasis on cyber and biosecurity risks. This creates a layer of verifiable testing that did not exist before.
Which concrete controls must you be able to demonstrate per workflow?
- Record the model and its purpose — document which frontier model each workflow uses and the lawful basis for the data it processes.
- Identify the external tests applied — establish which safety assessments the model passed before deployment and who conducted them.
- Retain the test results — keep copies of evaluation reports, risk assessments and any published findings relevant to your use case.
- Document your local thresholds — record how you translated external test results into usage limits, access restrictions and approval gates specific to your organisation.
- Track updates and re-evaluations — maintain a record of when models are updated and which new tests, if any, have been run since you began using them.
Why does it matter that multiple assessors may now evaluate the same model?
A future industry standards body will not replace government evaluation programmes like those run by the US Center for AI Standards and Innovation. Instead, you may find yourself dealing with results from multiple sources: industry-led assessments, government pre-deployment evaluations, and potentially regional or sectoral reviewers. This creates both opportunity and complexity. The opportunity is that no single lab controls the narrative about safety. The complexity is that you must be able to reconcile different assessment frameworks, understand what each one covers, and decide how to weight their findings in your own risk decisions. This is why per-workflow documentation becomes essential: you need to know not just that a model passed tests, but which tests, run by whom, and what gaps remain.
What remains your responsibility after external testing is complete?
External safety testing tells you about the model in isolation. It does not tell you whether your organisation's deployment of that model is safe. The translation from general test results to your specific workflows, data types and user populations is work you cannot outsource. If you use a frontier model to analyse confidential client information, a test that found the model resistant to prompt injection does not guarantee it will not leak that information. If you use it in a decision that affects individuals, a test that found it performs well on a benchmark does not mean it will perform well on your data. The standards body and government evaluators can make testing verifiable and comparable. They cannot make it local. That remains your professional judgement.
How should this affect your procurement decisions?
When you assess suppliers or model providers, ask which external tests have been run, by whom, and what the results show. This is now a concrete procurement criterion, not a nice-to-have. Suppliers who cannot point to independent evaluations or who rely only on their own internal testing are asking you to carry more risk than those who can show results from external assessors. At the same time, the existence of external testing does not mean you can skip your own due diligence. It means you have a starting point: a documented baseline of what the model can and cannot do, against which you can then assess whether it is fit for your purposes.
Tooling can help you map which model is used in which workflow, track which tests apply to each one, and flag when updates require re-evaluation. What tooling cannot do is decide whether a given test result is acceptable for your risk tolerance, or whether the gaps a test reveals matter for your use case. That judgement stays with you.
Sources: This article draws on reporting and guidance from Tweakers, Pondero, Cloud Security Alliance and The Guardian.
Written by
Marit Halversen
Covers AI governance and regulatory design, with a focus on how compliance obligations land on architecture rather than on paperwork.