Deploying GenAI safely in government: what you must be able to demonstrate per workflow
New EU reports on generative AI in government shift the question from policy memo to workflow governance: here is what to record per use case.
You must now document, per separate workflow, which model, data and controls support each AI-assisted decision in your organisation. That burden of proof has shifted from a general policy statement to demonstrable governance per use case.
The prompt is an analysis of 10 September 2026 of safe deployment of generative AI in European public administrations, which argues that governments should structure safe deployment along infrastructure, governance and use-case selection rather than through abstract principles alone. The European Commission and Joint Research Centre have mapped 33 guidelines and 61 concrete use cases across European governments, identifying stated applications such as policy summaries, citizen services and internal knowledge assistants. In our assessment, the most significant shift is organisational rather than normative: the question is no longer only whether an application is permitted, but how a public authority can demonstrate, per situation, that an AI-assisted outcome was arrived at with the right controls in place.
What must you record for each workflow?
The frameworks set out by the Commission and Joint Research Centre converge on a single requirement: visibility of the decisions and infrastructure that support an outcome. For each workflow, you must be able to show which model is in use, which data it processes, which privacy and security controls are built in, and which moments of human oversight and audit trails support the decision. This is not a single compliance checklist but a set of questions you answer per use case.
The distribution of responsibilities across the phases of your workflow—from design through deployment to audit—is itself a point of attention. Each phase carries distinct obligations for documentation and oversight.
Which failure modes and risks does this framework address?
- Data sovereignty and leakage — personal data or sensitive government information reproduced or inferred from model outputs or retained by external providers.
- Security and infrastructure dependency — reliance on external systems without clear exit routes or contractual safeguards on data handling.
- Inconsistent policy application — different workflows applying different rules to the same decisions, creating unequal treatment of citizens.
- Absence of human oversight — AI-assisted outcomes recorded without evidence of which decisions remained with civil servants and which were delegated to the model.
- Audit trail failure — no logging of what the model actually did per run, making subsequent scrutiny and accountability impossible.
What concrete controls must you be able to demonstrate?
- Record the model and its purpose — document which model each workflow uses, the lawful basis for the data it processes, and the objective the workflow serves.
- Define the permitted risk level and decision boundary — determine in advance which decisions are and are not left to the AI, and what error rates or failure modes are acceptable.
- Log each run and its inputs — capture what data the model received, which sources it consulted, and what output it produced per transaction.
- Implement privacy controls before processing — apply data minimisation or synthetic equivalents on your own infrastructure before any data leaves your systems.
- Establish audit trails for human oversight — record which civil servant reviewed or overrode the model's recommendation, when, and on what grounds.
How do pilots fit into this framework?
The European Commission has announced new generative AI pilots for public administrations, focused on concrete services and linked to monitoring and evaluation within the AI Act framework. A pilot in the public sector should be more than a walled-off experimental environment. That means: determine the objective and the permitted risk level in advance, record which decisions are and are not left to the AI, and log per run what actually happened. Without that logging, no subsequent audit, no parliamentary scrutiny and no citizen accountability is possible. The design of such audit trails and evidence per AI action determines whether a pilot can later be accounted for.
What role does tooling play in meeting these requirements?
A verification layer can support this work by making verification steps, consulted sources, applied transparency signals and moments of human oversight visible per workflow. However, no tool can replace the professional final judgement of the civil servant or public authority themselves. Tooling can enforce a structure for documentation and make controls visible; it cannot guarantee privacy protection or compliance. The responsibility for every AI-assisted outcome remains with the person or organisation that deploys it.
Sources: This article draws on reporting and guidance from Europese Commissie (AI Watch), Joint Research Centre, Public Sector Tech Watch, Interoperable Europe Portal and Europese Commissie (DG CONNECT).
Written by
Marit Halversen
Covers AI governance and regulatory design, with a focus on how compliance obligations land on architecture rather than on paperwork.