GPT-6 Astra in legal workflows: what you must be able to demonstrate per task
Astra raises agentic legal workflows at Harvey and Legora to a higher level. Here you can read what you must be able to reconstruct per matter, document and check.
You must now be able to reconstruct and document every step an agentic AI system takes through a legal workflow: which model ran the task, which source documents it read, which checks it performed, and where human sign-off occurred. This is no longer optional for trust-sensitive work.
The prompt is an analysis of 7 September 2026 of agentic legal workflows powered by GPT-6 Astra, which argues that the model's quality gains on complex multi-step legal tasks are only usable once every task becomes fully auditable and reconstructable. The case study describes an agent carrying out financial tie-out across 41 documents in a single run, checking balances against trial balances and consolidation schedules, and flagging discrepancies as structured evidence for human review. In our assessment, the shift towards agentic workflows that execute complete checking tasks under human oversight changes what you have to record and how you have to structure the work itself.
What makes Astra different from earlier models in legal practice?
Astra represents a step change in how legal AI handles complex, multi-document tasks. The model distinguishes source documents from drafts more reliably, separates supported from unsupported assumptions, and turns gaps into concrete drafting suggestions. On financial workflows, it reportedly delivers roughly a 40% improvement over the previous model on specific tie-out tasks and around 3% improvement on broader internal benchmarks.
The practical shift is this: a lawyer no longer asks a question of an assistant and receives an answer. Instead, you delegate a multi-step task to an agent that works across documents, data sources and tools, maintaining state and carrying out actions in sequence. That distinction—between isolated prompts and governed workflows—is where the real operational change lies.
Which failure modes and risks does agentic legal work introduce?
- Untraced decision paths — an agent completes a task but the steps it took, the documents it consulted, and the reasoning it applied remain opaque to the lawyer who must sign off.
- Model substitution without record — the agent runs on one model during development and another in production, with no documented change in capability or behaviour.
- Hallucinated sources — the agent cites or relies on documents, data points or precedents that do not exist in the matter file.
- Unchecked intermediate outputs — the agent produces working outputs, drafts or calculations that are never reviewed before being fed into the next step.
- Loss of audit trail — the lawyer signs off on a final output but cannot later reconstruct which agent version, which documents, and which checks led to that outcome.
- Unmonitored tool use — the agent accesses external data sources, APIs or systems without logging what it retrieved or how it used the result.
What must you be able to demonstrate per matter?
- Record the model and its purpose — document which model each workflow uses, when it was deployed, and the lawful basis for the data it processes.
- Log every document the agent consulted — maintain a list of source materials the agent read, in what order, and which sections it extracted or cited.
- Capture each discrete check or calculation — record what verification the agent performed, what criteria it applied, and what result it reached at each step.
- Document human review points — note where a lawyer reviewed the agent's output, what they checked, what they approved or rejected, and when they signed off.
- Preserve the agent's reasoning or evidence — retain the structured output, working papers, or audit log the agent generated so the task remains reconstructable after completion.
How do you operationalise this in practice?
These requirements align with existing expectations around working papers and audit trails in legal practice. Astra's agentic use simply turns those expectations into technical design requirements. You need to think of each workflow as a governed process, not a one-off query.
For outcomes that rest on text—a legal opinion, a due diligence summary, a contract review—claim-by-claim checking of the agent's outputs remains a sensible addition to the agent's logging. The model raises the ceiling on what the agent can do, but the auditability of every step determines whether you can use it in high-trust work.
Tooling can help operationalise this accountability. A verification layer can route a task through selected independent models and expose verification steps, corrections, disagreements and sources for inspection. That supports review and gives insight into how an outcome came about; it does not promise correctness and does not remove the need to check for hallucinations. For sensitive documents, privacy-preserving architectures can replace sensitive values with synthetic equivalents before AI processing, sending only anonymised content onward and blocking transmission if the privacy check fails.
The professional final judgement always remains with the lawyer. Astra raises what the agent can accomplish, but you remain accountable for what you delegate and what you sign off. That is exactly the role that suits agentic AI: the model does the work, but you control whether it is usable in your practice.
Sources: This article draws on reporting and guidance from Artificial Lawyer, OpenAI, Flank, Legora and WebProNews.
Written by
Marit Halversen
Covers AI governance and regulatory design, with a focus on how compliance obligations land on architecture rather than on paperwork.