🔍 Read the full analysis: Is OpenAI Training Agents Inside Your Software? Look At Ironclad on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI described training a frontier model, GPT-6 Astra, on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. Astra met an average 55% of rubric criteria; the reported task times were simulated estimates, not measured customer productivity gains. OpenAI says it used no non-public Ironclad customer data and says human oversight remains necessary.
OpenAI has reported training its frontier model GPT-6 Astra on legal, commercial and procurement tasks inside hosted copies of Ironclad’s contract-management software. The company says Astra met an average of 55% of task rubric criteria, while stressing that its time estimates were simulated and that the work still requires human oversight.
OpenAI’s report, titled “Advancing computer use with Ironclad,” describes a collaboration with the contract-management software company. Ironclad staff and OpenAI employees who use the product selected 11 tasks, including creating nondisclosure agreements, configuring procurement approval processes and updating a reusable contract clause based on a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task.
The tasks were graded against rubrics of 8 to 50 criteria, depending on complexity. OpenAI says Ironclad provided hosted product copies for model practice, and that the training tasks were synthetic, built using publicly filed contracts from the SEC’s EDGAR database and filtered to remove personal information. The company says it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data.
OpenAI reported that an earlier model, GPT-5.6 Sol, met an average 41.6% of criteria, compared with 55.0% for GPT-6 Astra. It also listed simulated estimated completion times of 37.0 minutes and 19.2 minutes, respectively. An internal model used during Astra’s development reached 63.7%, the report says. On one example task, Astra met about 94% of the criteria; that result is not the average across the 11 tasks.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Workflow Accuracy Matters
The result points to a potential way for AI systems to learn tasks inside specialised business software, rather than only answering questions or operating generic interfaces. OpenAI is also inviting a small number of software companies to work with it on tasks current agents cannot reliably complete. If that effort expands, software vendors could help shape how future models handle workflows specific to their products.
But the reported score does not establish that Astra can safely run contract or procurement processes without review. A percentage of criteria met is not a percentage of tasks completed successfully. If a workflow misses a required Finance approval, a Security review or a Legal check, the steps it did complete may not make the process fit for use. OpenAI’s own report says losing track of a business rule can limit what agents can be trusted to do, and that human oversight remains necessary.
The reported reduction in task time also is not evidence of customer savings: OpenAI says the estimates are simulated, based on assumed processing and generation speeds, and apply to these research tasks. For businesses considering agents, the relevant questions include which individual requirements failed, how errors are caught, and whether a person must verify every result.
contract management software with AI integration
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How OpenAI Tested Ironclad Tasks
Ironclad is a contract-management software company, not, as some news trackers reportedly inferred from the project name, a new hardened agent framework. OpenAI framed the work around teaching models to understand business rules, complete multi-step tasks in specialised software, and check whether the finished work meets the original requirements.
The report describes a research evaluation rather than a general product release or a claim that an autonomous agent is ready to handle customer contracts. Its 11 tasks were selected with people familiar with the product and its legal, commercial and procurement workflows. The reported scores measure how many rubric criteria the model met, not whether users adopted the system or whether it produced verified productivity gains in live operations.
OpenAI’s report also presents the work as an invitation to other software companies. Potential partners are asked to provide a concrete example of a task agents struggle with, experts who understand the work, a secure test environment and data that can safely be used for research. The report argues that a full contracting platform remains important because its business rules and controls govern the work an agent performs.
AI-powered legal document automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Evaluation Does Not Show
The report does not establish how Astra would perform across Ironclad’s full product, other companies’ software, or live customer workflows. It does not provide enough information here to identify which specific rubric criteria were missed across all 11 tasks, how often errors could cause a material business or legal problem, or how much review each task would require.
The 19.2-minute figure is a simulation, not a measured time for customers using Astra. OpenAI says it rests on assumed processing and generation speeds and covers the research tasks, so it should not be read as a demonstrated reduction in real-world contract work. The report also does not establish that the model can be deployed without human review; OpenAI says oversight remains necessary.
OpenAI says the work used public contract filings and no non-public Ironclad customer data. The supplied report does not detail the full data-filtering process or provide independent verification of that account. It also does not specify which software companies may join the proposed research effort or when further results will be released.
As an affiliate, we earn on qualifying purchases.
Further Tests and Vendor Partnerships
OpenAI says it plans to work with a small number of software companies on tasks current agents cannot reliably complete. The next useful evidence would include results from additional tasks and products, clearer reporting of which criteria agents miss, and details on how human reviewers detect and correct failures.
For businesses using contract or procurement software, the report offers a reason to ask vendors for task-level performance data, safeguards and review procedures before relying on agents in consequential workflows. OpenAI has not provided a public deployment schedule for Astra in Ironclad, a broader rollout plan, or a timetable for the next evaluation. Those details remain pending.
enterprise AI legal workflow tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI do with Ironclad?
OpenAI says it trained GPT-6 Astra on 11 legal, commercial and procurement tasks using hosted copies of Ironclad’s contract-management software.
What does Astra’s 55% score mean?
It is the average share of rubric criteria met across the evaluation, not the share of tasks completed successfully. The score does not show that the workflows are ready to run without review.
Did OpenAI measure a 19-minute time saving for customers?
No. OpenAI described the task times as simulated estimates based on assumed processing and generation speeds. They are not measured customer productivity results.
Did the training use private Ironclad customer contracts?
OpenAI says it used synthetic tasks based on publicly filed SEC EDGAR contracts, filtered to remove personal information, and did not use non-public Ironclad customer data. The report does not provide independent verification of that statement.
Can companies use Astra to handle contracts without human review?
The report does not establish that. OpenAI says human oversight remains necessary, and the average rubric score shows that the model did not meet all listed requirements across the evaluation.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
