🔍 Read the full analysis: Could OpenAI Train Agents In Your Software? Read Ironclad’s Fine Print on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI’s October 6 post describes training GPT-6 Astra in hosted copies of Ironclad’s contract-management software using 11 legal, commercial and procurement tasks. Astra met an average 55% of task criteria, while its estimated completion times were simulated rather than measured customer savings. The work signals a push to train agents with software vendors, but does not establish that agents can safely run contract workflows without human review.
OpenAI said on October 6 that it worked with contract-management company Ironclad to train and evaluate a frontier model on tasks inside hosted copies of Ironclad’s software. The model, identified in the source material as GPT-6 Astra, met an average 55% of the criteria across 11 tasks, a result that shows progress on specialized workflows but does not establish readiness for unsupervised contract work.
Ironclad staff and OpenAI employees selected 11 tasks spanning legal, commercial and procurement work. Examples included setting up nondisclosure agreements, building procurement approval processes and updating a reusable contract clause to match a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task. The tasks were graded against rubrics containing 8 to 50 criteria, depending on complexity.
OpenAI reported that GPT-6 Astra met an average 55.0% of the criteria, compared with 41.6% for GPT-5.6 Sol in its high setting. An internal model used during Astra’s development reached 63.7%. On one showcase task, Astra met about 94% of the criteria. These figures describe rubric criteria met across the evaluation; they do not mean Astra completed 55% of the tasks successfully.
OpenAI also reported estimated times of 19.2 minutes per attempt for Astra and 37.0 minutes for GPT-5.6 Sol. The company’s footnote says those figures are simulated estimates based on assumed processing and generation speeds, not measured customer time savings. OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information, and did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Workflow Accuracy Matters
The result matters because contract and procurement workflows are judged by whether they meet all required controls, not by an average score alone. A procurement process may need Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. Missing even one such rule can send a purchase down an unauthorized path. A score showing that a model met just over half of the rubric criteria does not identify which requirements it missed or whether any individual output was safe to use.
For companies buying software, the announcement is a prompt to ask how agents are evaluated before they handle consequential work. Buyers need to know which checks were passed, which failed, how exceptions are handled and who reviews the output. OpenAI’s description says human oversight still matters when an agent loses track of a business rule. The reported results do not show a customer deployment or prove that the system can complete these workflows reliably without review.
For software vendors, working with model developers could help improve agent performance on difficult tasks in their products. It also has a strategic implication: as agents become better at operating software, customers may interact less with its screens. A vendor’s lasting value may depend more on the reliability of its rules, records, permissions and audit trails than on the interface an agent can operate.
As an affiliate, we earn on qualifying purchases.
How the Ironclad Evaluation Worked
The OpenAI post was one of two publications the source says appeared on October 6; the other described 722 mathematics manuscripts and drew more attention. Some AI news trackers reportedly mistook “Ironclad” in the title for a new agent framework. The company involved is instead a contract-management software provider, and the post describes an effort to train and test models using its product and workflows.
Ironclad provided hosted copies of its software for models to practise in. OpenAI and Ironclad chose tasks meant to reflect work carried out in legal, commercial and procurement settings, then assessed model outputs against task-specific criteria. The evaluation therefore offers evidence about performance on a limited set of defined tasks in a test environment; it is not a general benchmark of contract software, a live customer trial or proof of broad workplace productivity gains.
OpenAI’s post also invited a small number of software companies to work with it on tasks that agents cannot reliably complete today. It said potential partners should bring a concrete example of a failing task, people with deep knowledge of the work, a secure test environment and data that can safely be used for research. That invitation suggests the Ironclad project is being presented as a possible model for future collaborations, though the source provides no list of other partners or timetable.
As an affiliate, we earn on qualifying purchases.
What the Scores Do Not Show
The published averages do not disclose which criteria Astra missed on each task, how often an individual workflow met every required condition, or how performance varied across task types. The roughly 94% result applies to one showcase task and should not be read as the model’s general accuracy. The source also does not describe the size or statistical uncertainty of the evaluation beyond the 11 selected tasks.
It remains unclear whether Astra has been deployed for Ironclad customers, how it would perform on live or unfamiliar matters, and what review or permission controls would govern any real use. The time figures are simulated, so they cannot establish that customers save time. OpenAI’s statement about data use is its account of the training process; the source material does not provide an independent audit of that claim.
The announcement also leaves open how future vendor partnerships would handle data access, model improvement, customer consent and responsibility for errors. No specific additional partner, product release date or commercial deployment was identified in the source.
AI-powered contract review software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What OpenAI’s Partner Call Signals
OpenAI said it is seeking a small number of software-company partners to work on tasks current agents struggle to complete. The next step described is not a general product launch, but identifying concrete failure cases and providing expert input, a secure environment and research-appropriate data. The company has not said when those partnerships will begin or when any resulting models or tools would become available.
For businesses considering agents in contract, finance or customer-record systems, the immediate next step is careful evaluation before granting access. They can ask vendors for task-level results, the criteria used to grade work, the handling of failed checks, the human review process and the boundaries on data use. OpenAI’s reported experiment points to continued development, but any deployment decision should account for the possibility of missed rules and the need to protect sensitive records.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI and Ironclad announce?
OpenAI described work with Ironclad, a contract-management software company, to train and evaluate a model on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s product.
Does a 55% score mean the agent completed 55% of the tasks?
No. OpenAI reported that GPT-6 Astra met an average 55% of the rubric criteria across the tasks. That is not the share of tasks fully completed, and the published average does not specify which requirements were missed.
Did the experiment prove that agents save customers time?
No. OpenAI said the reported times—19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol per attempt—were simulated estimates based on assumed processing and generation speeds, not measured customer savings.
Can companies use these agents without human review?
The results do not establish that. The source says human oversight remains important when an agent fails to retain a business rule, and a missed approval or contract condition could have practical consequences. The announcement does not describe a customer deployment without review.
What data did OpenAI say it used?
OpenAI said it built synthetic tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It also said it used no OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
