🔍 Read the full analysis: A Guide To Ironclad’s Fine Print On OpenAI Training Agents In Software on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
OpenAI described training a frontier model in hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task criteria, while its reported time estimates are simulated—not measured customer savings—and the results do not establish that the workflows are ready to run without human review.
OpenAI published details on October 6 of work with contract-management company Ironclad to train a frontier model on workflows inside the company’s software. In tests covering 11 legal, commercial and procurement tasks, GPT-6 Astra met an average 55% of the listed criteria; OpenAI says the accompanying task-time figures are simulated estimates, not measured customer savings.
Ironclad staff and OpenAI employees who use its product selected tasks including setting up nondisclosure agreements, building procurement approval processes and adapting reusable contract clauses to a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. Depending on difficulty, evaluators scored each task against 8 to 50 criteria, so the reported result is a share of requirements met—not a count of tasks completed successfully.
Ironclad supplied hosted copies of its software for model practice. OpenAI says it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It also says it used no OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data. The post identifies GPT-6 Astra as the first frontier model trained through this approach.
OpenAI reports that GPT-5.6 Sol, at a high setting, met an average 41.6% of criteria, compared with 55.0% for GPT-6 Astra at a maximum setting. The Astra attempts were estimated at 19.2 minutes, versus 37.0 minutes for Sol. An internal OpenAI model used during Astra’s development reached 63.7%. On one highlighted task, Astra met about 94% of criteria. These are results from the stated research tasks, not evidence of performance across all Ironclad work.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Workflow Accuracy Matters
The results matter because contract and procurement workflows depend on specific controls, not just a broadly plausible outcome. A procurement process may require Finance approval above a spending threshold, Security review for certain requests and Legal review for nonstandard terms. Missing one of those rules can send a request down an improper path, even if an agent gets the rest of the workflow right. A 55% average criteria score therefore should not be read as 55% of work being safely complete.
OpenAI’s post says human oversight remains necessary when an agent can lose track of a business rule during a multi-step task. That distinction is material to companies considering access for AI agents to contracts, financial records or customer systems: evaluation needs to reveal which requirements failed, not just provide one average score. The report describes a research result, not a deployment recommendation or a guarantee of reliability.
There is also a business question for software providers. The work suggests vendors may help train models to operate directly inside specialised products. That could make software more useful, but it may also shift the customer’s interaction away from screens and toward agents. In that case, a provider’s value may depend more on its business rules, records, audit trails and controls than on the interface an agent learns to use. That is an implication of the collaboration, not a reported change in how Ironclad customers currently work.
As an affiliate, we earn on qualifying purchases.
How OpenAI Tested Ironclad Workflows
The Ironclad work sits within OpenAI’s effort to develop models that can handle multi-step work in specialised software, understand business rules and check whether completed work matches its requirements. The October 6 post described a concrete test setting rather than a general benchmark: a small set of tasks in one contract-management product, scored against task-specific rubrics.
OpenAI’s reported time comparison needs particular care. The company says the figures are simulated estimates based on assumed processing and generation speeds, and that they cover these 11 research tasks. They are not observed time savings for Ironclad customers or a measurement of end-to-end work in ordinary business use. A faster attempt that leaves important criteria unmet may still require substantial human checking.
The post also invites a small number of software companies to partner on tasks current agents cannot reliably complete. It asks prospective partners to bring a concrete example of failure, people with deep knowledge of the work, a secure testing environment and data that can safely be used for research. This frames the Ironclad project as a possible model for further software-specific training, rather than evidence that a broad program or set of partnerships is already in place.
What the Test Scores Cannot Show
The published average does not identify, in the supplied material, which individual criteria the model missed across the tasks or how often a miss would create a serious business or compliance problem. The 94% result applies to one showcase task and does not establish consistent performance across other tasks. It is also unclear how the models would perform on different companies’ policies, software configurations or real customer work.
The source gives no measured customer time savings, production deployment results or evidence that an agent can complete these workflows without review. OpenAI says it did not use non-public Ironclad customer data, but the material provided does not describe every security, monitoring or access-control arrangement for future partner projects. The scope and timing of any additional collaborations are also not specified.
What Future Software Partnerships Require
OpenAI says it is seeking a small number of software-company partners to test difficult tasks that current agents do not reliably complete. Prospective partners are expected to provide specific failure examples, domain experts, secure test environments and research-appropriate data. The company has not announced a schedule, named further partners in the supplied material or set out a deployment milestone.
For software buyers, the immediate practical question is how vendors disclose and evaluate agent performance: which requirements were missed, what human review is required and what safeguards prevent an agent from bypassing approvals. Further public results would help establish whether the gains reported in Ironclad’s limited task set hold up across more workflows and real operating conditions.
Key Questions
What did OpenAI and Ironclad announce?
OpenAI published details of research training and evaluating a model in hosted copies of Ironclad’s contract-management software. The work used 11 tasks spanning legal, commercial and procurement workflows.
Does a 55% score mean GPT-6 Astra completed 55% of the tasks?
No. OpenAI reports that Astra met an average 55% of rubric criteria across the tasks. That is not the share of tasks completed, and it does not show that the missed requirements were unimportant.
Were the reported time savings measured with customers?
No. OpenAI describes the times as simulated estimates based on assumed processing and generation speeds. They cover the 11 research tasks and are not measured customer productivity gains.
Did OpenAI use Ironclad customer contracts to train the model?
OpenAI says it used synthetic tasks based on publicly filed contracts in the SEC’s EDGAR database and filtered out personal information. It says it did not use non-public Ironclad customer data, OpenAI customer data or OpenAI internal contracts.
Can businesses rely on agents to run these workflows without review?
The reported results do not establish that. The average criteria score is incomplete for judging risk, and OpenAI’s post says human oversight remains necessary when agents may fail to retain business rules.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
