OpenAI has entered into a strategic partnership with contract management platform Ironclad to train and test artificial intelligence agents on complex contracting procedures. This initiative aims to transform high-value legal and business tasks into advanced research and application environments.
- Advanced research program to train agents on complex business tasks
- New model performance and outperformance of previous versions
- Synthetic training environment and customer data privacy assurance
- Need for human oversight and broader software partnership evolution
- Frequently asked questions
Advanced research program to train agents on complex business tasks
OpenAI announced on Tuesday, October 6, that it has concluded a qualitative cooperation agreement with Ironclad, a company specializing in the development of digital contract management and auditing software, with the goal of training and testing independent AI agents on complex legal and business workflows. Ironclad is the first partner to join a new research program launched by the developer, aimed at close collaboration with selected specialized software companies to transform arduous, high-economic-value business tasks into research and training challenges to upgrade the capabilities of generative models.
OpenAI explained that its advanced model, “GPT-6 Astra,” represents the first flagship model trained directly and intensively on Ironclad’s contracting tasks. Internal research evaluation results showed a remarkable performance superiority for the new model, achieving an average score 32% higher compared to the previous generation, “GPT-5.6 Sol,” alongside a decisive drop in the estimated time taken to execute a single attempt by 48%, reflecting a clear leap in processing efficiency and transaction completion speed.
New model performance and outperformance of previous versions
OpenAI researchers collaborated with Ironclad employees, as well as legal teams who use the platform daily, to identify 11 core business tasks covering legal affairs, commercial procurement, and administrative compliance. These tasks included preparing non-disclosure and confidentiality agreements, building regulatory approval chains for major procurement, and updating reused legal terms and conditions to accurately comply with user-defined local jurisdictions and laws—tasks that typically take an experienced human specialist 30 to 40 minutes.
Each task underwent rigorous evaluation based on strict criteria ranging from 8 to 50 verification checks for performance integrity. Across these 11 tests, the new Astra model recorded an average success rate of 55.0%, compared to about 41.6% for the Sol version, while the average time required per attempt dropped from 37 minutes to just 19.2 minutes. An advanced internal model used during Astra’s development phases also achieved a 63.7% success rate, though the company noted that these measurements represent specific virtual simulations and do not constitute calculated time savings for the platform’s commercial enterprise customers generally.
Synthetic training environment and customer data privacy assurance
Ironclad provided dedicated, simulated cloud workspaces mirroring its software to enable smart models to practice decision-making mechanisms and procedural step execution. OpenAI researchers relied on designing advanced synthetic training tasks and applying reinforcement learning techniques to guide agent behavior. Virtual training operations relied on publicly available commercial contracts and documents officially accessible within the SEC’s EDGAR database, with content carefully filtered to remove any sensitive personal data or information.
Both companies emphasized in a joint statement that none of OpenAI’s customer data, internal contracts and communications, or Ironclad’s confidential customer data were used in training and algorithm testing operations. This strict commitment to privacy protection forms an essential element for building trust in automated solutions within legal environments that demand rigorous confidentiality for mutual agreements and business transactions.
Need for human oversight and broader software partnership evolution
Research findings indicated that the loss of a single business rule by a software agent during multi-stage execution tangibly limits institutional trust in relying on it entirely without supervision. For this reason, human oversight and auditing remain the decisive safety valve in legal workflows even with steady improvements in algorithmic performance. Sunita Verma, CTO at Ironclad, emphasized that successful contract AI must comprehend the full lifecycle of a contract and how operational stages interconnect while maintaining the regulatory controls that departments rely on.
This partnership builds on an extended history between the two sides, as Ironclad built its smart contract review tool on OpenAI models starting in 2023 and launched a technical integration enabling contract data querying via ChatGPT earlier this year. This step coincides with OpenAI announcing an expanded partnership with Atlassian, issuing an open invitation for software companies to join the program and submit precise business tasks that current agents cannot reliably complete, contributing to the innovation of the next generation of automated solutions.
Frequently asked questions
Question: What results did the Astra model achieve in legal contract tasks?
Answer: The model outperformed its predecessor Sol by 32%, achieving 55% execution accuracy, and cut the estimated time required to execute tasks by 48%.
Question: How were data confidentiality and documents protected during algorithm training?
Answer: Training relied on synthetic tasks built on public documents from the US Securities and Exchange Commission, without using customer data or private contracts.
Question: Why is auditing and human oversight necessary in smart contract management?
Answer: Because any omission of a single business rule could disrupt the entire contractual path, necessitating human personnel to ensure the integrity of legal controls.