Our client is a US-based insurtech organization operating as a Managing General Agent (MGA) platform across multiple commercial insurance programs.
The company manages complex contractual relationships involving MGAs, carriers, underwriting programs, products, territories, and compliance obligations. Its operations depend heavily on large volumes of insurance agreements, endorsements, amendments, underwriting guidelines, and coverage documents that drive underwriting, sales, governance, and compliance workflows.
As contract volumes increased, our client needed a more scalable way, through AI integration and adoption, to structure, manage, and operationalize the intelligence locked inside these documents.
Our client managed large volumes of highly detailed MGA agreements, endorsements, underwriting guidelines, amendments, exclusions, and territory-specific clauses.
Critical business information was buried within lengthy, unstructured insurance agreements that frequently exceeded 100 pages and evolved continuously through amendments, renewals, and guideline revisions. The data inside the documents drove decisions across sales, underwriting, and compliance, yet none of it was structured, stored, or queryable.
Each contract was read, interpreted, and manually extracted by trained resources who spent several hundred hours a month.
The issue our client faced wasn’t limited to extracting text from PDFs using LLMs. They wanted visibility into the economics of production-scale AI before committing to deployment. Token usage, infrastructure sizing, model costs, and document processing expenses all needed to be validated upfront.
The contracts contained highly variable insurance language, amendment relationships, nested clauses, and multiple versions of the same contractual obligations. Traditional extraction approaches could not reliably structure the information without losing context.
The obvious answer was AI, yes. But that opened a different can of worms.
As long as contract data lived only in PDFs, our client could not build reliable downstream intelligence. Sales teams could not query coverage details. Underwriting could not access historical terms. Compliance had no audit trail. The platform was sitting on a goldmine of structured intelligence locked inside documents that no system could read.
Our client faced four compounding problems with no clear path forward
100+ page documents · monthly amendments · 40+ business-critical attributes locked inside
Each contract required 4+ hours to review.
100+ contracts monthly = 800+ person-hours.
Volume was growing with no ceiling in sight.
Early tests cost ~$20 per contract.
Token usage in production couldn’t be forecast.
No cost model meant no deployment approval.
Financial risk made scaling impossible.
Insurance is a regulated environment.
Every extracted value needed a citation — page-level, clause-level, section-level.
Uncited AI outputs were unusable in audits.
New contracts and amendments arrive daily.
AI had no connection to live workflows.
Underwriting and sales ran on separate rails.
No pipeline meant results stayed siloed.
Trigent approached the engagement as an AI cost validation and production engineering exercise, limited not only to document and contract intelligence.
We began by first analyzing the client’s existing contract ecosystem, including MGA agreements, amendments, underwriting guidelines, exclusions, and territory-specific documents. Our team studied how information moved across underwriting, compliance, sales, and governance workflows and identified the exact contract attributes required downstream.
Our team standardized more than 40 business-critical contract attributes, including underwriting rules, carrier information, product definitions, amendment histories, clause language, territorial restrictions, and eligibility conditions. A relational data structure was created to preserve relationships between contracts, amendments, clauses, and downstream operational entities.
During the evaluation phase, we tested multiple AI approaches and hyperscaler platforms against real 90–100-page MGA contracts.
Azure AI Foundry and AWS Bedrock were evaluated first using agentic workflows and commercially available LLMs. While both platforms demonstrated acceptable early extraction capabilities, several production limitations became clear during testing:
The initial AI workflows our client used processed contracts at nearly $20 per document during evaluation. It became apparent that processing costs could increase rapidly across large contract volumes without tighter orchestration and token optimization.
Trigent then shifted the evaluation framework to ArkOS AI Workbench to gain greater flexibility across models, orchestration logic, prompt engineering, and workflow routing.
Using ArkOS, Trigent benchmarked multiple frontier models against the complete 40-attribute extraction workload. Each model was evaluated on:
Trigent’s AI team evaluated multiple Frontier Models that were commercially available at that time and documented the performance in terms of its accuracy and cost involved to process a given contract document. The findings after testing multiple LLMs are detailed below:
| Frontier Model | Data Extraction Accuracy % | Speed | Comparative Cost |
|---|---|---|---|
| Llama 3.2, 3.3 | Average – 85% | Slow | Low |
| Qwen 2.5 | Average – 80% | Very Slow | Low |
| DeepSeek R1 / V3 | Poor – <70% | Slow | Low |
| Claude 4.5 Sonnet | Very Good – 95% | Fast | Very High |
| Open AI GPT 4o mini | Poor – <70% | Average | Low |
| Open AI GPT nano | Poor – <70% | Average | Low |
| Open AI GPT 5 nano | Average – 82% | Average | Medium |
| Open AI GPT 5.1 | Good – 89% | Fast | High |
| GPT 5.2Â (Selected) | Very Good – 95% | Fast | Medium |
OpenAI GPT 5.2 ultimately delivered the strongest balance between extraction accuracy and processing economics.
Even with the right model selected, processing performance remained a challenge initially.
Early single-agent workflows forced the model to process entire 90–100-page contracts sequentially. This increased token usage, slowed processing times, and introduced contextual inefficiencies because the model was analyzing large sections of irrelevant information during every extraction cycle.
To solve this, Trigent redesigned the system around a specialized multi-agent architecture.
Instead of using a single extraction workflow, the platform deployed dedicated AI agents responsible for specific contract functions such as:
Contract classification
Amendment reconciliation
Company and carrier extraction
Clause extraction
Product extraction
Underwriting guideline extraction
Contract-level metadata processing
Each agent processed only the sections of the contract relevant to its responsibility.
This architectural shift fundamentally changed the platform’s economics and reliability.
By narrowing context windows and reducing unnecessary token processing:
Trigent then optimized the platform further through specialized prompt engineering and workflow refinement.
Each agent was equipped with custom prompts tailored to specific insurance extraction tasks, amendment logic, and clause interpretation requirements. The system was also designed to preserve full source traceability, allowing every extracted field to map directly back to the originating contract language and page reference.
The final solution delivered a production-grade contract intelligence platform capable of transforming large unstructured insurance agreements into structured, queryable, audit-ready business data while maintaining predictable operating economics at scale.

We are happy to answer any questions you may have.
Fill in your details below and our team will get back to you shortly.