AI engineering · Insurance
AI engineering for insurance: how production agents are built
AI agents that make insurance decisions in production are built from five parts: document reading that survives real-world files, a rules layer that sets each agent's authority, models that do the reasoning, integrations into the policy system, and a record of every action. This is how Native builds them, using the system that now decides most of TIE's endorsements as the example.
Updated · 3 min read
The architecture
Each file goes through the same four stages. Every stage is logged, so any decision can be traced back to the documents and the rules that produced it.
| Stage | What happens | Built with (at TIE) |
|---|---|---|
| Arrive | The document pack lands in the policy system or an inbox | Surefyre |
| Extract | Each document is classified from its contents and read | Mistral OCR, with a vision model as fallback |
| Route and check | The file goes to the right agent, which runs its checks and queries outside data | Claude Sonnet agents, FMCSA SAFER, n8n |
| Decide or refer | Settled by the rules, or handed to an underwriter pre-analysed | Rules engine, review queue, audit log on AWS |
Reading documents that were never meant for machines
Insurance documents arrive as scanned ACORD forms, loss runs from a dozen carriers, and driver licences photographed on a phone. Single files often mix several document types. A classifier works out what each page is from its contents rather than its file name, and a second model reads anything the first one cannot.
Missing documents are detected the moment a file arrives, so the request goes out the same day instead of after an underwriter opens it.
Outside data, checked mid-process
Agents query external sources while they work. At TIE they check the federal SAFER database on every relevant file, which catches carriers that re-register under a new DOT number to shed a bad safety record. The same pattern applies to MVR providers, sanctions lists and licensing databases.
Testing before production, and after
- Historical decisions become the test set: the agent must reach the decision an underwriter reached, or refer.
- Every change to a prompt, model or rule is run against that set before release.
- In production, samples of automated decisions are reviewed by underwriters each week.
- Cost and time per decision are tracked alongside accuracy.
Security and ownership
The system runs in your cloud account, under your credentials, with your data never used to train a model. The code is yours, documented, and your team is trained to run and extend it.

Case study · TIE
TIE decides transportation insurance endorsements in under a minute.
Native helped TIE, a US transportation insurance MGA, put agents on every submission and endorsement pack. They check each one and either decide it or hand the underwriter a pre-filled decision, and 74% of decisions now need no underwriter.
Frequently asked questions
What models do insurance AI agents use?
Usually a mix: an OCR model for documents, a vision model for poor scans, and a reasoning model such as Claude for the checks. Models are chosen per task and can be swapped without rebuilding the system.
How do you stop an AI agent from making a decision it shouldn't?
Each agent's authority is written down and enforced in a rules layer outside the model. Anything outside its scope is referred to a person with the analysis attached.
Does the AI integrate with our policy administration system?
Yes. Agents work through the system's API or its inbox, reading files where they land and writing decisions back. At TIE that system is Surefyre; the same approach works with platforms such as Guidewire, Duck Creek and Applied Epic.
How is accuracy tested?
Against your own historical decisions before go-live, against every change before release, and by weekly underwriter review of sampled decisions in production.
Where does the data live?
In your cloud account, under your credentials. Nothing is used to train third-party models.
