We cut underwriting review by 89% and made it audit-ready
A Series B insurance platform was capped by how fast underwriters could read unstructured documents to pull a handful of decision-critical fields. We built a document-intelligence pipeline with OCR, a fine-tuned extraction model, and a RAG layer that grounds every field in its source passage. A human-in-the-loop review UI lets underwriters confirm in seconds.
- 89%
- Less manual review
- 99.2%
- Extraction accuracy
- <2s
- Per document
What they came with
Every submission arrived as a stack of documents, and someone had to read all of it to pull the handful of fields a decision actually turned on. How much business the platform could take on was set by how fast underwriters could read, not by how much came in. Scans, photographs and re-keyed forms all landed in the same queue at the same priority. When a decision was questioned months later, the reasoning behind it sat in a person's memory rather than in the file.
What the engagement covered
- Fine-tuned extraction grounded to source passages for full auditability
- Confidence scoring that routes only genuine edge cases to a human
- Every decision logged to continuously improve the model
Technical detail
Every field points at its source
Extraction returns each value with a pointer to the passage it came from, retrieved through a RAG layer over embeddings in pgvector. The reviewer sees the value next to the sentence that produced it, and the same pointer is what makes a decision auditable later.
Confidence decides what a human sees
Each field carries a confidence score from the fine-tuned model, and only fields below threshold are routed into the review queue. A submission that is clean throughout passes without anyone opening it, so reading time goes to the pages that are genuinely ambiguous.
OCR keeps the page layout
Scanned and photographed pages go through OCR with positional information retained, so tables and multi-column forms are not flattened into a single run of text. Field extraction can then use position as a signal, not just wording.
Reviews feed back into training
Every confirmation and correction is logged against the source passage and the model version that produced it. That gives a labelled set that grows with use and a record of how accuracy moved between fine-tunes.
The stack
AI
Data
Backend
Frontend
- Practice
- AI Document Intelligence
- Sector
- Insurance
- Shape
- Client engagement
- Stack
- OpenAI, RAG, pgvector
Something like this to build?
Tell us what runs today and where it hurts. An engineer reads it and replies.