Enterprise automation
Concept studyAutomating the first pass of claims adjudication
An internal system that handles the mechanical work of first-pass adjudication so human reviewers spend their time on the cases that actually need judgement.
- Engagement
- Dedicated engineering team
- Duration
- Four months to production, then embedded support
- Year
- 2025
- Client
- Not applicable
The challenge
What made this difficult.
A large share of first-pass adjudication was mechanical: coverage verification, duplicate detection, document completeness, routine policy lookups. Highly trained reviewers were spending most of their time on that work, and a backlog was growing faster than headcount. Manual-only automation was not the answer, because the cases that genuinely required judgement were being mixed in with the ones that did not.
Constraints
Non-negotiables we designed around.
A human must remain accountable
Every decision needs an attributable reviewer. The system prepares and recommends; it does not adjudicate.
Unstructured, inconsistent source documents
Claims arrive with documents of widely varying quality and structure, and 'complete' means different things in different queues.
Policy changes take effect immediately
A rules engine baked into a deploy cycle could not keep pace with policy updates.
Legacy systems of record
The authoritative record lived in systems that could be read but not modified, so the new system had to be correct on first write.
Approach
How we would build it.
Separated automation from authority
The system produces a prepared case: extracted facts, matched policy, coverage reasoning, and a recommended disposition with confidence. Reviewers approve or override. Authority stays with a person; the mechanical load does not.
Extracted facts with source spans, not summaries
Structured extractions carry the exact location they came from. A reviewer can verify a field without re-reading the document, and a wrong extraction is visibly wrong rather than plausibly wrong.
Made policy an externalised, versioned artefact
Policy rules live outside the deploy cycle, are versioned independently, and every decision records the policy version applied. Changing policy stops being an engineering event.
Measured override rate as the primary metric
Accuracy and speed both matter, but override rate on recommended dispositions is the signal that tells you whether the system is learning the job or confidently guessing. It was the number we reviewed weekly.
Deployed beside the systems of record
The new system read authoritative state and produced prepared cases without taking on write authority, which avoided a migration while still making the reviewer experience coherent.
Architecture
How the pieces fit together.
Preparation pipeline. The system produces a prepared case with attributable extractions; authority for disposition stays with a human reviewer.
Intake
- Document ingestion
- Queue routing
- Completeness assessment
- Duplicate detection
Extraction
- Layout-aware parsing
- Field extraction with spans
- Confidence scoring
- Manual review queue
Policy
- Externalised rules
- Versioned policy artefacts
- Coverage evaluation
- Effective-dating
Preparation
- Recommendation
- Reasoning trace
- Prepared case assembly
- Reviewer assignment
Decision
- Reviewer approval
- Override capture
- Authoritative record write
- Feedback into evaluation set
Stack
What it would run on.
Interface
- TypeScript
- React
- Python
Data
- PostgreSQL
- Document object storage
- Queue-backed workers
- Append-only decision log
Infrastructure
- Container services
- Secrets management
- Structured logging
- Infrastructure as code
Expected outcomes
What success would look like.
Accountability
Human-held
Every disposition is attributable to a named reviewer. The system prepares; it does not adjudicate.
Policy updates
Decoupled
Policy rules are versioned outside the deploy cycle, and each decision records the policy version applied.
Quality tracking
Override rate
Override rate on recommended dispositions was the primary weekly signal — it exposes confident guessing faster than any accuracy metric.
Correction loop
Closed
Overrides were fed back into the evaluation set, so the system was measured against the cases practitioners actually rejected.
What we learned
The conclusions we would carry forward.
Separating preparation from authority made the automation acceptable to the reviewers who had to live with it.
Span-bound extractions made verification cheap, which is what determined whether reviewers actually checked the system's work.
Externalising policy out of the deploy cycle removed engineering from the critical path of routine business changes.
Related
