Skip to content
KAZES.studio — home
All projects

Enterprise automation

Concept study

Automating the first pass of claims adjudication

An internal system that handles the mechanical work of first-pass adjudication so human reviewers spend their time on the cases that actually need judgement.

Engagement
Dedicated engineering team
Duration
Four months to production, then embedded support
Year
2025
Client
Not applicable
0255075100125150175200MS010203040506
Request lifecycle and timing budget.. This is a generated schematic, not a product screenshot.

The challenge

What made this difficult.

A large share of first-pass adjudication was mechanical: coverage verification, duplicate detection, document completeness, routine policy lookups. Highly trained reviewers were spending most of their time on that work, and a backlog was growing faster than headcount. Manual-only automation was not the answer, because the cases that genuinely required judgement were being mixed in with the ones that did not.

Constraints

Non-negotiables we designed around.

01

A human must remain accountable

Every decision needs an attributable reviewer. The system prepares and recommends; it does not adjudicate.

02

Unstructured, inconsistent source documents

Claims arrive with documents of widely varying quality and structure, and 'complete' means different things in different queues.

03

Policy changes take effect immediately

A rules engine baked into a deploy cycle could not keep pace with policy updates.

04

Legacy systems of record

The authoritative record lived in systems that could be read but not modified, so the new system had to be correct on first write.

Approach

How we would build it.

Separated automation from authority

The system produces a prepared case: extracted facts, matched policy, coverage reasoning, and a recommended disposition with confidence. Reviewers approve or override. Authority stays with a person; the mechanical load does not.

Extracted facts with source spans, not summaries

Structured extractions carry the exact location they came from. A reviewer can verify a field without re-reading the document, and a wrong extraction is visibly wrong rather than plausibly wrong.

Made policy an externalised, versioned artefact

Policy rules live outside the deploy cycle, are versioned independently, and every decision records the policy version applied. Changing policy stops being an engineering event.

Measured override rate as the primary metric

Accuracy and speed both matter, but override rate on recommended dispositions is the signal that tells you whether the system is learning the job or confidently guessing. It was the number we reviewed weekly.

Deployed beside the systems of record

The new system read authoritative state and produced prepared cases without taking on write authority, which avoided a migration while still making the reviewer experience coherent.

Architecture

How the pieces fit together.

Architecture

Preparation pipeline. The system produces a prepared case with attributable extractions; authority for disposition stays with a human reviewer.

  1. Intake

    • Document ingestion
    • Queue routing
    • Completeness assessment
    • Duplicate detection
  2. Extraction

    • Layout-aware parsing
    • Field extraction with spans
    • Confidence scoring
    • Manual review queue
  3. Policy

    • Externalised rules
    • Versioned policy artefacts
    • Coverage evaluation
    • Effective-dating
  4. Preparation

    • Recommendation
    • Reasoning trace
    • Prepared case assembly
    • Reviewer assignment
  5. Decision

    • Reviewer approval
    • Override capture
    • Authoritative record write
    • Feedback into evaluation set

Stack

What it would run on.

Interface

  • TypeScript
  • React
  • Python

Data

  • PostgreSQL
  • Document object storage
  • Queue-backed workers
  • Append-only decision log

Infrastructure

  • Container services
  • Secrets management
  • Structured logging
  • Infrastructure as code

Expected outcomes

What success would look like.

Accountability

Human-held

Every disposition is attributable to a named reviewer. The system prepares; it does not adjudicate.

Policy updates

Decoupled

Policy rules are versioned outside the deploy cycle, and each decision records the policy version applied.

Quality tracking

Override rate

Override rate on recommended dispositions was the primary weekly signal — it exposes confident guessing faster than any accuracy metric.

Correction loop

Closed

Overrides were fed back into the evaluation set, so the system was measured against the cases practitioners actually rejected.

What we learned

The conclusions we would carry forward.

Separating preparation from authority made the automation acceptable to the reviewers who had to live with it.

Span-bound extractions made verification cheap, which is what determined whether reviewers actually checked the system's work.

Externalising policy out of the deploy cycle removed engineering from the critical path of routine business changes.