AI systems
Concept studyA retrieval layer that clinical reviewers actually trusted
An evidence-grounded assistant for trial protocol review, built so that every answer could be traced to source and every change to a prompt could be measured before release.
- Engagement
- Dedicated engineering team
- Duration
- Five months, discovery through deployment
- Year
- 2025
- Client
- Not applicable
The challenge
What made this difficult.
Reviewers were spending hours cross-checking protocol language against prior amendments and precedent documents. A first-generation assistant had been abandoned because it produced fluent answers that could not be verified, and reviewers had no way to tell a reasonable inference from an invented one. The technical problem was not summarisation — it was provenance. Without traceability, no amount of answer quality would make the system usable in a regulated workflow.
Constraints
Non-negotiables we designed around.
Every claim must be attributable
Answers are only useful if a reviewer can jump to the exact clause that supports them, and confirm nothing was paraphrased in transit.
The corpus is versioned and amended
Protocols accrue amendments. A retrieval index that ignores document lineage will happily return superseded guidance.
Regulated domain, low tolerance for drift
Behaviour changes between model versions are a compliance concern, not just a quality concern.
Existing document infrastructure
Content already lived in a permissions-aware repository that had to remain the source of truth.
Approach
How we would build it.
Made provenance a hard requirement, not a feature
The retrieval layer returns structured spans — document, version, clause range — rather than free text. The interface is designed so a claim without a resolvable span cannot be displayed, which removes the failure mode structurally instead of asking the model to behave.
Modelled document lineage as first-class metadata
Amendment chains are represented explicitly in the index. Superseded clauses are retained but demoted, and the citation surfaces the amendment that introduced the language.
Built the evaluation harness before the assistant
A fixed set of reviewer-authored question-and-answer pairs, with unanswerable questions deliberately included. Every prompt, retrieval, and model change runs against this set as a gate. Quality became a number that could be compared across releases.
Isolated model providers behind a boundary
Inference sits behind an internal interface with prompt and retrieval versioning, so a provider upgrade is a controlled operation with a rollback rather than a surprise.
Designed the autonomy boundary explicitly
The system retrieves, cites, and drafts. A qualified reviewer approves anything that reaches the record. That boundary is enforced in code and stated in the interface, not left to convention.
Architecture
How the pieces fit together.
Retrieval and evaluation architecture. Document lineage is represented in the index rather than inferred at query time.
Sources
- Protocol repository
- Amendment history
- Precedent set
- Reviewer annotations
Ingestion
- Structure-aware parsing
- Lineage graph
- Clause segmentation
- Change detection
Retrieval
- Hybrid keyword + dense
- Version-aware filtering
- Span extraction
- Ranking with recency prior
Generation
- Provider abstraction
- Prompt versioning
- Output schema validation
- Span enforcement
Assurance
- Evaluation set
- Regression gate in CI
- Trace store
- Reviewer checkpoint
Stack
What it would run on.
Interface
- TypeScript
- React
- Python
- PostgreSQL
Retrieval
- Hybrid lexical + dense search
- Structured clause index
- Document lineage graph
Infrastructure
- Containerised services
- Object storage for raw documents
- Tracing and cost telemetry
- CI with evaluation gate
Expected outcomes
What success would look like.
Answer attribution
100%
Design target: every displayed claim resolves to a citable source span. Structurally enforced, not model-dependent.
Unanswerable handling
Refusal tested
Deliberately included unanswerable cases in the evaluation set to measure refusal behaviour rather than answer volume.
Change safety
Gated
Prompt, retrieval, and model changes must pass the fixed evaluation set before reaching reviewers.
What we learned
The conclusions we would carry forward.
In a regulated workflow, traceability is a design constraint on the interface, not a property you ask the model for.
Representing document lineage at index time was cheaper and more reliable than reconstructing version context at query time.
Building the evaluation harness first changed the conversation with the client about what 'good' meant — it stopped being a matter of taste.
Related
