Skip to content
← Selected work

Global Response Documents

An agentic AI system that drafts the evidence-cited documents a pharma company uses to answer scientific inquiries from doctors and researchers, with the medical writer in the loop.

Client work from a former company. The organisation and the product are kept deliberately general, and nothing shown here reproduces a real interface. Happy to talk through the engineering in person.

Client
A global pharma and R&D company
Role
AI Engineer
Timeline
Client engagement
Discipline
Agentic AI · Regulated Documentation
Illustrative artwork. Not a client interface, and no real content is shown.

Overview

When a physician or researcher asks a pharmaceutical company a scientific question about one of its products, the answer comes back as a Global Response Document. It is a structured, non-promotional paper in which every statement is grounded in and cited from published literature. The system drafted these documents from the publications a medical writer supplied, section by section and in conversation with the writer, and the dominant engineering constraints throughout were citation integrity and privacy.

Challenge & approach

01

The challenge

These documents have to satisfy three demands at once. They must read like scientific writing, they must be produced at the speed inquiries actually arrive, and they must hold up under review by experts whose job is to check every line. Writing one by hand is slow, highly skilled work, and inquiries arrive faster than that. The bottleneck was never the science. It was the time between a question arriving and a defensible document existing.

02

The approach

The medical writer stays in charge throughout. They bring the inquiry, the publications that can answer it and the settings for the document set; the system drafts the first sections, and the writer critiques them in a chat. Each round of feedback updates the draft, later sections are built on the earlier ones the same way, and at the end a finished document with its citations is ready to download and send. Behind the interface is an agentic AI system: several expert language models, each an agent with a small task and only the context it needs, working in parallel and handing their results to one another, shaped by a great deal of prompt engineering. Every output also carries a quality assessment, from a model acting as judge together with plain statistics anyone can check.

How it was built

Privacy set the rules before anything was built. The publications and the inquiries are sensitive, so nothing could be stored: no questions, no answers, no finished documents, and no fine-tuning on any of it. That ruled out most of the usual ways to make a system better over time and pushed the work into the architecture and the prompts instead, which is where the agentic AI design came from: instead of one model learning from data, a set of agents that reason well from the publications in front of them. Speed was the hard stretch. Drafting a full document through one long chain of model calls was too slow to live with. Splitting the work into smaller tasks with less context each, running the expert models in parallel, and letting the writer choose the number of improvement iterations in the interface changed both the waiting time and the quality of the result. The client was genuinely satisfied with what came out at the end. The stack was Python throughout: LangGraph for the agentic flow between the expert models, Langfuse to trace every step of it, all running on AWS. The interface began in Streamlit, which was the right call for prototyping with the writers and, honestly, stayed a little too long: we ended up building more complexity into it than Streamlit is usually asked to carry. Later the whole thing was integrated into another tool and rebuilt in React and TypeScript, which is where an interface like this belongs. Publications arrive as PDFs full of tables and formulas, so getting reliable text out of them mattered as much as anything downstream; Mathpix, Mistral OCR and a few other OCR tools were compared and combined for that. Quality metrics and usage KPIs were kept in DynamoDB and shown with Plotly, so the client could see how the system was doing without asking. We tested with the experts themselves: medical writers used the tool on their own inquiries, gave feedback in the chat as they went, and filled in a form afterwards with both scores and written comments. Those rounds shaped the sections, the feedback loop and the assessment.

How it worked

05 stages
  1. 01

    Inquiry

    The medical writer brings the question, the publications that can answer it, and the settings for the document set.

  2. 02

    Draft

    Expert models draft the first sections in parallel, every claim tied to the publication behind it.

  3. 03

    Refine

    The writer critiques each draft in a chat; every round updates it, and later sections build on the earlier ones.

  4. 04

    Assess

    Each output gets a quality assessment: a model as judge, plus statistics that are easy to check.

  5. 05

    Deliver

    A finished document with its citations, ready to download and send out.

In principle

Response document

Draft
1
2
3

Cited publications

1
2
3
Claim traced to source
Illustrative schematic. Not a client interface, and no real content is shown.

What it does

  • 01Section by section drafting from the publications the writer supplies
  • 02A feedback chat on every draft, with each iteration updating the text on the critique
  • 03Later sections built on the approved earlier ones, reviewed the same way
  • 04Every claim bound to its source, with citations carried into the final document
  • 05Quality assessment on every output: a model as judge plus checkable statistics
  • 06Improvement iterations selectable in the interface, with expert models running in parallel
  • 07A finished document, ready to download and send

Guiding principles

  • Nothing is stored: no inquiries, no answers, no documents, and no fine-tuning on any of it.
  • Every statement has to trace to published literature. No claim travels without its source.
  • Every reference has to be real and match its claim exactly.
  • The output is non-promotional and has to stay inside that boundary.
  • The medical writer stays in the loop at every step and has the last word.

Built with

  • Python
  • Streamlit (prototype)
  • React and TypeScript (final interface)
  • LangGraph
  • Langfuse
  • AWS
  • DynamoDB
  • Mathpix and Mistral OCR
  • Plotly
  • Prompt engineering
  • LLM as a judge
  • Human-in-the-loop review

Have something like this in mind?

Start a conversation→