Immigration Assistant
A RAG chatbot over official government information, answering the questions people actually have when moving to a new country.
Client work from a former company. The organisation and the product are kept deliberately general, and nothing shown here reproduces a real interface. Happy to talk through the engineering in person.
- Client
- A public-sector information service
- Role
- AI Engineer
- Timeline
- Client engagement
- Discipline
- Applied AI · Public Information
Overview
The information someone needs in order to move countries already exists. It is published, it is public, and it is hard to use. It sits spread across ministries and agencies, written in bureaucratic register, structured for the institution rather than for the person. This assistant sat over a large corpus of official governmental material and answered in conversation instead of in forms: what applies to me, in my situation, right now.
Challenge & approach
The challenge
Someone relocating is usually working in a second language, against a deadline, with their legal status depending on getting the answer right. The official material is complete and accurate, and it is organised by which agency owns it rather than by what a person is trying to do. So the challenge is not missing information. It is helping a person find the paragraph that applies to them, and be sure it is the current one.
The approach
Retrieval quality was the whole product, so that is where the engineering went. Answers were grounded in the source documents and carried their provenance back to the reader, so a claim could always be checked against the text it came from. Just as much work went into the boundaries: knowing when a question fell outside the corpus, saying so rather than improvising, and handing off to a human instead of guessing. Getting this right matters, because a fluent wrong answer could mean a missed deadline or the wrong application. The whole system ran serverless, which kept it remarkably cheap to host and easy to keep running.
How it was built
We built the proof of concept as a small team, deliberately mixed: another data scientist and me on the system, a designer on the conversation and the interface, and two project managers keeping a tight timeline honest. The pace was fast from the first week, and it stayed that way. The interface was written in TypeScript and React, the retrieval and orchestration with LangChain over a LanceDB vector store, and everything ran on AWS Lambda, so there was no server to mind and the running costs stayed close to nothing. The testing was the team's finest hour, and I say that as the one who watched it from the engineering side: my colleagues took the assistant onto the streets and asked people to use it there, with their own questions, thinking aloud while they listened. That feedback shaped the answers, the tone and the interface directly. The response was warm, and the work drew real attention from government bodies and aid organisations, who saw what it could do for the people they serve. After the proof of concept I was pulled onto another client project, the Global Response Documents. The team carried the assistant forward, built it out and founded a start-up around it. Some time later they stopped: once they ran the numbers on what a large language model costs per question at the volume such a service would attract, together with the complexity of the material, the economics did not hold. I think that is worth saying on a portfolio. An AI product lives or dies on the cost of each answer, so it pays to do that maths early, ideally before the first line of code, and again as soon as real people are using it. And the maths keeps changing: prices per token keep falling, and capable open models that run on hardware of your own keep arriving. An idea that does not add up today may add up soon, and the work of getting the answers right does not go to waste.
How it worked
- 01
Ingest
Official material collected from across agencies and normalised into a single searchable corpus.
- 02
Retrieve
Vector search surfaces the passages that actually bear on the question being asked.
- 03
Ground
The answer is composed only from retrieved passages, with each claim tied to its source.
- 04
Check
Low confidence or out of scope questions are handed to a human instead of guessed.
In principle
Official corpus
Question
Grounded answer
What it does
- 01Retrieval over a large corpus of official governmental information
- 02Answers grounded in source documents, with provenance surfaced to the reader
- 03Conversational access to material otherwise buried in institutional structure
- 04Honest uncertainty: the assistant says when it does not know instead of guessing
- 05A firm line between explaining official information and giving legal advice
- 06Evaluation harness to keep retrieval quality high, release after release
- 07Serverless from end to end, so hosting stayed close to free
Guiding principles
- People's legal status can depend on the answer, so an honest 'I do not know' beats a guess.
- Explaining official information is not the same as giving legal advice, and the line had to hold.
- Source material changes, so answers had to reflect the current text rather than a stale snapshot.
- Readers are often not working in their first language.
- A tight timeline, so every round had to teach something and ship.
Built with
- TypeScript
- React
- LangChain
- LanceDB
- AWS Lambda
- Retrieval-augmented generation
- Evaluation harnesses