Building a Production RAG System Across 3,000+ Pages of Mortgage Guidance
Building a Production RAG System Across 3,000+ Pages of Mortgage Guidance
Retrieval-augmented generation can look straightforward in a prototype: ingest documents, create embeddings, retrieve relevant passages and send them to a language model. Production systems are different. Retrieval quality, document structure, permissions, traceability, latency and evaluation all become engineering problems.
While building Vellum, a mortgage underwriting platform, our team worked with more than 3,000 pages of mortgage guidance and documentation. The experience reinforced a simple lesson: production RAG is primarily a retrieval and systems-engineering problem, not just an LLM integration.
Why mortgage guidance is difficult for RAG
Mortgage guidance contains dense rules, exceptions, eligibility conditions and overlapping guidance across different loan programmes. A useful system cannot simply find a paragraph containing similar words. It needs to retrieve the right evidence for the specific underwriting question and preserve enough context for the answer to be checked.
1. Ingestion is more than uploading PDFs
Documents need to be normalised and prepared so their structure survives ingestion. Headings, sections, tables and surrounding context can materially affect retrieval. Poor document preparation creates problems that no prompt can reliably repair later.
2. Chunking changes retrieval quality
Chunks that are too small lose the conditions surrounding a rule. Chunks that are too large introduce unrelated information and weaken retrieval precision. The right strategy depends on document structure and the questions users actually ask, so chunking needs to be tested against representative queries rather than selected by a fixed token count.
3. Vector search is only part of retrieval
For Vellum, Pinecone forms part of the retrieval architecture. But a vector database alone does not make a RAG system accurate. Metadata, filtering, query construction, ranking and the amount of context passed downstream all influence whether the model receives the evidence it needs.
4. Answers need evidence
In a high-stakes workflow, a fluent answer is not enough. Users need to understand where an answer came from. Production RAG should preserve source information and make the retrieved evidence available for review rather than asking users to trust model output blindly.
5. Evaluation has to test retrieval and generation separately
When an answer is wrong, the failure may have happened before the model generated anything. The system may have retrieved the wrong passage, failed to retrieve the correct one, supplied too much context or interpreted good evidence incorrectly. Evaluating retrieval and generation separately makes these failures easier to diagnose. We cover the measurement framework in more detail in our guide to RAG evaluation, retrieval quality and groundedness.
6. Production architecture matters
Real RAG products also need authentication, document storage, background processing, queues, observability and resilient infrastructure. Our Vellum architecture combines a NestJS backend and Next.js application with Pinecone, Redis, RabbitMQ, S3 and Auth0 alongside the AI layer. These components matter because users experience the complete system, not the retrieval demo.
7. Cost and latency become product decisions
More retrieved context is not automatically better. Larger prompts increase model cost and can increase latency while introducing irrelevant material. Production optimisation means finding the smallest amount of high-quality evidence needed to answer reliably.
What separates production RAG from a prototype?
A prototype proves that retrieval can work. A production system must prove that it can work repeatedly, securely and transparently across real documents and real user questions. That means treating ingestion, retrieval, evaluation, permissions, infrastructure and monitoring as first-class parts of the product.
For organisations exploring this architecture, see our RAG development services, our broader AI development capabilities, and our UK AI development cost guide.
Production RAG FAQs
What is a production RAG system?
A production RAG system connects an AI model to trusted internal knowledge and retrieves relevant evidence at query time, with the security, evaluation and infrastructure needed for real users.
How is production RAG different from a prototype?
A prototype proves retrieval can work. Production RAG also needs reliable ingestion, permissions, source traceability, monitoring, latency control and repeatable evaluation.
What affects RAG accuracy most?
Document quality, chunking, metadata, retrieval strategy, ranking and evaluation usually matter more than simply choosing a larger language model.
Can RAG be used in regulated workflows?
Yes, but the system should preserve evidence, access controls and auditability so users can verify where answers came from and review important decisions.
Planning a production RAG system?
If you are working with complex documents, internal knowledge or regulated workflows, WeUno can help you assess the retrieval architecture, evaluation approach and engineering required to take the system into production.