RAG in Production: Retrieval Augmented Generation Done Properly

RAG that works for real users: chunking, hybrid search, reranking, permission-aware retrieval, freshness, evaluation, and when to skip RAG entirely.

S
Softzee EngineeringSeptember 29, 2026 · 6 min read

A RAG demo takes an afternoon. A RAG system that gives correct, current answers to real users, without leaking documents they should not see, takes considerably more work. Most teams discover the gap after launch, when the assistant confidently quotes an outdated policy or answers from a file the user was never allowed to open.

Retrieval augmented generation is simple in principle: find the relevant text, put it in the prompt, let the model answer from it. In practice, nearly every quality problem in a RAG system is a retrieval problem, not a model problem. This guide covers the parts that decide whether RAG works in production.

Chunking: the decision that quietly sets your ceiling

If the right answer is split across two chunks, or buried in a chunk full of unrelated text, no model will fix it. Chunking is where retrieval quality is won or lost, and the default "split every 500 tokens" approach is rarely the best choice.

Better options, roughly in order of effort:

  • Structure-aware chunking. Split on headings, sections, list items and table boundaries instead of character counts. A policy document chunked by section keeps each rule intact.
  • Overlap. A modest overlap between adjacent chunks prevents sentences from being cut in half at the boundary.
  • Context headers. Prepend the document title and section path to every chunk ("Employee Handbook > Leave > Annual leave"). A chunk that says "this applies to all full-time staff" is useless without knowing what "this" refers to.
  • Parent-child retrieval. Index small chunks for precise matching, but pass the larger parent section to the model so it has enough surrounding context.

Tables, scanned PDFs and slide decks need specific handling. A table flattened into plain text loses its meaning. Convert tables into row-level statements or keep them as Markdown, and run OCR quality checks on scanned material before it ever reaches the index. For Arabic and mixed Arabic-English content, which is common in Gulf deployments, check that your text extraction preserves word order and that your embedding model actually handles Arabic well.

Hybrid search and reranking

Pure vector search is good at meaning and bad at exact terms. It will happily match "refund policy" to "return policy" but can miss a query for invoice number INV-20931 or a specific product code. Keyword search (BM25) has the opposite strengths. Production systems should use both.

Hybrid retrieval

Run a vector query and a keyword query in parallel, then merge the results. Reciprocal rank fusion is a simple, effective way to combine the two lists without tuning score scales. Most modern search engines and vector databases support hybrid queries natively, so this is usually configuration rather than custom code.

Query rewriting

Users do not type search queries. They type "and what about for part-time staff?" in the middle of a conversation. Before retrieval, have a small, cheap model rewrite the latest message into a standalone query using the conversation so far, and for broad questions, generate two or three variants and retrieve for each. This step is inexpensive and often fixes a surprising share of follow-up questions that otherwise retrieve nothing useful. Log the rewritten query next to the original so you can debug retrieval misses later.

Reranking

Retrieve a wide set of candidates (say 30 to 50) and then use a cross-encoder reranker to score each candidate against the query and keep the top few. Rerankers read the query and passage together, so they judge relevance far more accurately than embedding similarity alone. The trade-off is added latency, which is why you rerank a shortlist instead of the whole corpus.

candidates = hybrid_search(query, k=40, filters=user_acl_filter)
ranked     = reranker.score(query, candidates)
context    = [c for c in ranked[:5] if c.score > MIN_RELEVANCE]
if not context:
    return "I could not find this in the knowledge base."

Note the last two lines. A relevance threshold lets the system say "I don't know" instead of forcing an answer from weak matches, which is one of the most effective ways to reduce hallucination in RAG.

Permissions-aware retrieval

This is where internal RAG projects most often fail a security review. If your index contains HR files, contracts and board papers, and your retrieval does not enforce who can see what, the assistant becomes a search engine for confidential data.

The rule is simple: filter at retrieval time, not after generation. Telling the model "do not reveal confidential information" is not access control. Instead:

  1. Store access metadata (owner, groups, tenant, classification) on every chunk when you index it.
  2. Resolve the current user's identity and group memberships on each request.
  3. Apply those as hard filters in the search query, so unauthorized chunks are never retrieved.
  4. Sync permission changes from the source system (SharePoint, Google Drive, your document store) on a schedule short enough that revoked access actually takes effect.

For multi-tenant SaaS products, use a tenant filter on every query at minimum, and consider separate indexes or namespaces per tenant when contracts or regulation require stronger isolation. In Saudi Arabia, the PDPL has been fully enforced since September 2024, so retrieval that exposes personal data to the wrong user is a compliance issue as well as an embarrassment.

Freshness: keeping the index honest

An index built once and never updated is a liability. Policies change, prices change, documents get deleted. The model will keep quoting whatever is in the index.

  • Use incremental ingestion driven by change events or modified timestamps rather than full rebuilds.
  • Handle deletions explicitly. Removing a document from the source must remove its chunks from the index.
  • Store a version or last-updated date on each chunk, and give the model that date so it can say "according to the policy updated in March".
  • For fast-changing facts such as stock levels, order status or account balances, do not use RAG at all. Call the live system through a tool.

Monitor ingestion like any other pipeline: documents processed, failures, lag between a source change and the index reflecting it. Silent ingestion failures are a common cause of "the bot is giving old answers" tickets.

Evaluation: how you know it works

Without evaluation, every change to chunk size, embedding model or prompt is a guess. Build a test set early, ideally 100 or more real questions collected from users or support logs, each with the documents that should be retrieved and a reference answer.

Measure retrieval and generation separately:

LayerWhat to measureWhy it matters
RetrievalRecall at k: did the right chunk appear in the top results?If retrieval misses, the answer cannot be right
RetrievalPrecision: how much of the context is relevant?Irrelevant context wastes tokens and confuses the model
GenerationFaithfulness: is every claim supported by the context?Catches hallucination on top of good retrieval
GenerationAnswer correctness against the referenceThe user-facing quality number
SystemCorrect refusals on out-of-scope questionsSaying "I don't know" is a feature

LLM-as-judge scoring is useful for faithfulness and correctness at scale, but calibrate it against human ratings on a sample before you trust it. Run the suite on every change to the pipeline, and add every bad answer reported in production to the test set.

When not to use RAG

RAG is often the default answer to "make the AI know our stuff", and it is frequently the wrong one.

  • The data is structured. Questions like "How many orders shipped late last week?" need a query against your database, not semantic search over exported text. Give the model a tool that runs a safe, parameterized query.
  • The data is live. Balances, availability and appointment slots should come from the system of record through an API call.
  • The knowledge base is small. If everything fits comfortably in the context window, and prompt caching makes repeating it cheap, putting the whole thing in the prompt may be simpler and more accurate than retrieval.
  • You need behavior, not facts. Tone, format and domain style are better handled by prompting, examples or fine-tuning than by retrieval.

Many good assistants combine approaches: RAG for policies and documentation, tools for live data, and a well-written system prompt for behavior. Connecting those sources cleanly is usually more of an AI integration problem than a model problem.

How Softzee can help

We build RAG-backed assistants for web, WhatsApp and internal teams, with permission filtering, evaluation and freshness built in from the first release. If you are planning an AI assistant on top of your own documents or fixing one that already gives wrong answers, book a call and we will review your setup.

RAGRetrieval augmented generationVector searchLLM evaluationAI engineering

Have a project in mind?

Tell us what you are trying to build. You will get an honest take on scope, timeline and cost, usually within one business day.

Keep reading

All articles