Almost every RAG project I've seen follows the same emotional arc. Week one: someone wires up a vector database, drops in 50 PDFs, asks a question, gets a shockingly good answer, and forwards a screenshot to the CTO. Week six: the same system is confidently citing a policy document that was superseded in 2021, and nobody trusts it enough to put it in front of a customer.
The gap between those two weeks is what this article is about. Not the tutorial version of retrieval augmented generation — the version where 40,000 documents live in four different systems, half of them are scanned, and legal wants to know why the bot said what it said.
First, an honest question: do you actually need RAG?
I'd rather lose a project than build the wrong thing. So before architecture, three cases where a production RAG system is the wrong answer:
- Your corpus is small and stable. If everything the model needs fits comfortably in a modern context window — say, under 100–200 pages that rarely change — just put it in the prompt. Prompt caching makes this cheap. You've saved yourself an entire retrieval stack.
- Your users ask a narrow, repeating set of questions. Twenty well-written canned answers plus keyword search will outperform a mediocre RAG pipeline, cost nothing to run, and never hallucinate.
- What you actually need is structured querying. "How many invoices over ₹5 lakh are unpaid past 60 days?" is not a retrieval problem. That's SQL. Semantic search over invoice text will give you plausible-sounding garbage. Text-to-SQL with a validated schema is a different, better-suited pattern.
RAG earns its complexity when you have a large, changing, heterogeneous body of knowledge and questions you can't enumerate in advance. Support knowledge bases, engineering documentation, regulatory archives, contract repositories, decades of field-service reports. That's the sweet spot.
The architecture that survives production
A prototype has two moving parts: embed and retrieve. A production system has roughly seven. Each one exists because something broke without it.
1. An ingestion pipeline you can re-run
This is the single most underrated component. Your first chunking strategy will be wrong. Your first embedding model will be superseded in eight months. If re-indexing means a developer running a notebook by hand for six hours, you will simply never improve the system.
Build ingestion as an idempotent, versioned job from day one. Source document → parsed text → chunks → embeddings, with content hashes at every step so unchanged documents are skipped. We usually run this as a queue-backed Python worker service; for clients already on AWS or Azure it slots into their existing batch infrastructure.
2. Document parsing that respects structure
Text extraction is where most quality is silently lost. A table in a PDF flattened into a run-on line of numbers is worse than useless — it's actively misleading. Scanned drawings need OCR. Long technical manuals need heading hierarchy preserved so a chunk knows it belongs to "Section 7.3: Calibration Faults".
Budget real engineering time here. On document-heavy projects, parsing and normalisation is routinely 30–40% of the total effort. It is the least glamorous and highest-leverage work in the whole build.
3. Chunking with context, not just character counts
Splitting every 500 characters is a starting point, not a strategy. What consistently works better in our projects:
- Split on semantic boundaries — headings, clauses, list items — then merge small pieces up to a target size.
- Prepend lightweight context to each chunk: document title, section path, effective date.
- Keep a parent-child relationship so you retrieve on the small precise chunk but feed the model the larger surrounding block.
4. Hybrid retrieval, always
Pure vector search fails on exactly the queries enterprise users care about: part numbers, error codes, statute references, proper nouns. Embeddings smooth away the specificity you need. BM25 keyword search nails those and fails at paraphrase.
Run both, fuse the results, then rerank the top 30–50 candidates with a cross-encoder before passing maybe 8–12 to the LLM. In our experience reranking is the highest return-on-effort upgrade available — it usually does more for answer quality than switching to a more expensive generation model.
5. Metadata filtering and access control at query time
This is a security requirement, not a feature. If retrieval doesn't respect who the user is, your RAG system becomes a very efficient data leak: the salesperson asks a vague question and gets a chunk of the HR compensation review. Permissions must be applied inside the retrieval filter, not by asking the model politely to keep secrets.
The same mechanism handles freshness — filter by effective date, prefer current revisions, mark superseded documents.
6. Grounded generation with citations
Instruct the model to answer only from provided context and to say plainly when the context is insufficient. Return citations down to the chunk, with a link back to the source document and page. Users forgive a system that says "I don't have this". They do not forgive one that invents a warranty term.
7. Evaluation and logging
You cannot improve what you don't measure, and "the demo felt good" is not measurement. Build a test set of 100–300 real questions with acceptable answers, sourced from actual support tickets or user interviews. Track retrieval recall separately from answer quality — when something goes wrong you need to know whether the right chunk wasn't found or was found and ignored.
Log every query, retrieved chunks, and final answer. Six weeks of production logs will teach you more about your users than six months of planning.
The pitfalls that cost the most
Debugging the wrong layer. A bad answer is almost always retrieval, not generation. Teams burn weeks on prompt engineering when the correct chunk never entered the context. Check retrieval first, every time.
Stale content with no lifecycle. Documents get deleted, revised, expired. If your index has no deletion path, the system will keep quoting a 2019 pricing sheet forever. This is the most common quiet failure I see in inherited projects.
Ignoring conversational context. "What about for the 400 series?" is meaningless as a standalone search query. You need a query-rewriting step that resolves follow-ups against conversation history.
No fallback behaviour. Real systems need a graceful path when confidence is low: escalate to a human, offer document links instead of an answer, ask a clarifying question.
Treating it as a project, not a product. RAG quality is a maintenance commitment. Content drifts, users find new phrasings, models change.
Real costs, with real numbers
Approximate figures based on projects we've delivered — treat them as ranges, not quotes.
Build. A credible internal production RAG system over a defined corpus, including ingestion, hybrid retrieval, evaluation harness and a usable interface, is typically a 10–16 week engagement with 2–3 engineers. Customer-facing systems with strict compliance, multi-tenancy and SSO run longer.
Inference. With a mid-tier commercial model, a typical query at 4,000–8,000 context tokens costs somewhere between a fraction of a cent and a few cents. At 20,000 queries a month that's plausibly $100–600 — genuinely modest. Embeddings are cheaper still; a one-time index of a million chunks is often under $100.
Infrastructure. Managed vector databases start around $70–100/month at small scale and climb with vector count and replicas. Self-hosting pgvector or Qdrant on your own hardware is frequently the better call, especially in regulated environments.
The part people forget. Ongoing evaluation, content pipeline fixes, model upgrades and prompt tuning realistically run 15–25% of the original build cost annually. Skip it and quality decays quietly until people stop using the thing.
The pattern is consistent: model tokens are the small line item. Engineering and data preparation are the real cost, and that's where the returns are too.
Where this leaves you
A production RAG system isn't a clever prompt on top of a vector store. It's a data pipeline with a language model at the end of it — and it should be resourced like one. Get parsing and retrieval right and even a modest model will feel sharp. Get them wrong and no frontier model will save you.
We've been building Python and AI/ML systems at FoogleTech since 2012, for clients across the USA, UK, Europe and the Middle East. If you're weighing up a retrieval augmented generation project — or you have a prototype that impressed everyone in the demo and nobody since — we're happy to look at it and tell you honestly what it needs, including if the answer is "less than you think". Reach us at foogletech.com/contact-us.