Fine-Tuning vs RAG vs Prompt Engineering: How to Choose

A practical decision framework for LLM customization — from someone who has watched teams spend six figures solving the wrong problem.

Verified Top Talent
FoogleTech Engineering Team
By

FoogleTech Software •

EXPERTISE
Fine-Tuning vs RAG vs Prompt Engineering: How to Choose
Article Contents

A client came to us last year with a clear ask: "We need to fine-tune a model on our internal documentation." They had a budget approved, a GPU quote from a cloud vendor, and a six-week timeline.

We spent two days looking at their actual problem. It turned out their support team needed accurate answers from a knowledge base that changed every week. Fine-tuning would have baked last week's answers permanently into model weights — and they would have had to retrain constantly to stay current. We shipped a retrieval system instead. It took eleven days and cost roughly a tenth of the original plan.

That conversation happens more often than you would think. The debate around fine-tuning vs RAG has become weirdly tribal online, when in practice it is a fairly boring engineering decision with clear signals. This post is my attempt to lay out those signals honestly, including the cases where you should skip all three approaches and do nothing at all.

The three levers, described plainly

Before comparing them, it helps to be precise about what each technique actually changes.

Prompt engineering changes the instructions you send at request time. The model's weights are untouched. You are shaping behaviour through context: system prompts, few-shot examples, output schemas, chain-of-thought scaffolding, tool definitions. This includes the more sophisticated end of the spectrum — prompt chaining, self-critique loops, structured output enforcement — which people sometimes dismiss as "just prompting." It isn't just anything. A well-engineered prompt pipeline is real software.

RAG (Retrieval-Augmented Generation) changes what information the model can see. You keep a searchable store of your data — vector database, keyword index, SQL, or usually some combination — retrieve the relevant slices at query time, and inject them into the prompt. The model doesn't "know" your data; it reads it fresh on every request.

Fine-tuning changes the model itself. You take a base model and continue training it on your examples so the weights shift. Full fine-tuning updates everything; parameter-efficient methods like LoRA update a small adapter layer, which is what most teams actually use in production because it is dramatically cheaper and easier to version.

The critical distinction, and the one that resolves most arguments: RAG gives a model knowledge. Fine-tuning gives it behaviour. Prompt engineering gives it instructions. Those are three different problems.

Start with the question you're actually solving

Here is the diagnostic I run with clients. Ask which of these best describes your failure mode:

  • "The model doesn't know things it needs to know." — Facts about your products, your customers, your policies, this month's pricing. This is a knowledge problem. RAG.
  • "The model knows enough but responds in the wrong style, format or structure." — Wrong tone, ignores your taxonomy, won't reliably produce the JSON shape your downstream service expects. This is a behaviour problem. Start with prompting; fine-tune if prompting plateaus.
  • "The model is right but too slow or too expensive at our volume." — This is an economics problem. Fine-tuning a smaller model to imitate a larger one often works beautifully here.
  • "The model is wrong in ways I can't quite characterise." — This is an evaluation problem, and you cannot fix it with any of the three until you can measure it. More on this below.

Nine times out of ten, naming the failure mode precisely tells you the answer. The trouble is that most teams skip that step and go straight to picking a technology.

When RAG is the right call

RAG wins whenever your ground truth changes faster than you can retrain, or whenever you need to show your work.

Go with RAG if:

  • Your source data updates weekly, daily, or continuously.
  • You need citations — the user must be able to click through to the source document. This is non-negotiable in legal, medical, finance and compliance contexts.
  • You need access control. Different users should see answers drawn from different documents. You can enforce that at retrieval time; you cannot enforce it inside model weights.
  • Your corpus is large. Millions of documents cannot be memorised usefully, but they can be indexed.
  • You need auditability. When something goes wrong, you want to know which chunk caused it.

What people underestimate about RAG is that the hard part is not the LLM — it is the retrieval. Chunking strategy, hybrid search, re-ranking, query rewriting, handling documents with tables and diagrams, dealing with near-duplicate content across document versions. In our projects, retrieval quality accounts for the large majority of the engineering effort and almost all of the accuracy gains. If your RAG system is underperforming, the model is rarely the culprit.

When fine-tuning genuinely earns its cost

I am not anti-fine-tuning. It is the right tool in specific, identifiable situations:

  • Consistent structured output at scale. If you need a hundred thousand documents parsed into an identical schema, a fine-tuned smaller model will beat a prompted large model on both cost and consistency.
  • Domain language and reasoning patterns. Medical coding, semiconductor test logs, legal drafting conventions, industrial sensor fault classification. Where the style of reasoning is specialised, examples teach better than instructions.
  • Latency and cost pressure. Distilling a frontier model's behaviour into a small open-weight model you host yourself can cut per-request cost by an order of magnitude. At high volume this pays for itself quickly.
  • Data residency or air-gapped deployment. Several of our embedded and industrial clients cannot send data to an external API at all. A fine-tuned open model running on-premise is the only option.
  • Prompting has plateaued. You have iterated seriously, you have a real eval set, and you are stuck at 82% when you need 94%. This is the legitimate signal.

The honest cost picture: the compute for a LoRA fine-tune is often trivial — tens to a few hundred dollars. The expensive part is the dataset. Curating, labelling and cleaning a few thousand high-quality examples is weeks of skilled human work, and it is where fine-tuning projects die. Then you own a model artefact forever: versioning, regression testing, retraining when the base model deprecates. Budget for the lifecycle, not the training run.

When you need none of this

This is the part most agencies won't tell you.

If you are still validating whether users want the feature, do not build a RAG pipeline. Put your twenty most important documents into a long context window with a carefully written system prompt and ship it behind a feature flag. Modern context windows are enormous. A surprising number of "we need RAG" problems are actually "we need to paste the handbook in" problems.

And if your task is well-represented in public data — summarising, translating, classifying common categories, general drafting — a good prompt against a strong general model may already be at or near your quality ceiling. Spending three months on LLM customization to gain two percentage points you cannot even measure is not engineering. It is theatre.

The hybrid reality of production systems

In practice, the fine-tuning vs RAG framing is a false binary. Most serious systems we build end up using all three:

  1. Prompt engineering for orchestration — routing, tool selection, output contracts, guardrails.
  2. RAG for anything factual, current or access-controlled.
  3. A small fine-tuned model for one or two high-volume, narrow sub-tasks — query rewriting, document classification, extraction — where a specialist beats a generalist on cost and consistency.

That layering is the norm, not the exception. The question is never "which one" but "which one for which part of the pipeline."

Build your evaluation set before you build anything else

If you take one thing from this post, take this. Before choosing an approach, write 100 to 300 real test cases with expected outputs, drawn from actual user behaviour rather than your imagination. Score them. Get a baseline number with plain prompting.

Without that, you cannot tell whether fine-tuning helped, whether your retrieval changes improved anything, or whether last week's prompt tweak silently broke a category of queries. Teams without eval sets ship on vibes and then discover problems in production. Teams with eval sets make decisions in an afternoon that would otherwise take a month of arguing.

We generally spend the first week of an LLM engagement building evals, not features. Clients occasionally push back. None of them have regretted it.

A rough decision path

  1. Define the failure mode in one sentence.
  2. Build an eval set and get a prompt-only baseline.
  3. Knowledge gaps? Add retrieval. Iterate on retrieval quality before touching the model.
  4. Behaviour or format problems? Push prompting hard first — few-shot examples, structured output, decomposition.
  5. Still short of target, or facing real cost/latency/residency constraints? Now fine-tune, on data you have curated deliberately.
  6. Re-measure after every change. Keep the eval set growing with real production failures.

Where we come in

At FoogleTech we have been building software since 2012, and Python and AI/ML work now sits at the centre of what we do — RAG systems over messy enterprise document estates, fine-tuned models running on-premise for clients with strict data rules, and LLM features embedded into products that also have to talk to hardware and IoT fleets. We also run our own SaaS products, which means we live with our own architectural decisions rather than handing them over and walking away.

If you are weighing fine-tuning vs RAG for a real product and want a straight answer rather than a proposal for the most expensive option, get in touch at foogletech.com/contact-us. Tell us the failure mode you are seeing. Sometimes the most useful thing we do in that first conversation is talk a client out of a build — and if that is the honest answer for you, we will say so.