Start with a measured prompt baseline. Use retrieval for accessible, current evidence; consider fine-tuning for recurring task behavior that examples can teach. Combine them only when each solves an independently observed failure.
RAG changes context; fine-tuning changes parameters
A typical RAG application searches a document collection and sends selected passages to the model alongside the question. Google Cloud describes ingestion, transformation, indexing, retrieval, and generation as distinct stages. Updating a policy can therefore update the available evidence without requiring a new model-training run, once ingestion and indexing have caught up.
Supervised fine-tuning uses examples of inputs and desired outputs to adapt model parameters. It can help with recurring tasks and output conventions. That is different from making a document searchable. The original RAG paper also illustrates that parametric and retrieved knowledge can coexist; RAG and fine-tuning are not mutually exclusive categories of product.
Diagnose the mistake before buying a solution
Take a failed answer and supply the correct, complete evidence manually. If the answer becomes correct, the first suspect is the information path: missing documents, poor retrieval, stale indexing, or incomplete context. If the model still confuses the rules, investigate the prompt, evidence layout, task difficulty, and examples. A failure after retrieval is not automatically a reason to fine-tune.
Use this as a diagnostic experiment, not a guarantee. One carefully chosen passage may be easier to read than real production results. Repeat the experiment on a representative set, including ambiguous questions. Separate cannot find the rule from found the rule but used it incorrectly; otherwise an apparent model improvement may simply reflect a better search result.
Worked comparison: a subscription support assistant
Consider an original example: a software company has region-specific refund rules and wants a support assistant to return an answer plus an internal case category. A customer asks whether yesterday's annual renewal is refundable. The relevant policy depends on region, purchase channel, and effective date. These details must be obtained from authorized records or clarified, rather than guessed from the language of the message.
A prompt-only baseline may explain a familiar refund policy fluently while applying the wrong region. A RAG version retrieves the applicable policy revision and cites it. A fine-tuned model trained on last quarter's conversations may learn the company's concise tone and case categories, yet those examples do not establish today's policy. We would choose retrieval for the policy-evidence problem in this scenario.
Suppose retrieval now supplies the correct rule, but the assistant repeatedly labels renewal cases as new purchases. First test clearer category definitions and examples in the prompt. If that behavior remains a measurable problem and you have reviewed training examples, a fine-tuning experiment may be justified. A combined version can retrieve policy evidence and use an adapted model for the recurring classification task. Compare each change separately.
RAG quality depends on evidence selection
For this assistant, store policy identity, revision date, effective period, region, and purchase channel with each searchable passage. Preserve headings that connect an exception to its rule. Search results containing a refund deadline but omitting its renewal exception can produce a plausible wrong answer. More retrieved text is not a reliable substitute for the right evidence.
Enforce document access before content enters the model's context. When a policy changes, verify the new revision can be found and the superseded revision is handled correctly. Add an explicit outcome for insufficient or contradictory evidence: ask a focused question or escalate. A citation helps the reader inspect a claim, but the presence of a link does not establish that the linked passage supports it.
Fine-tuning quality depends on examples
Google's tuning guidance recommends establishing a prompt baseline, examining errors, and using high-quality examples that resemble production inputs. In our scenario, a useful training example includes the question, applicable evidence, and correct case category. A transcript with an agent's initial mistake left uncorrected teaches the wrong target.
Separate training and evaluation data by meaningful groups, such as customer case or policy scenario. Randomly splitting near-duplicate messages from the same incident can exaggerate apparent generalization. Keep changing policy facts in the supplied context when that is how production works. Evaluate whether the model follows new evidence when it conflicts with older training examples; do not assume adaptation has taught it a dependable update mechanism.
Evaluate retrieval and the final answer separately
Microsoft's RAG evaluation documentation separates retrieval quality from answer properties such as groundedness, relevance, and completeness. For the subscription assistant, label which policy passages are needed, whether the final decision is correct, and whether the internal category is correct. This reveals an improvement in classification that might otherwise hide a regression in policy accuracy.
Compare prompt-only, prompt plus retrieval, and any tuned variants on the same held-out questions. Include missing region, unavailable policy, conflicting revisions, and a newly changed renewal rule. Measure justified abstention as well as successful answers. Record end-to-end latency and cost with failed attempts included. Review critical policy mistakes manually and calibrate model judges against those reviews. Treat illustrative scores as hypotheses until you have run the experiment.
Compare operating costs and maintenance
RAG adds document ingestion, search, and context tokens to an application. Fine-tuning adds dataset preparation, training, evaluation, and model-version management. Neither is universally cheaper. A shorter tuned prompt might reduce generation costs while a maintained retrieval layer still remains necessary. Price the full request path at your expected traffic and refresh rate.
Both approaches can generate unsupported answers. Retrieval can return irrelevant or outdated passages; an adapted model can learn undesirable patterns or perform worse on unrepresented cases. For changing customer-specific facts such as whether a refund has already been issued, use an authorized system lookup. A document index and training examples are both unsuitable substitutes for the live transaction record.
A decision checklist you can actually test
Collect twenty realistic support questions and write their expected evidence and outcomes. This small set is a starting diagnostic exercise, not sufficient production validation. Run the prompt baseline, then add the correct evidence manually. Classify the remaining mistakes before changing the architecture. In an interview, explain the failure, proposed fix, and measurement rather than declaring one technique better in every case.
- Missing or changing policy evidence: test retrieval, freshness, and document access.
- Wrong answer despite correct evidence: test instructions and evidence structure before training.
- Repeated category or format errors: compare prompt examples with a fine-tuning experiment.
- Customer-specific live state: query the authorized operational service.
- Two independent problems: test a combined design and verify each component's contribution.
Quick answers
Frequently asked questions
Can fine-tuning replace a knowledge base?
It can encode patterns and some information, but it is not a dependable replacement for searchable, current, permission-scoped evidence. Test how changing facts are supplied and verified in your application.
Does RAG prevent hallucinations?
No. Relevant evidence can help, but retrieval and generation can both fail. Check whether the answer is supported, handles missing evidence, and uses the applicable document version.
Can I use RAG with a fine-tuned model?
Yes. A retrieved policy can supply facts while an adapted model performs a learned classification or output task. Compare the combination with simpler baselines to establish whether both components add value.
Should I start with RAG or fine-tuning?
Start with an evaluated prompt baseline. Add retrieval when errors come from missing evidence; test fine-tuning when a recurring behavior problem remains and representative, reviewed examples are available.
Source notes
References and review policy
Information checked on October 2, 2026. Section links identify sources for factual claims and technical explanations. Interpretations, practice scenarios and preparation recommendations are RecallDeck’s editorial work.
From reading to recall
Practice the full interview loop.
RecallDeck schedules the concepts you miss and keeps coding, design, and behavioral fundamentals available when the interviewer changes direction.