InsightsArchitecture decisions

RAG vs fine-tuning: choosing for company knowledge

Decide whether your application needs fresher evidence, more consistent behavior, or both before investing in retrieval or model fine-tuning.

RAG and fine-tuning solve different parts of an AI application's problem. Retrieval-augmented generation supplies relevant information at request time. Fine-tuning adapts model behavior through additional training. The useful starting question is which failure you need to fix.

If an assistant cannot find the current policy, investigate sources and retrieval. If it receives suitable information but repeatedly fails a task-specific behavior, examine prompts, examples, validation, and whether fine-tuning is justified. Sometimes the application needs both; sometimes a simpler implementation is enough.

Begin with failure evidence

Collect representative requests and inspect the result. Was the necessary fact available? Did retrieval find it? Did the model use it correctly? Was the requested output valid for the next application step?

These are distinct questions. A wrong answer caused by an obsolete document is not the same problem as a failure to follow a stable output convention. Record the observed cause rather than choosing an architecture from a general claim that one approach is better.

Compare the application requirements

Scroll sideways to view the full table.

RequirementFirst area to investigate
Information changes frequentlySource updates and request-time retrieval
Answers need inspectable citationsPassage provenance and grounded generation
Different users have different document accessAuthorization in retrieval and source delivery
A stable task needs more consistent behaviorPrompt examples, validation, and possible fine-tuning
Relevant context is found but used poorlyContext selection and model behavior evaluation
Both facts and behavior varySeparate evaluations before combining approaches

This table is a diagnostic starting point. It does not establish that any approach will meet your acceptance criteria without testing.

A fictional policy-assistant decision

Imagine a company whose travel policy changes between versions. An employee asks which expenses require approval. The assistant needs the policy that applies to that employee and the relevant effective date.

A retrieval design can select the permitted source and return the supporting passage with the answer. Updating an approved document does not require teaching every policy change through a new training run. Fine-tuning would not remove the application's obligation to select the current, authorized information.

Now imagine a separate stable task: converting an approved expense description into an internal categorization format. The team can compare prompt examples and structured validation first. If recurring behavior failures remain and suitable training examples exist, fine-tuning can be evaluated against a held-out baseline.

The examples are illustrative. Neither establishes a clinical, financial, or organizational decision policy for a real deployment.

Keep knowledge ownership outside the model

Narravo records company knowledge with sources, relationships, and review states. A model-generated suggestion does not automatically become an approved fact. That product decision remains useful regardless of which model serves the response.

Fine-tuning changes model weights; it does not give an application document-level permissions or a source-review workflow. Microsoft's overview explicitly distinguishes fine-tuning from retrieval for current information and application-level controls.

Compare the whole operating system

For retrieval, account for ingestion, indexing, source updates, permission checks, and answer evaluation. For fine-tuning, account for training examples, evaluation data, supported models, training work, and serving arrangements. Provider availability and costs should be checked when making the implementation decision.

Do not compare only a per-request model bill. An approach that is cheaper to call may create additional work maintaining data or investigating unsupported answers.

Make the decision experimentally

Create a baseline that uses the simplest suitable approach. Evaluate factual support, source access, output validity, and task completion separately. Change one part at a time so you can explain what improved and what became more expensive or difficult to maintain.

Read production RAG for the source and evaluation foundations. For a project-specific architecture decision, explore RAG consulting.

Prepared with AI assistance using Paul’s documented project work and the linked sources. Examples are illustrative unless identified as project records.

Sources & further reading

Your next useful system

What could work
better?

Bring the business problem.
We’ll figure out the right next move.

Discuss a project