ai · llm · architecture

RAG vs long context vs fine-tuning: pick one on purpose.

Ramswaroop
5 Sept 2026
aillmarchitecture

"Should we fine-tune it?" is the "should we rewrite it in Rust?" of AI projects. Sometimes yes. Usually: not yet. Here is the table I wish people printed.

ApproachBest forWatch out for
Long context
paste it all in
Small, stable sets of documents; one-off analysis; prototypesCost and latency per call; attention gets fuzzy on huge inputs
RAG
retrieve, then answer
Large or changing knowledge; private data; answers that need citationsRetrieval quality is your quality. Bad chunks in, bad answers out
Fine-tuning
change the model
Consistent style, format or behaviour; narrow tasks at scalePoor at injecting fresh facts; needs good data and evals; goes stale
Need new facts?Small + stable?Big or changing?Long contextRAG
and if the answer is "change how it behaves", that is when fine-tuning gets its turn

My default order

  1. Write a clear prompt with a few good examples.
  2. Add retrieval if the knowledge is big, private or fresh.
  3. Try long context if the whole thing fits and you only need it occasionally.
  4. Fine-tune last, and only when you have an evaluation that proves the others are not enough.
the unglamorous secret

Whichever you pick, build the evaluation first: 30 real questions with known-good answers beats any amount of vibes.

They also stack. A fine-tuned model reading retrieved documents is a perfectly respectable architecture. The mistake is choosing the shiniest tool before you have measured the problem.


© Ramswaroop Patelwritten as one plain .html file