ai · llm · architecture
RAG vs long context vs fine-tuning: pick one on purpose.
"Should we fine-tune it?" is the "should we rewrite it in Rust?" of AI projects. Sometimes yes. Usually: not yet. Here is the table I wish people printed.
| Approach | Best for | Watch out for |
|---|---|---|
| Long context paste it all in | Small, stable sets of documents; one-off analysis; prototypes | Cost and latency per call; attention gets fuzzy on huge inputs |
| RAG retrieve, then answer | Large or changing knowledge; private data; answers that need citations | Retrieval quality is your quality. Bad chunks in, bad answers out |
| Fine-tuning change the model | Consistent style, format or behaviour; narrow tasks at scale | Poor at injecting fresh facts; needs good data and evals; goes stale |
My default order
- Write a clear prompt with a few good examples.
- Add retrieval if the knowledge is big, private or fresh.
- Try long context if the whole thing fits and you only need it occasionally.
- Fine-tune last, and only when you have an evaluation that proves the others are not enough.
the unglamorous secret
Whichever you pick, build the evaluation first: 30 real questions with known-good answers beats any amount of vibes.
They also stack. A fine-tuned model reading retrieved documents is a perfectly respectable architecture. The mistake is choosing the shiniest tool before you have measured the problem.
© Ramswaroop Patelwritten as one plain .html file