conceptEntrenamiento y Adaptación~1 min de lecturaActualizado 2026-06-07#fine-tuning#rag#prompting#decision
Esta nota todavía no está traducida, así que se muestra la fuente en inglés.

When to fine-tune vs prompt vs RAG

Fine-tuning is expensive because it changes weights, deployment shape, and evaluation surface. Before doing it, prove that cheaper adaptation — prompt or retrieval — cannot produce the behavior you need.

The adaptation ladder

Need First tool Why
Better instructions, tone, or output shape Prompting Behavior can be specified at runtime.
Fresh or private knowledge RAG Knowledge stays outside the weights and can update.
Stable style or task skill across many calls SFT The behavior becomes default.
Smaller, cheaper specialized model Distillation Transfer a large model's behavior into a smaller one.

Fine-tuning should be a response to repeated evidence, not optimism. If a prompt change fixes the failure, you do not need a fine-tune yet.

Good reasons to fine-tune

  • You have many examples of the exact behavior you want.
  • The desired format must be reliable across thousands of calls.
  • The base model understands the domain but fails the style, policy, or procedure.
  • You need a smaller model to imitate a stronger one for cost or latency.
  • You can evaluate the fine-tuned behavior with held-out examples.

Bad reasons to fine-tune

  • "The model does not know our latest docs" → use retrieval.
  • "The prompt is messy" → fix context engineering first.
  • "We want fewer hallucinations" → evaluate grounding and evidence use first.
  • "We have a pile of raw logs" → raw logs are not a training dataset.

Pitfall

Fine-tuning for facts bakes stale knowledge into the model and makes updates slow. Use retrieval for knowledge, and fine-tune for behavior.

Connects to: context engineering · why RAG · fine-tune evaluation