Fine-tuning has a reputation for being the "advanced" option, which leads teams to reach for it before they've exhausted what a well-engineered prompt can do.
In practice, we only recommend fine-tuning once a client has a stable, well-tested prompt and is still hitting an accuracy or consistency ceiling, or when per-request token costs from a long system prompt start to dominate the unit economics at scale.
The deciding factor is usually volume: below a few hundred thousand requests a month, a good prompt plus retrieval usually wins on total cost once you include training and evaluation time. Above that, a fine-tuned, smaller model often wins on both cost and latency.

