RAG (retrieval-augmented generation) lets a language model read the right passages from your business documents before it answers; fine-tuning trains the model further on your data to change how it responds. For most tasks in small and mid-sized businesses (answering customers from policies, looking up procedures, summarizing files), RAG is the right choice because it is cheaper, faster, updated by editing documents, and always able to cite a source. Fine-tuning is only worth it when the model must speak in a very specific voice or format, or handle a specialized kind of data that instructions cannot cover.
The shared problem: the model knows nothing about your business
Large language models are trained on public text. They do not know your store's return policy, this month's price list, or your internal warranty procedure. Ask about those and they answer vaguely or fabricate confidently.
There are two ways to bring business knowledge in, and they are fundamentally different.
RAG: read the documents, then answer
How it works, simply:
- Business documents (policies, price lists, guides, records) are split into passages and indexed.
- When a question arrives, the system finds the most relevant passages.
- The model receives the question plus those passages and answers based on the documents, with the ability to cite them.
Like a new employee who is fluent but does not know the company yet: every time a customer asks, they open the right page of the handbook and answer.
Advantages: deployed in weeks; updated by editing documents, no retraining; sources can be cited for human checking; uses the best commercial models without any training data.
Limits: quality depends on the documents (contradictory documents give contradictory answers); initial effort to organize documents; questions requiring synthesis across many long documents remain hard.
Fine-tuning: teach the model how to answer
How it works: prepare thousands of "question, model answer" pairs that match how the business wants to respond, then train the model further on them. The tuned model answers in the learned voice, format and habits.
Like training an employee for months so they speak in the company's style without opening the handbook.
Advantages: stable voice and format; handles specialized data well (product codes, industry jargon, custom document structures); can run a smaller, cheaper-per-call model.
Limits: needs quality training data, usually thousands of examples; every knowledge change requires retraining, costing money and time; the model can still fabricate because fine-tuning teaches how to speak, not new facts; hard to cite sources.
The most common misunderstanding: fine-tuning is not a way to load new knowledge into a model. If you want the model to know this month's price list, use RAG.
Head-to-head comparison
- Time to first usable version: RAG 2 to 4 weeks; fine-tuning 6 to 12 weeks (mostly data preparation).
- Upfront cost: RAG low (document organization, search component); fine-tuning higher (data, training, evaluation).
- When knowledge changes: RAG edit the documents; fine-tuning retrain.
- Source citation: RAG yes; fine-tuning no.
- Fabrication risk: RAG lower when documents are complete; fine-tuning still present.
- Custom voice and format: RAG decent via instructions; fine-tuning strong.
- Cost per call: RAG slightly higher because passages are sent along; fine-tuning can be lower with a smaller model.
Decision table for 6 common situations
- Answering customers from policies, prices, guides: RAG. Knowledge changes often, citations needed.
- Internal procedure lookups for staff: RAG. Documents exist, accuracy matters more than style.
- Summarizing files, contracts, tickets: RAG or just a well-instructed model; no fine-tuning needed.
- Classifying thousands of daily requests against a custom label set: instructions with examples first; if accuracy is insufficient and you have a large labeled dataset, consider fine-tuning a small model for lower cost.
- Writing content in a very distinctive brand voice: instructions with samples first; fine-tuning when high consistency is needed at large volume.
- Handling specialized data (technical codes, narrow industry terms) that commercial models misread: fine-tuning, usually combined with RAG for current knowledge.
Short rule: start with good instructions, add RAG when you need proprietary knowledge, fine-tune only after both have been tried and fall short. In practice, most businesses that begin with fine-tuning end up adding RAG anyway, because the knowledge they wanted the model to have keeps changing.
What decides success is not the method
Whichever you choose, three things determine whether the system is still usable after six months:
- Clean documents and data: RAG on contradictory documents and fine-tuning on wrong examples both produce poor results. From spreadsheets to smart systems is the foundation to lay first.
- A fixed test set: 50 to 100 sample questions with correct answers, rerun whenever documents change, instructions change, or the provider updates the model version. Without it you cannot tell whether the system is improving or degrading.
- Someone operating it: reading logs, reviewing answers customers complained about, updating documents, tracking cost. This is what Siri9 does under AI system management after building under AI integration.
Foundational concepts on accuracy, overfitting and drift are in machine learning fundamentals for business leaders.
Frequently asked questions
Is RAG secure if documents are sent to the model?
The system sends only a few passages relevant to the question, never the whole document store; use a provider that commits not to train on customer data; mask identifying information when not needed. For highly sensitive data, the model can run in private infrastructure at higher cost.
Does a small business ever need fine-tuning?
Rarely. Only when volume is so high that per-call cost becomes an issue, or the data type is extremely specialized. Most stop at good instructions plus RAG.
What does RAG cost for a small business?
Building from a few tens of millions of dong depending on the number of document sources and systems to connect; monthly operations include model fees (1 to 3 million dong at small volumes) plus AI system management.
How do we start?
Send Siri9 the kind of questions you want AI to answer and the documents you have; we assess within 24 hours whether instructions, RAG or fine-tuning fits, with a fixed scope and price.
