8 min read

AI chatbot development: RAG vs fine-tuning vs rules

Train it on our data can mean three very different things. How to choose between rules, retrieval and fine-tuning for each kind of question your chatbot has to answer.

In AI chatbot development, choose rules for fixed, high-risk answers, retrieval-augmented generation (RAG) for answers that depend on your changing documents and data, and fine-tuning only for consistent style, format or narrow classification that prompting cannot achieve. Most production chatbots combine rules and RAG; fine-tuning is rarely the first step.

This guide explains how to make that choice per question type rather than per project. It is based on implementation work and on inspected code in the ImadDhin portal and the public FoCoCo case study. It does not claim accuracy figures, because accuracy depends on your content, your users and how you evaluate.

Why the choice gets confusing

Buyers often ask for a chatbot trained on their data. Vendors say yes, and each means something different. One means uploading documents to a retrieval index. Another means fine-tuning a model on example conversations. A third means a scripted flow with a language model polishing the wording. All three get called training.

The confusion matters because the approaches fail differently. Rules fail by being rigid: anything outside the script gets a dead end. RAG fails when retrieval misses the right passage or finds an outdated one. Fine-tuning fails quietly: the model sounds confident and on-brand while stating facts that were true at training time, or were never true at all.

What teams often get wrong

The most expensive mistake is fine-tuning to teach facts. Fine-tuning adjusts how a model responds; it is a poor way to store knowledge you need to update, delete or cite. When a price or policy changes, a fine-tuned model keeps the old answer until you retrain, and it cannot show where an answer came from.

The second mistake is treating RAG as upload and forget. Dumping every PDF, wiki page and old email into a vector database produces a bot that retrieves contradictory versions of the same policy. Retrieval quality is mostly a content problem: which sources are authoritative, how they are split, and how they are kept current.

The third mistake is letting generation handle answers that must be exact. Prices, legal disclaimers, eligibility rules and emergency instructions should come from rules or structured data, with the model at most formatting the result.

  • Fine-tuning to inject knowledge that changes
  • Indexing every document without choosing authoritative sources
  • Relying on the prompt to hide documents a user should not see
  • Generating exact values instead of reading them from data
  • Shipping without a test set of real questions

A practical approach: decide per question type

List the questions your chatbot must handle, grouped by type, and choose a mechanism for each group. A single bot usually contains all three.

Rules for fixed and high-risk answers

Use deterministic handling for intents where the answer must not vary: opening a refund request, emergency or safety instructions, legal notices, escalation to a person, and anything regulated. Rules also act as guardrails around the model, such as blocking topics the bot should never discuss and forcing a handoff on specific triggers.

RAG for knowledge that lives in documents

Use retrieval for questions answered by your policies, manuals, product documentation and help articles. A production RAG pipeline includes ingestion from named authoritative sources, cleaning, splitting by document structure rather than fixed character counts, metadata such as product, region and effective date, access filters applied at retrieval time, and often a combination of keyword and semantic search with reranking. The answer should cite its sources, and the bot should say it does not know when nothing relevant is retrieved.

Tools for live data

Questions such as where is my order or what is my balance should not be answered from an index at all. Give the bot read-only tools that query the system of record with the verified user's identity. Exporting live data into documents for retrieval creates stale answers and privacy problems.

Fine-tuning for behavior, not facts

Consider fine-tuning when you need consistent output structure at high volume, a narrow classification task such as routing tickets, a distinctive tone that prompting cannot hold, or a smaller and cheaper model for one well-defined job. It needs a good set of labeled examples, an evaluation set, and a plan to repeat the work when you change base models.

Combining them in one conversation

A typical turn runs through the layers in order. Rules check first for triggers that must be handled deterministically, such as a request for a person or a restricted topic. If none match, the model decides whether the question needs a live tool, retrieval, or neither. Retrieved passages and tool results are passed to the model with instructions to answer only from them, cite the source, and say so when they do not cover the question. A final rule layer can check the output for forbidden content or missing citations before it reaches the user. Logging which layer produced each answer makes failures much easier to diagnose later.

Implementation considerations

Assign content ownership. Every indexed source needs an owner who knows when it changes. Without that, the index drifts away from reality, and the bot confidently quotes last year's policy.

Handle document versions explicitly. Store an effective date and a superseded flag with each source so retrieval prefers the current policy, and remove replaced documents rather than leaving both in the index. Where regional or product variants exist, tag them and filter by the user's context, so a customer in one market is not quoted another market's terms.

Enforce permissions outside the model. If some documents are internal or customer-specific, filter them at retrieval based on the user's identity. A prompt instruction such as do not reveal internal documents is not access control.

Treat retrieved content as untrusted input. Documents, web pages and user uploads can contain instructions aimed at the model. The model should not be able to take actions or reveal data just because a retrieved passage told it to.

Build an evaluation set before launch: real questions, the expected answer, and the source that should support it. Run it whenever you change content processing, prompts or models. Include questions the bot should refuse or escalate.

Understand the cost drivers. RAG adds ingestion, embedding and storage, plus the tokens of retrieved passages in every request. Fine-tuning adds training runs and repeated effort on model upgrades. For a small, stable document set, placing the documents directly in a long-context prompt can be simpler than either.

Trade-offs

  • Rules: predictable and auditable, but brittle outside their script and costly to maintain for many intents
  • RAG: updatable, citable and permission-aware, but only as good as your content and retrieval tuning
  • Tools over live data: accurate and current, but require secure integration and identity verification
  • Fine-tuning: consistent behavior and potentially cheaper inference for narrow tasks, but slow to update and weak for facts

The combination most teams end up with is rules for safety and escalation, RAG for documents, tools for live data, and a general-purpose model doing the language work. Fine-tuning comes later, if evaluation shows a specific behavior that prompting and retrieval cannot fix.

Lessons from ImadDhin work

These are implementation and code-level observations, not measured accuracy or business results.

The public FoCoCo case study describes a coaching system that connects a golfer's conversations, practice, rounds and reflection across a phone app and a web app. That kind of personal context has to be retrieved per user at request time. It cannot sensibly be fine-tuned into a shared model, both because it changes after every round and because one user's data must never shape another user's answers.

In the ImadDhin portal, the chat agent's instructions and model selection come from remote configuration rather than from code or a trained model. Changing how the assistant behaves is a configuration change, reviewed and reversible, not a retraining project. Saved memory and chat history are stored server-side with direct client access denied by security rules, so the browser cannot read or rewrite what the assistant remembers.

Web research is also a separate, gated tool rather than something the model is assumed to know. Keeping live lookup as an explicit tool makes it clear when an answer came from a source and when it came from the model.

Common mistakes to test for

  • Ask a question whose answer is not in your content and confirm the bot says so
  • Update a policy and confirm the old version is no longer retrieved
  • Ask, as a user without access, about a restricted document
  • Plant an instruction inside a test document and confirm the bot ignores it
  • Check that prices and eligibility rules match your data exactly
  • Verify that each cited source actually supports the answer given

When a simpler solution is better

If your questions are few and stable, a searchable FAQ or a guided decision tree may serve users better than any generative chatbot. If your documentation is small, a single well-written prompt containing it can outperform a complex retrieval pipeline and is easier to maintain.

Many helpdesk products already suggest articles as customers type. Before building a custom bot, check whether that plus a clear route to a person solves most of the problem.

Choose the mechanism before the model

Group your questions, decide which need rules, retrieval, live tools or fine-tuning, and build an evaluation set before you pick a vendor. If you want help designing a chatbot that answers from the right sources, explore ImadDhin's AI automation consulting or AI product development, or talk it through in a 30-minute call.

Frequently asked questions

Is RAG or fine-tuning better for a company chatbot?

For answers based on company documents and policies, RAG is usually the better starting point because content can be updated, cited and filtered by permission. Fine-tuning suits consistent behavior, format or narrow classification, not changing facts.

Does fine-tuning stop a chatbot from hallucinating?

No. A fine-tuned model can still state incorrect facts, often in a more convincing tone. Grounding answers in retrieved sources, refusing when nothing relevant is found, and evaluating against real questions are more effective controls.

Can a chatbot answer questions about a customer's own orders?

Yes, through read-only tools that query your system of record after the customer is verified. Exporting live order data into a retrieval index creates stale answers and privacy risk.

How do we keep a RAG chatbot up to date?

Index only named authoritative sources, give each source an owner, re-ingest on change, and remove superseded versions. Run an evaluation set after content changes to catch regressions.

When is a rules-based chatbot enough?

When the set of questions is small and stable, or when answers must be exact and regulated. Many production bots use rules for these intents and a model only for the rest.

Build a chatbot that answers from the right sources

Bring your question types and content sources.

Book a 30-minute call

Chatbot architecture, retrieval and evaluation.

AI automation consulting

Production AI features inside your product.

AI product development

Keep reading