All articles

Building an AI Customer Service Chatbot with RAG

How retrieval-augmented generation lets a chatbot answer from your own FAQs, catalogue, and policies: preparing documents, chunking, hybrid search, live system data, human handover, and measuring quality.

AI Development|Published |10 min read
An illustration of artificial intelligence connected to a network of data

A customer service chatbot built on a language model alone will answer confidently about prices, policies, and products it has never seen. Retrieval-augmented generation (RAG) fixes this by retrieving the relevant passages from the business's own material for every question and instructing the model to answer only from them. As a result, most of the work in a good chatbot goes into the documents and the search around them, and comparatively little into the model.

Why RAG Instead of Fine-Tuning

Fine-tuning changes how a model writes, but it is a poor way to teach it facts that change. Prices, stock, opening hours, and policies are updated regularly, and each change would require another training run. With RAG, updating a document updates the chatbot's answers immediately. Every answer can also be traced back to the passage it came from, which makes errors much easier to diagnose.

ApproachGood forLimitation
Prompt onlyA quick demo with a handful of factsBreaks once the knowledge no longer fits in the prompt
RAGFAQs, catalogues, policies, and SOPs that change over timeAnswer quality depends on document quality and retrieval
Fine-tuningTone, format, and domain-specific phrasingExpensive to update and cannot cite sources for facts
Tool calls to live systemsOrder status, availability, account-specific dataRequires an API and careful access control

Prepare the Knowledge Base Before Writing Code

The chatbot can only be as accurate as the material it retrieves. Collect the questions customer service actually receives, usually from chat history, email, and the team's own notes, and check that each common question has a clear, current answer somewhere. Gaps found here are cheaper to fix than wrong answers found after launch.

  • Remove outdated and contradictory documents. If two price lists disagree, the chatbot will eventually quote the wrong one.
  • Rewrite answers that only make sense with internal context, such as references to an earlier email or to an internal system name.
  • Give each document an owner and a review date so the knowledge base does not quietly drift out of date after launch.
  • Keep structured data such as product catalogues in a structured form, with names, codes, and prices as fields, rather than flattening them into long prose.

Chunking: How Documents Are Split for Retrieval

Documents are split into chunks, each chunk is converted into an embedding, and the embeddings are stored in a vector database such as pgvector or Qdrant. Chunk size is a trade-off. Chunks that are too small lose the context that makes them meaningful, and chunks that are too large dilute the match and fill the context window with irrelevant text.

  • Split along the document's own structure, such as headings, FAQ entries, or product records, rather than at a fixed number of characters.
  • Keep one question and its answer together in the same chunk for FAQ material.
  • Prefix each chunk with its document title and section heading, so a passage that says "the fee is waived" still carries what the fee is for.
  • Store metadata such as document type, product category, language, and last-updated date with each chunk, so retrieval can filter before it ranks.

Use Hybrid Search for Product Names and Codes

Vector search matches meaning, which handles questions phrased in many different ways. It is weaker on exact tokens such as product codes, model numbers, and proper names, where a customer expects an exact match. Combining vector search with keyword search, then merging or re-ranking the results, covers both cases. For an Indonesian business this also helps with mixed-language questions and informal spelling, which are common in chat.

When the chatbot gives a wrong answer, check the retrieved passages first. More often than not, the model answered faithfully from the wrong passage.

Ground the Answer and Allow "I Don't Know"

The system prompt should instruct the model to answer only from the retrieved passages, to say when the information is not available, and to offer a handover to a person in that case. A chatbot that admits it does not know is far less damaging than one that invents a refund policy. Returning the source passages with each answer also makes review and debugging straightforward.

Live Data Belongs in Tool Calls, Not in the Index

Order status, stock levels, booking availability, and account balances change by the minute and are specific to one customer. Indexing them as documents produces stale answers and risks showing one customer's data to another. Instead, expose these as tools that the model can call through your existing APIs, with the customer's identity verified by the application, not by the model. The model decides when to call the tool; the application decides what that customer is allowed to see.

Design the Handover to Human Staff

A chatbot should resolve routine questions and pass everything else to the team cleanly. Handover is triggered when the customer asks for a person, when the answer is not in the knowledge base, when the topic is sensitive such as complaints or payment disputes, or when the conversation goes in circles. The staff member receiving the handover should see the full conversation so the customer does not have to repeat themselves.

Measure Quality and Improve Continuously

  • Build an evaluation set of real customer questions with expected answers, and run it before every change to prompts, chunking, retrieval, or model.
  • Track the share of conversations resolved without handover, and read a sample of those conversations, because a resolved conversation can still contain a wrong answer.
  • Log the questions the chatbot could not answer. That list is the most direct guide to what the knowledge base is missing.
  • Monitor token usage and cost per conversation, and set limits on conversation length and retrieved context so a single long session cannot generate an unexpected bill.
  • Review answers on sensitive topics such as pricing, refunds, and legal terms more often than general questions.

How to Launch a RAG Chatbot Step by Step

  1. 1Collect the most common customer questions from chat history and agree which topics the chatbot will handle and which go straight to staff.
  2. 2Clean up the knowledge base, assign owners to documents, and fill the gaps found in the question list.
  3. 3Build the retrieval pipeline with structure-aware chunking, metadata, and hybrid search, and test retrieval on its own before connecting a model.
  4. 4Write the grounding prompt, add tool calls for live data, and implement handover with full conversation context.
  5. 5Run the evaluation set, fix the failures, then launch to a limited audience or a single channel such as the website or Telegram.
  6. 6Review unanswered questions and sampled conversations weekly, update the knowledge base, and expand to more channels once quality is stable.

After the Chatbot Goes Live

Most of the work after launch is keeping the documents correct. Prices change, promotions end, policies are revised, and the chatbot only knows what the documents say. Make document updates part of the team's routine, for example whenever a price changes, and use the unanswered-question log to decide what to add next.

Key takeaways

  • Use RAG for facts that change, such as prices, stock, and policies. Updating a document updates the answer, and each answer can be traced to its source.
  • Clean up and assign ownership of the knowledge base first. Retrieval cannot compensate for contradictory or outdated documents.
  • Split documents along their own structure, add titles and metadata to each chunk, and use hybrid search for product codes and names.
  • Fetch customer-specific and live data through tool calls with access control enforced by the application, not by indexing it.
  • Let the chatbot say "I don't know", hand over with full context, and use the unanswered-question log to improve the knowledge base.

Related articles

More articles on software development, AI, cloud, and infrastructure.

A smartphone showing a folder of messaging apps
AI Development|

AI Chatbots for Business: Use Cases, Channels, and Costs

A guide for business owners considering an AI chatbot: how it differs from a rule-based bot, when it pays off, examples by industry, choosing between website, Telegram, and WhatsApp, what drives the cost, and how to measure the results.

A Linux terminal showing an Ubuntu prompt with the sudo command
Developer Tools|

Setting Up WSL 2 for Software Development on Windows

A practical guide to developing on Windows with WSL 2: installation, where to keep project files, VS Code and Git, limiting memory with .wslconfig, Docker and systemd, networking, backups, and fixes for common problems.

Looking for a software development partner?

Tell us about your project, what you need to build, and the challenges you are facing. We can discuss the technical approach, scope, timeline, and estimated cost.

Start a conversation