AI and automation · 9 min read
AI chatbots for business: what they cost in India and how RAG works
An FAQ bot, a RAG assistant and an agent that acts cost very different amounts to build and to run. Here is what each one costs in India, how RAG keeps answers honest, and where the hand off to a person belongs.
Most AI briefs reach our Mumbai studio with a screenshot attached. A chatbot, usually a quick pilot on a general model, has told a customer something confident and wrong: a return window that does not exist, a discount nobody approved, a delivery date pulled from thin air. The founder wants to know two things. What would it cost to do this properly? And can it ever be trusted?
Both answers depend on which kind of chatbot you need. There are three, they cost very different amounts, and most businesses need a simpler one than they expect. Here is what each costs to build and run, how retrieval augmented generation keeps answers honest, and the two things that decide whether customers trust the result: guardrails, and the hand off to a person.
The short answer: AI chatbot cost in India
These are the ranges we quote in 2026 for production work, with testing, guardrails and hand off included. USD figures are rounded at roughly ₹85 to the dollar.
| Tier | What it does | Build cost | Running cost a month | Typical build time |
|---|---|---|---|---|
| FAQ bot | Answers a fixed set of approved questions | ₹2 to 5 lakh (US$2,500 to 6,000) | ₹5,000 to ₹15,000 (US$60 to 175) | 3 to 5 weeks |
| RAG assistant | Searches your documents and catalogue, answers with sources | ₹6 to 15 lakh (US$7,000 to 17,500) | ₹15,000 to ₹50,000 (US$175 to 600) | 6 to 10 weeks |
| Agent that acts | Checks orders, books slots and raises tickets in your systems | ₹15 to 40 lakh (US$17,500 to 47,000) | ₹40,000 to ₹1.5 lakh (US$500 to 1,800) | 10 to 16 weeks |
Running costs cover model usage, hosting, search and monitoring. WhatsApp charges and ongoing improvement sit on top, and we come back to both below.
Three tiers, and how to tell which one you need
Tier one: the FAQ bot
An FAQ bot answers the twenty questions your team answers every day, from content you have approved. Delivery times, payment options, store hours. It may use a language model to recognise the hundred ways people ask the same thing, but its answers come from a short, fixed list.
It is quick, cheap and predictable, and for a small business with stable questions it is often enough. Ask it anything outside the list, though, and the best it can do is offer a person.
Tier two: the RAG assistant
A RAG assistant reads before it speaks. Ask what the warranty covers on a model sold two years ago, and it searches your product data and policies, answers from what it found and links to the source.
This is where most serious work on an AI chatbot for customer service, or an AI chatbot for ecommerce, belongs, because it grows with your knowledge rather than with a list someone maintains by hand. Most of the assistants in our AI chatbot and automation work sit in this tier.
Tier three: the agent that acts
An agent does not just answer. It checks an order, books a service slot, raises a return or updates a lead in your CRM. That makes it the most useful tier and the one that needs the strongest fences: a wrong answer can be corrected, but a wrong refund has already left the account.
Agents cost more because most of the effort goes into integration and safety: proper APIs, tight limits on what the agent may touch, a log of every action and a person approving anything that spends money or changes a customer’s record. We usually begin agent programmes with a three week Discovery Sprint, from ₹3.5 lakh (US$4,000), before anyone commits to a build.
A quick way to choose
Take a week of real customer messages, names removed, and sort them into three piles. Questions with one fixed answer point to an FAQ bot. Questions that need something looked up point to a RAG assistant. Messages that end with someone opening another system point to an agent. I still do this with a printout and three coloured pens. It takes an afternoon, and the biggest pile nearly always tells you where to start.
How RAG works, in plain words
A general language model answers from what it absorbed in training: a vast, slightly dated memory of the public internet. It knows nothing about your return policy, and unless told otherwise it will not admit that. RAG, short for retrieval augmented generation, hands the model an open book. Only your book.
- Collect. Gather the sources you approve: product sheets, policies, help articles and manuals.
- Prepare. Split each document into short passages, tagged with their source and the date they were last updated.
- Index. Turn every passage into an embedding, a list of numbers that captures its meaning, and store it in an index that also supports plain keyword search.
- Retrieve. When a customer asks something, find the handful of passages most likely to hold the answer.
- Answer. Give the model the question and those passages, with instructions to answer only from them, cite the source and say plainly when the answer is not there.
So what is RAG in AI, really? It is the difference between an assistant that sounds right and one that can show its working. It also beats fine-tuning a model on company data for customer questions: cheaper, current the day a document changes, and clear about its source.
Where RAG goes wrong
Nearly every broken assistant we are asked to rescue fails because of its documents, not its model. Three returns policies that disagree. A price list last touched in 2023. Scanned PDFs nobody can search. No named owner for the knowledge. A knowledge audit in the first fortnight, listing what is current and who looks after it, prevents most of this.
Where the money goes
The build
Chatbot development cost is mostly skilled people’s time. The model is the cheapest part. A typical RAG assistant budget divides roughly like this:
| Work | Share of the build |
|---|---|
| Discovery and knowledge audit | 10% |
| Conversation design, fallbacks and hand off moments | 10% |
| Retrieval pipeline and integration with your store, CRM or help desk | 35% |
| Guardrails, testing and an evaluation set | 20% |
| Channels and the shared inbox | 15% |
| Staged launch and training | 10% |
If a quote has no line for testing, the testing will be done by your customers.
The monthly bill
Model providers charge per token, roughly a fragment of a word, in and out. A RAG reply sends the question plus the retrieved passages, perhaps 3,000 tokens, and receives about 300 back. At 2026 list prices for a capable mid-sized model, that is 10 to 40 paise a reply. Prices vary by provider and keep falling, so we estimate per thousand conversations and set spending alerts.
Around that sit hosting, the search index, monitoring and Meta’s WhatsApp charges. Meta bills business-initiated template messages by category, replies inside the 24 hour customer service window are free, solution partners add a platform fee and GST applies. Check Meta’s current rate card before you budget.
The last line is improvement: a monthly read of real conversations, new answers where the assistant hesitated and tighter rules where it was too sure. Our Care & Grow plan covers that from ₹40,000 a month (US$500). A simple FAQ bot usually needs only a quarterly look.
A worked example: 6,000 conversations a month
Take a hypothetical D2C home fragrance brand in Bengaluru with sixty products on its own Shopify store. It handles about 6,000 conversations a month, a third of them after 9pm, and two people answer them. It wants a RAG assistant on its website and WhatsApp, in English and Hindi, that answers from product sheets and policies, checks order status and hands everything else to the team.
The build comes to about ₹9 lakh (US$10,500) over eight weeks, including a sprint spent trying to make the assistant misbehave. Running costs look like this:
| Item | Assumption | Monthly cost |
|---|---|---|
| Model usage | 36,000 replies at about 20 paise each | ₹7,000 (US$85) |
| Search index and hosting | PostgreSQL with pgvector, one application server | ₹9,000 (US$105) |
| Monitoring | Logs, evaluation runs and cost alerts | ₹3,500 (US$40) |
| Order updates and the partner’s platform fee | ₹4,500 (US$55) | |
| Total | About ₹24,000 (US$285) |
Now the return. Suppose the assistant resolves 40% of conversations end to end, close to the 41% that the bilingual concierge we built for a D2C skincare brand in Dubai resolves today. That is 2,400 conversations at about four minutes each, or 160 hours. The other 3,600 arrive already summarised, saving perhaps a minute apiece. Call it 220 hours a month, worth about ₹55,000 at a loaded cost of ₹250 an hour.
Net of running costs, staff time alone repays the build in roughly two and a half years. That is why we rarely make the case on headcount. The stronger case is the third of messages that arrive after 9pm and used to wait until morning. If instant answers turn just 50 of them into orders a month at ₹1,500 each and a 60% gross margin, the build repays in about a year. We would still prove the resolution rate first, with a two to three week pilot on one channel.
Guardrails: write the never list first
The first document on every AI project we run is not a prompt. It is a short list headed things it must never do. Never give medical or legal advice. Never promise a refund. Never guess a delivery date. Never reveal another customer’s details. Then the list is enforced in layers:
- Grounding. Answers come only from approved sources, and “I do not know, let me find someone who does” counts as a correct reply.
- Scope. It politely declines topics outside its job, however cleverly they are phrased.
- Injection testing. For a sprint before launch, we try to trick it with hidden instructions, fake policy text and role play.
- Evaluation. A scored set of a few hundred of your real questions runs before every release and before any model switch.
- Limits. Agents get the narrowest permissions, spending caps and a person approving anything that cannot be undone.
- Monitoring. A monthly read of conversations, with cost and error alerts in between.
This is the slowest part of our work on AI assistants and agents, and the one part we will not skip.
The human hand off
The debate about chatbot vs live chat is mostly a false choice. The setups that work use both in one inbox: the assistant takes the first reply at any hour, and a person takes over when needed. Agree the rules before launch. A person should step in when:
- the customer asks, in any words: “human”, “agent” or “kisi se baat karao”;
- the topic is sensitive, such as a complaint, a health or legal question or a failed payment;
- an order or refund passes a value you set;
- the assistant has misunderstood twice, or found nothing reliable to answer from;
- the tone turns angry.
The hand off is part of the product. The person receives a two line summary and the conversation so far, so the customer never repeats themselves. After hours, the assistant says honestly when someone will reply and offers a callback, rather than promising what nobody is awake to deliver.
Data, privacy and the DPDP Act
India’s Digital Personal Data Protection Act, 2023, and the rules notified under it in 2025, are being phased in through 2026 and 2027. For a chatbot the practical duties are clear: say what you collect and why, collect only what the job needs, obtain consent where required, and delete data once the purpose ends or the customer asks.
In practice, that means masking phone numbers and addresses before text reaches a model, choosing providers whose terms exclude training on your data, keeping logs for a fixed period and recording where data travels. Brands selling across the Gulf face similar questions under UAE and Saudi law, which is why our work with Gulf teams settles where data lives before anything is built. This is how we build, not legal advice; have your counsel review the privacy notice.
A short checklist before you brief anyone
- Collect a week of real customer messages, with names removed.
- Sort them into fixed answers, look-ups and actions, and count each pile.
- List the documents the assistant may use, and name an owner for each.
- Write your never list: five to ten things it must not do.
- Decide when a person takes over, and who that person is at 11pm.
- Choose one first channel: website, app or WhatsApp.
- Agree three measures, such as resolution rate, reply time and customer rating.
- Ask every vendor for running costs per thousand conversations, separate from the build.
The takeaway
Start with the smallest tier that clears your biggest pile of messages, prove it on one channel with your own conversations, and only then give it more to do. The model is the cheap part. The knowledge, the guardrails and the hand off are where a business chatbot earns trust, and where the budget should go. For a second pair of eyes on your piles, or on a quote you already hold, send us a week of messages; a partner replies within one working day.