AI Chatbot Development Cost: What It Really Costs in 2026
A custom chatbot or AI assistant costs $8,000 to $22,000 for a scoped single-channel build (4 to 7 weeks), $25,000 to $70,000 for a production assistant wired into your real systems (3 to 5 months), and $80,000 to $250,000+ for regulated, multi-channel or voice-enabled deployments (6 to 12 months). Most mid-market buyers land between $45,000 and $85,000. The number moves on integration count, compliance scope and the accuracy bar you set. Model choice barely touches it.
What a custom chatbot or AI assistant actually costs
Across 2,000+ projects at Digital Heroes, chatbot and assistant work sorts cleanly into three bands. The bands are set by how much of your business the assistant has to touch, not by the model behind it.
Tier 1: $8,000 to $22,000. Timeline 4 to 7 weeks. One channel, usually a web widget. One knowledge source, typically your docs, help centre or a few hundred PDFs. Retrieval over that content, a decent system prompt, guardrails against off-topic answers, and an email or ticket handoff when the bot is stuck. Roles: one senior engineer part-time, one conversation designer for two weeks, light QA. What falls OUT at this tier, and this is where cheap builds hurt: no write-backs into your CRM (Customer Relationship Management) or booking system, so the bot can answer but cannot do anything. No admin console, so every prompt or knowledge change comes back to the vendor as billable work. No evaluation harness, so nobody can prove accuracy beyond vibes. No multilingual. No mobile or WhatsApp. No analytics past raw transcript dumps. Tier 1 is a good answer engine and a bad employee.
Tier 2: $25,000 to $70,000. Timeline 3 to 5 months. This is the honest middle and where most serious buyers land. Two or three channels (web, in-app, WhatsApp or Slack). Three to six integrations, at least two of them read AND write, so the assistant books, updates, refunds or escalates instead of just talking. A retrieval pipeline over messy internal data with a real cleanup pass. An evaluation set of 150 to 400 graded questions so you can measure regression before every release. An admin console your ops team uses to edit prompts and knowledge without a ticket. Human handoff with full context. Roles: tech lead, two engineers, conversation designer, QA, part-time PM.
Tier 3: $80,000 to $250,000+. Timeline 6 to 12 months. Regulated data (health, finance, legal), or voice, or eight-plus integrations, or a multi-tenant assistant you resell to your own customers. Add security review, audit logging on every model call, PII redaction, role-based answers where a rep and a customer asking the same question get different responses, and a fallback path when the model provider has an outage. Roles expand to include a solutions architect, a security engineer and a data engineer.
What actually drives the number
Six variables move a chatbot quote more than everything else combined. Ask any vendor to price these explicitly.
1. Integration count and direction. $2,500 to $6,000 per read-only integration, $5,000 to $12,000 per read-write integration. Reading from a documented REST API is a week. Writing into a 2014 on-prem system with no sandbox, no docs and a rate limit is a month. A five-integration assistant carries $20,000 to $45,000 of pure integration work before a single conversation is designed. This is the single biggest line item in most builds.
2. Data preparation. $3,000 to $18,000. Nobody's knowledge base is ready. Scanned PDFs, conflicting policy docs, three versions of the pricing sheet. In every assistant we have shipped, content cleanup moved answer quality more than model choice did, and cleaning 2,000 documents is real labour. If a quote has no data prep line, the vendor is planning to dump raw files into a vector store and hope.
3. Compliance and security. Adds 18 to 35 percent to the whole build. HIPAA, SOC 2 alignment, GDPR data residency or financial audit requirements add PII redaction, encrypted logging, data processing agreements, access controls and a penetration test. On an $80,000 build that is $14,000 to $28,000. It is not optional and it is not a checkbox at the end.
4. Channels. $4,000 to $12,000 per additional text channel, $15,000 to $40,000 for voice. Web plus mobile is not one build with a different skin: session handling, auth, push and app store review are separate work. Voice is a different animal entirely, with speech-to-text, barge-in handling, latency budgets under 800ms and telephony infrastructure.
5. Accuracy bar and evaluation. $6,000 to $25,000. "Good enough for internal FAQ" and "safe enough to quote prices to customers" are different projects. The second one needs a graded test set, adversarial testing, confidence thresholds, refusal behaviour and a human escalation path. Skipping this saves $15,000 and costs you the first time the bot invents a refund policy.
6. Design depth. $2,000 to $20,000. A templated widget is $2,000. Branded UI with custom components, rich cards, file upload, typing states and full accessibility work is $9,000 to $20,000. Conversation design, meaning the actual scripting of failure paths and clarifying questions, is another $4,000 to $9,000 and is the difference between an assistant people use twice and one they use daily.
A worked example: support and booking assistant, 40-location services company
Real shape of a Tier 2 build that ran slightly hot. Web widget plus WhatsApp, answers service questions from 1,900 internal documents, books and reschedules jobs in the CRM, hands off to a human with context.
- Discovery, intent mapping, conversation design: $4,800
- Knowledge ingestion and cleanup, 1,900 documents: $7,200
- Retrieval pipeline, chunking strategy, evaluation harness: $11,500
- Core chat backend, orchestration, guardrails, refusal logic: $13,000
- CRM integration, read and write: $5,400
- Scheduling system integration: $4,800
- Ticketing handoff with transcript context: $3,200
- Branded web widget, accessible, responsive: $6,500
- WhatsApp channel: $4,200
- Admin console: prompt editing, knowledge updates, analytics: $8,000
- Evaluation set of 280 questions, red-teaming, UAT: $5,600
- Deployment, infrastructure as code, monitoring, alerting: $4,300
Subtotal: $78,500. Project management and cross-team QA at 10 percent: $7,850. Quoted total: $86,350, delivered in 19 weeks with an $8,600 contingency reserve that got half spent when the scheduling vendor's API turned out to have no sandbox. That build sits at the top of Tier 2. Strip WhatsApp ($4,200), the admin console ($8,000) and the evaluation work ($5,600) and the same assistant quotes at $66,770, and you pay most of that $19,580 back within a year in vendor tickets and bad answers.
The ongoing costs nobody puts in the quote
Model inference. Do the arithmetic yourself. Take conversations per month, multiply by turns per conversation, multiply by tokens per turn. A retrieval-heavy assistant sends roughly 3,000 to 5,000 input tokens per turn because it is stuffing retrieved documents into context. At 25,000 conversations a month and 8 turns each, that is around 800 million input tokens. Providers publish per-million-token rates, and those published rates differ by roughly an order of magnitude between a small model and a frontier one. For that same traffic it works out at a few hundred dollars a month on a small model and around $3,000 to $3,500 on a frontier model. Prototype on the expensive model, ship on the cheap one for routing and the expensive one only for hard turns.
Infrastructure and third-party services. Hosting, vector database, queueing and logging run $150 to $900 a month at this scale. Observability and evaluation tooling adds $100 to $600. WhatsApp and telephony are per-message or per-minute pass-through and go straight onto your bill.
Maintenance at 15 to 20 percent of build cost per year. On the $86,350 example, $13,000 to $17,300 annually. That is not padding. Model versions get deprecated, your CRM ships a breaking API change, retrieval quality drifts as content grows, and someone has to watch the failure transcripts weekly.
Year one change requests: budget 20 to 40 percent of build. This one gets ignored and it is the most predictable cost in the project. The business will ask for things. Once the assistant works, sales wants it on the pricing page, ops wants a new intent, someone wants Spanish. On an $86,350 build, reserve $17,300 to $34,500 for year one. If you do not reserve it, you will spend it anyway and it will feel like a failure instead of a plan.
How to not get burned on price
The cheapest quote is usually the most expensive outcome, because it is cheap for a boring reason: it is incomplete. A $19,000 quote against a $70,000 scope has silently removed data cleanup, write integrations, evaluation and the admin console. You discover this in month three, and the change orders to add them back cost 1.4 to 1.8 times what including them upfront would have, because retrofitting write access and audit logging into an assistant that was built read-only means rework, not addition. We have rebuilt enough of these to say it plainly: rescuing a failed cheap build costs 60 to 90 percent of a fresh build, and you paid for the first one.
What a change request should cost. Blended senior rates for this work run $85 to $150 an hour depending on region and seniority. A typical in-flight change (a new intent, an extra field written to the CRM, a copy and tone pass) is 6 to 16 hours, so $500 to $2,400. If your vendor quotes $9,000 for a new intent, either the architecture is wrong or the relationship is.
Contract terms that protect the number. Fixed scope with a written, itemised scope document, not a proposal paragraph. IP assignment on payment, in writing, covering prompts, evaluation sets and fine-tuning data, not just code. Source code in your repository from week one, not delivered at the end. Model and vendor keys in your accounts, so switching providers is a config change and not a hostage negotiation. A named change request rate agreed before kickoff. A 10 percent contingency line you both acknowledge exists.
How to brief a vendor so the quotes come back comparable
Most quote variance is brief variance. Two vendors pricing the same vague paragraph will differ by 3x because they are pricing different projects. Send every vendor the same one-page brief with these six answers and the spread we see tightens to about 25 percent, which is the real difference between them.
One: the exact list of systems the assistant must read from and write to, named, with a note on whether each has a documented API and a sandbox. Two: expected conversations per month and peak concurrency. Three: channels, named, ranked, with which are phase one. Four: your data. How many documents, what format, who owns them, how often they change. Five: your compliance reality, stated up front, including any data residency requirement. Six: what "working" means, written as three example conversations that must succeed and two that must be refused or escalated.
Then ask every vendor for the same three things back: a line-item breakdown, not a single number. A named team with named roles and hours. And their assumed change request rate. Any vendor who will not itemise is guessing, and you find out which parts they guessed wrong when it is too late to matter.
The evidence behind this guide
Independent findings on why this investment pays off. Every link goes to the primary source.
- McKinsey emphasizes that most L&D functions still fail to tie training to business outcomes, recommending organizations track 2-3 business-relevant indicators (such as time-to-proficiency, redeployment into priority roles, or frontline productivity) rather than participation metrics to demonstrate training effectiveness. Source: McKinsey & Company (2025) →
- In the Flexera 2025 State of ITAM report, respondents reported roughly 33% of SaaS spend is wasted, underscoring how paying for off-the-shelf seats and tiers that go unused erodes the supposed cost advantage of generic SaaS. Source: Flexera (2025) →
- Deloitte's research found that digitally advanced small businesses experienced revenue growth nearly 4x as high as the prior year, were about 3x as likely to have exported, were nearly 3x as likely to have created new jobs, and were more than 3x as likely to have seen more sales inquiries in the last year. Source: Deloitte (research summarized by Google) (2017) →
- Companies in the top quartile of McKinsey's Developer Velocity Index had 2014-18 revenue growth four to five times faster than bottom-quartile peers, showing that software-building capability is a driver of business performance, not just a support function. Source: McKinsey & Company (2020) →
Rohan advises mid-market and enterprise teams on ERP, CRM and custom software, and has led delivery on dozens of business-software builds.
Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.