What an AI Support Agent Actually Costs in 2026, and When Intercom Fin's $0.99 Is the Cheaper Choice
In short
Below about 1,500 conversations a month, buy Intercom Fin: at $0.99 per outcome, handoffs included, no $12k to $25k build pays back within a year. Above about 3,100 a month with a stable question mix, build: tokens on Claude Sonnet 5 come to about eight cents a conversation, so even a $25k build pays back inside twelve months. At 20,000 a month, build and spend the saving on evals, a refuse path and monitoring, the costs most quotes leave out.

On this page
- What does Intercom Fin actually charge, and what counts as an outcome?
- What does one support conversation cost to run on your own stack?
- At what volume does a custom build pay for itself?
- What costs are missing from most AI support agent quotes?
- How do you cut the running cost of an AI support agent?
- Why do published AI chatbot cost estimates vary so much?
- What should a fixed-price AI support agent quote include?
- Should you build your own support agent or buy Fin?
Short answer: the model tokens are the cheap part of a custom AI support agent. At Claude Sonnet 5's list price ↗ of $2 per million input tokens and $10 per million output tokens, a six-turn support conversation costs about eight cents on my assumptions, while Intercom Fin charges $0.99 per outcome ↗, and handoffs count as outcomes. What decides build versus buy is the fixed cost around those tokens: the build itself, hosting, evals and maintenance, none of which shrink when your volume does. Counting the build price over a 12-month horizon, Fin is cheaper under roughly 1,500 conversations a month, a $12k to $25k build pays for itself inside a year from roughly 1,650 to 3,100 conversations a month, and at 20,000 a month building wins easily. If a build runs over budget, look at evals, guardrails and model migrations before you blame the tokens.
What does Intercom Fin actually charge, and what counts as an outcome?
Intercom's pricing page ↗ lists Fin at $0.99 per outcome, where an outcome is a resolution, a procedure handoff or a disqualification. Sales qualifications cost $9.99 each, and standalone use carries a 50-outcome monthly minimum.
The word people skim past is handoff. When Fin passes a conversation to a human through a procedure, you pay $0.99 and then you pay your human. The price does not mean "you only pay when the bot solves it", which is how it first reads.
For the maths below I treat Fin as $0.99 per conversation. That is the worst case for Fin: every conversation ends in exactly one billable outcome. If some of yours end with nothing billable, Fin costs less than my numbers say and a build takes longer to pay back. I've left your helpdesk subscription out of both sides, because you keep paying for one either way.
A third-party pricing roundup ↗ reports Zendesk at $1.50 to $2.00 per automated resolution. I haven't checked that against Zendesk's own pages, but if it holds, $0.99 is not the expensive end.
What does one support conversation cost to run on your own stack?
Apart from the list prices, every number below is an assumption. Replace them with yours.
ASSUMPTIONS (mine; replace with yours)
Model Claude Sonnet 5: $2 per 1M input, $10 per 1M output
Turns per conversation 6
Input per turn 3,000 tokens (system prompt, retrieved articles, history)
Output per turn 400 tokens
Tokenizer margin 30% (Claude 4.7 and later count more tokens per text)
Infra, flat $300 a month (vector store, server, tracing, uptime)
Maintenance 20% of the build price per year
PER CONVERSATION
Input 6 x 3,000 = 18,000 tokens at $2 per 1M $0.036
Output 6 x 400 = 2,400 tokens at $10 per 1M $0.024
Subtotal $0.060
With the 30% tokenizer margin ($0.060 x 1.3) $0.078
MONTHLY, 5,000 CONVERSATIONS $12k build $25k build
Tokens (5,000 x $0.078) $390 $390
Infra $300 $300
Maintenance (20% a year / 12) $200 $417
Build, running cost $890 $1,107
Fin (5,000 x $0.99) $4,950 $4,950
Monthly difference $4,060 $3,843
Payback on the build price 3.0 months 6.5 monthsThe tokenizer line is there because Anthropic's pricing page ↗ says Claude 4.7 and later models use a tokenizer that produces about 30% more tokens for the same text. If you sized your prompts with an older tokenizer or a words-per-token rule of thumb, you are short by roughly that much. Same list price, bigger bill.
On infra, a managed vector store has a floor (Pinecone's Standard plan has a $50 monthly minimum ↗), while pgvector on the Postgres you already run adds no new subscription. On maintenance, t3c.ai ↗ quotes 15 to 25% of build cost per year for iteration. That is a consultancy's estimate, not an invoice, so I took a number near the middle and labelled it as mine.
The assumption I trust least is 3,000 tokens a turn. If your agent calls tools on most turns (look up the order, check the plan, open the ticket), every call is another round trip that sends the context again. Anthropic reports ↗ that agents typically use about 4x the tokens of chat. At 4x, tokens for 5,000 conversations come to $1,560 a month, and payback moves from 3 months to about 4 for the $12k build and to a little over 9 for the $25k one. Swap in Opus 5 at $5 and $25 per million ↗ and the token line grows 2.5x. At 5,000 conversations that still doesn't flip the answer.
At what volume does a custom build pay for itself?
Same assumptions, three volumes:
1,000 a month 5,000 a month 20,000 a month
Fin at $0.99 $990 $4,950 $19,800
Tokens at $0.078 $78 $390 $1,560
Build running cost, $12k $578 $890 $2,060
Build running cost, $25k $795 $1,107 $2,277
Payback, $12k build 29 months 3 months 0.7 months
Payback, $25k build 128 months 6.5 months 1.4 months
Break-even volume for a 12-month payback:
(build price / 12 + infra + maintenance) / ($0.99 minus $0.078)
$12k build: ($1,000 + $300 + $200) / $0.912 = about 1,650 a month
$25k build: ($2,083 + $300 + $417) / $0.912 = about 3,070 a monthAt 1,000 conversations a month, buy Fin. The cheap build takes about two and a half years to pay back and the expensive one more than ten. I don't trust any payback longer than a year for something that sits on a model API, because the API may not last that long. OpenAI removed the Assistants API on 26 August 2026 ↗, which turned any support bot built on it into a migration project.
At 20,000 a month, build. Fin comes to $19,800 a month, $237,600 a year. The $25k build plus a year of running costs is about $52,300. With a gap that size you can afford to spend well past $25k on evals and monitoring, and you should.
In between, the build price decides it. If your agent is tool-heavy and you use the 4x token figure, the two thresholds move to about 2,200 and 4,100.
The payback figures leave two things out. They count from launch, and a build takes weeks before it answers anyone, during which you are paying Fin or people. They also don't price your own team's time reading transcripts and fixing help articles, which neither option does for you.
What costs are missing from most AI support agent quotes?
Five of them. Evals first, because they're the biggest.
Evals. The evals FAQ by Hamel Husain and Shreya Shankar ↗ says experienced practitioners report spending 60 to 80% of development time on error analysis and evaluation, and that an LLM-as-judge evaluator needs 100+ labelled examples plus weekly upkeep. LangChain's State of Agent Engineering survey ↗ found quality, not cost, is the top barrier to getting agents into production, yet only 52.4% of teams run offline evals while 89% have observability. For a support agent, the eval set is your real past tickets with the correct answer written down by your own support people, re-run on every prompt, retrieval or model change.
A refuse path, and the people behind it. A support agent that can't find the answer has to say so and hand off. The price of improvising is public: Cursor's support bot invented a one-device policy ↗ in April 2025, which led to cancellations and a public apology, and Air Canada was held liable ↗ for its chatbot's wrong fare advice. My own version came from a review-reply product I lead, where a fallback path once published generic English replies to Swedish businesses for days. Nothing errored. Every degraded path in that system now refuses instead of publishing. Refusing costs something too: a queue of held conversations that a person has to work through. That is staff time, and it is in nobody's quote. The pattern is in Citations or It Didn't Happen ↗.
Monitoring. For a live marketplace I built a 16-probe production monitor that runs every five minutes. The probes are the cheap part. The lasting cost is a person who answers the alert. An AI agent also needs checks an uptime monitor can't do: did a fallback model answer, is the reply in the customer's language, is the refusal rate drifting. In LiteLLM, a fallback is one line of config and shows up only in logs and the x-litellm-attempted-fallbacks header ↗, with no check on output quality, so alert on it.
Forced migrations. Besides the Assistants API, OpenAI's deprecations page ↗ lists whisper-1 shutting down on 26 February 2027. The tokenizer change above is the quiet version of the same thing. Put a model or API change in the budget as a certainty with an unknown date.
Hard caps on usage. On an AI avatar-video SaaS I built, each avatar-generation call cost roughly a dollar, and the first pricing page promised unlimited generation with no cap anywhere in the code. Every plan limit now has a check on the generation path that actually enforces it. A support agent carries the same risk at a smaller unit price: a tool call stuck in a loop, or two bots politely replying to each other. Cap turns per conversation, tokens per turn and spend per account per day, and log cost per account so you can see which customer is expensive.
Gartner predicted ↗ in June 2025 that over 40% of agentic AI projects will be cancelled by the end of 2027 over escalating costs, unclear value or inadequate risk controls. It's a forecast, not a count. The reasons still line up with this list uncomfortably well.
How do you cut the running cost of an AI support agent?
Caching and routing do most of the work. Deleting model calls you never needed does the rest.
Caching. Anthropic bills cache reads at 0.1x the input price ↗. In a support conversation the system prompt, tool definitions and policy text are identical on every turn, and so is the conversation so far. If two thirds of each turn's input came from cache, input per turn would drop from $0.006 to $0.0024 and token cost per conversation would fall by about a third. Cache writes are billed separately and I've left them out, so read that as the most you'd save.
Routing. On a multi-tenant RAG product I built, I run four models behind LiteLLM so each kind of query can go to the cheapest model that is good enough. Haiku 4.5 lists at $1 input and $5 output per million tokens ↗, half of Sonnet 5 on both. If half your conversations are "where is my invoice" and can go to Haiku, the token line drops by about a quarter on list prices. "Good enough" means it passes your eval set, one more reason evals come first.
Deleting calls. On that same product, every request used to make an LLM call just to detect the language of the question. I replaced it with a deterministic language detector, which took a whole model call out of every request and made the product noticeably faster. Language detection doesn't need a model. Neither does pulling an order number out of a message, which a regex does for free.
Batch and thinking tokens. Batch requests are 50% off ↗, useless for a live reply and fine for the nightly eval run. Watch reasoning models too. Google states that Gemini output prices include thinking tokens ↗, so a model that thinks for 1,500 tokens before a 400-token answer bills you for 1,900. Most support questions are lookups, not puzzles.
The mechanics are in How to Reduce LLM API Costs: Caching and Routing in 2026 ↗.
Don't count on falling prices to do this for you. Epoch AI ↗ measured per-token prices for a fixed capability level dropping 9x to 900x a year (the fastest drops are recent), and Stanford's AI Index ↗ found a GPT-3.5-level model got over 280 times cheaper to query between November 2022 and October 2024. Yet Menlo Ventures ↗ put enterprise LLM API spend at $3.5B in late 2024 and $8.4B by mid-2025. Cheaper tokens get spent on longer prompts and more agent steps.
Why do published AI chatbot cost estimates vary so much?
Because they are marketing, not invoices. t3c.ai ↗, a consultancy, puts a RAG chatbot build at $50K to $150K plus $2.5K to $15K a month to run. Ailog ↗, a RAG vendor, estimates a do-it-yourself build at $180.8K to $364K in year one. Kellton ↗ gives $15K to $300K+, with a median of $75K to $120K. None of them shows the invoice behind the number. I'm not publishing invoices here either, which is exactly why the model above shows every assumption.
They also price different jobs under one name. A bot that answers from one help centre is a smaller project than an agent that issues refunds across three internal systems. My numbers assume the first kind plus a few read-only actions like order lookup.
What should a fixed-price AI support agent quote include?
What I charge: a production RAG chatbot usually starts around $4k to $12k fixed. Once the bot takes actions it is agent work, which I price at $175 to $300 an hour or fixed per workflow. Integrating AI into an existing product is $120 to $200 an hour. I give a fixed quote after a short call.
Whoever builds it, a support agent quote should name these in writing. They are what I'd write into mine:
- Retrieval over your help centre and docs, with every answer citing the article it came from.
- The actions agreed on the call, such as order lookup or plan checks, wired to your APIs.
- An eval set built from your past tickets, and a script that re-runs it.
- A refuse path that hands low-confidence conversations to your helpdesk instead of guessing.
- Hard caps on turns, tokens and daily spend per account, with cost logged per account.
- Alerts when a fallback model answers or the refusal rate moves.
What sits outside it: your token bill (that goes on your own provider account, so you see every cent), your helpdesk or hosting, and rewriting your help content. An agent can only be as right as the articles it retrieves, and a wrong article becomes a wrong answer given politely to every customer who asks. Ongoing eval review and model migrations after handover are separate work, and should be agreed up front rather than discovered.
Should you build your own support agent or buy Fin?
The rule I'd apply, using the numbers above:
How many support conversations do you handle a month?
Under about 1,500
Buy Fin, or your helpdesk's own AI agent.
Look again when volume doubles.
About 1,500 to 3,100
Get a fixed quote for the build.
Quote near $12k, and most answers already in your docs?
yes: build
no: buy Fin
Over about 3,100
Is your help content accurate and your question mix stable?
yes: build, and spend part of the saving on evals and monitoring
no: fix the content first; bought or built, the agent repeats your docs
Tool call or action on most turns?
Move both thresholds up to about 2,200 and 4,100.If you sit between those numbers and want the assumptions checked against your real ticket volume, that is the short call I do before quoting. The build side is my AI agents and automation ↗ and RAG development ↗ work, and my contact page ↗ is open. At 1,000 conversations a month I would rather talk you out of a custom agent than sell you one.
FAQ
How much does Intercom Fin cost per resolution?
Intercom's pricing page lists Fin at $0.99 per outcome, and an outcome covers procedure handoffs to a human and disqualifications as well as resolutions. Sales qualifications are priced separately at $9.99, and standalone use has a 50-outcome monthly minimum. For planning I treat Fin as $0.99 per conversation, which is the worst case.
How much does it cost to run a custom AI support agent per conversation?
At Claude Sonnet 5's list price of $2 per million input tokens and $10 per million output tokens, a six-turn conversation with about 3,000 input and 400 output tokens per turn costs roughly $0.06, or $0.078 with a 30% margin for the newer Claude tokenizer. On top of that come fixed costs for hosting, retrieval, monitoring and maintenance, which at low volume cost more than the tokens. If the agent calls tools on most turns, budget for about four times the tokens.
At what volume is building an AI support agent cheaper than Intercom Fin?
On my worked assumptions, a $12k build pays back within 12 months from about 1,650 conversations a month, and a $25k build from about 3,070. Under roughly 1,500 conversations a month, Fin is the cheaper choice on a 12-month view. At 20,000 conversations a month, a build pays back in under two months on paper.
What hidden costs do AI chatbot quotes leave out?
Evals are the biggest one: in Hamel Husain and Shreya Shankar's evals FAQ, experienced practitioners report spending 60 to 80% of development time on error analysis and evaluation. Quotes also tend to leave out the refuse path and the staff who work its queue, production monitoring, hard usage caps, and forced migrations such as OpenAI removing the Assistants API in August 2026. Your own team's time reviewing transcripts is rarely priced at all.
How can I reduce the LLM cost of a support chatbot?
Cache the parts of the prompt that repeat on every turn, since Anthropic bills cache reads at a tenth of the input price. Send simple questions to a cheaper model such as Haiku 4.5, which lists at half of Sonnet 5's price, and replace model calls that don't need a model, like language detection, with plain code. Use batch pricing, which is 50% off, for offline work such as nightly eval runs.
Working on something like this?
I build web apps, AI features, and mobile products for clients. If this article matches a problem you have, tell me about it.
Start a conversationMalik Hamza Shabbir · Full-Stack & AI Engineer
I build full-stack and AI products solo: a reputation SaaS in production, RAG pipelines, and React Native apps. I write from what I ship, not from documentation summaries.
Run it on yours
Every number above came from code you can read.
The harness is open and the per-probe log ships with it. If you build one of these systems, or you are choosing between them with real money, the same method is available to you.
Send me an endpoint and I will write an adapter, run it on the same workload under the same rules, and publish the result. It goes in the table whether it wins or not.
Sponsored inclusion is disclosed in the article and the result publishes either way. Paying moves you up the queue, never up the table.
Related articles
I Benchmarked 5 AI Agent Memory Systems. The Best Scored 57.6%.
Mem0, Zep, Letta, LangMem and GoodMem on the same 419-turn conversation, graded only on the text they retrieved. Vendors publish 80-93% on LoCoMo. The best here scored 57.6%. Open harness, every number reproducible.
Private RAG on Local Models: Qwen3 vs Gemma 4 in 2026
Yes, you can ship private RAG on one 24GB GPU in 2026. I ran a 50-question eval: Gemma 4 26B MoE wins English corpora, Qwen3.6 27B wins multilingual.
RAG, Fine-Tune, or Just Prompt? A 2026 Decision Tree for Million-Token Context Windows
Cheaper long context in 2026 broke the old always-RAG advice. Here is the decision tree I use: when full-context prompting beats a pipeline, when RAG is mandatory, when fine-tuning earns its keep, plus the hybrid stack and a cost and latency comparison.