AI Assistant Cost: Build vs Buy

Published by Bles Software, a custom software and AI company based in Yehud-Monoson, Israel, building web apps, AI agents and API integrations for clients in Israel, the US, the UK and the EU.

What you are choosing between when you buy or build an assistant

Two bills arrive for an AI assistant and they are not the same shape. Buying one prices a seat for every person who touches it, or a flat fee for each ticket it resolves. Building one prices every token it reads and writes, plus the engineers who keep it honest. The numbers below are list prices read on 9 October 2026, so you can price your own case rather than ours.

The reason the answer flips is that one bill scales with headcount and the other scales with usage. A team of forty that each touch an assistant twice a day is cheap to buy for and expensive to build for. A support queue of sixty thousand conversations a month is the other way around.

The buy side, as the vendors publish it

Two details in that list decide more than the prices do. Intercom counts an outcome when the customer confirms the issue is resolved, does not come back after Fin answers, or a workflow completes, and it charges a conversation once [1]. That is a price per unit of work, so it tracks your volume rather than your payroll. The Microsoft figure is a promotional one that runs to 31 December 2026 for eligible existing Business customers [2], so a budget built on it has a known expiry date.

The build side, as the model vendors publish it

Anthropic publishes a caching multiplier rather than a second price list: a five minute cache write costs 1.25 times the base input rate, a one hour write costs 2 times, and a cache read costs 0.1 times the base rate, falling to 0.05 times on Opus 5.5 and Sonnet 5.5 [4]. Its Batch API takes 50 per cent off both input and output [4]. OpenAI publishes the cached input rate directly in the list above, which on gpt-6.1-sol is a twentieth of the uncached rate [5].

One warning that belongs next to any token estimate you were given before this year. Anthropic states that Claude 4.7 and later models use a tokenizer producing approximately 30 per cent more tokens for the same text [4]. A spreadsheet written against an older tokenizer understates the bill by roughly that much.

The part of the build bill that is not tokens

Token prices are the visible half. The other half is the people, and public wage data prices it without anyone having to take our word for it. The US Bureau of Labor Statistics puts the mean hourly wage for software developers at $71.20 and the mean annual wage at $148,100 in its May 2025 occupational survey, with a median annual wage of $135,980 [6][7]. Stack Overflow's 2025 survey gives United States medians by role rather than a single figure, including $189,500 for an AI and machine learning engineer and $180,000 for an architect [8].

Those are wages, not what an employer pays, and not what an agency charges. Use them as the floor under any build estimate: the engineering time around an assistant is evaluation harnesses, prompt and tool regressions, a retrieval layer, an escalation path to a human, and the logging that lets you prove what it said to whom. None of that is optional and none of it is in the token price.

Where the two bills cross

The crossover is arithmetic once you have your own two inputs: how many people would hold a seat, and how many units of work the assistant would actually complete.

Take a support queue. At the published $0.99 per outcome [1], ten thousand resolved conversations a month is $9,900. Pricing the same ten thousand conversations as tokens needs your own measured conversation length, and that measurement is the whole exercise: multiply average input and output tokens per conversation by the rates above [4][5], add the cache reads for the instructions you resend every time, and then add the engineering months. For most teams the token bill alone comes out below the per outcome bill well before ten thousand, and the total stays above it until the volume is large enough to amortise the engineering.

Take an internal assistant instead. At $18.00 per user per month [2] a department of sixty costs $1,080 a month with nothing to build, nothing to evaluate and nobody on call. A built equivalent has to beat that number including the people who maintain it, which at the wage figures above [6] is difficult below a few hundred users.

The questions that actually decide it

Ask these before you price anything, because each one changes the answer by more than any discount will.

What we would do with a week

Measure before choosing. One week of logging on the real queue gives you volume, average conversation length and the share of conversations a human has to finish. With those three numbers both bills above become arithmetic rather than opinion, and the decision usually makes itself. We build the measurement first for exactly this reason, and we would rather tell you to buy a seat than build something that costs more than it saves.

If you want that measured on your own numbers, the fastest route is a short call and read access to one month of conversation logs.

Sources

More Costs and Timelines from Bles Software