What does a voicebot cost — and why you should count your unanswered calls first

A voicebot price list has four lines, not one. We sum two published Polish price lists over twelve months and show which number to measure at your own company.

Asked what a voicebot costs, vendors answer with a single monthly figure. It is a real number and the least useful of the four that make up the bill: alongside the subscription you also pay for minutes above the bundle, for the launch, and for every integration with a system where the data actually lives. Below we set those lines out, sum two published Polish price lists over twelve months, and show what really moves the bill. We start, though, with a number that appears in no price list at all — because that is the one which decides whether the project makes sense.

Start with the number that is in no price list

Before comparing offers, pull a three-month report from your own PBX and answer four questions: how many inbound calls you take per day, how many of them ring out unanswered, what hours those unanswered calls cluster in, and what they are usually about. That is one afternoon’s work, not a research project.

That number, not the price list, decides whether an implementation is worth doing. If five calls a day go unanswered and three of them are sales reps, no arithmetic in the rest of this piece turns it into a viable project. If forty go unanswered and half concern one repeatable matter — order status, appointment time, opening hours — you have a concrete scope to cost out.

We will not offer an “industry benchmark for missed calls” here. The figures circulating online come from vendor marketing and carry no stated method. Your own switchboard knows the value more precisely than any benchmark does. We approach every AI automation the same way: measure the process at the client first, quote second.

A voicebot price list has four lines, not one

The platform subscription. A fixed monthly amount for access to the conversation engine, the console and the recordings. It is the only figure vendors put on display, and the only one that does not depend on how much you actually call.

The minute bundle and the overage rate. The subscription includes a set pool of conversation minutes. Every minute beyond it is billed separately, and that rate tends to be the biggest surprise in the second quarter of the engagement.

The one-off setup fee. Launching the platform, configuring the script, recording the prompts, testing. The spread across Polish price lists is enormous — from nothing to the low tens of thousands of PLN net.

Integrations. Connecting the bot to a calendar, an order system, a CRM or an appointment schedule. These are quoted separately because they depend on the state of your data, not on the vendor’s price list.

A fifth line never appears in a price list because it is not a product: running cost. Conversation scripts, the pronunciation lexicon and the escalation thresholds change along with the offer, the prices and the opening hours. A system nobody maintains will, six months later, be sending callers to a promotion that no longer exists.

Two published price lists and twelve months

We read the figures below from the vendors’ own pages on 31 August 2026. They carry that date because price lists change, and a number without a date is useless.

odbierze.ai — LITE: 4,990 PLN setup fee plus 1,190 PLN per month, 500 minutes included, overage 1.49 PLN per minute. GROWTH: from 9,990 PLN setup plus 2,490 PLN per month, 1,500 minutes, overage 1.19 PLN per minute. ENTERPRISE is quoted individually. All amounts net, with 23% VAT added.

Smart Asystenci — from a pricing article published on 13 June 2026: Starter 499 PLN net per month, 500 minutes included, overage 0.78 PLN per minute; Pro 899 PLN, 1,000 minutes, overage 0.75 PLN; Pro+ 1,499 PLN, 2,000 minutes, overage 0.73 PLN. No setup fee.

Summed over twelve months, with traffic staying inside a 500-minute bundle:

  • odbierze.ai LITE: 4,990 + 12 × 1,190 = 19,270 PLN net
  • Smart Asystenci Starter: 0 + 12 × 499 = 5,988 PLN net

And with real traffic of 800 minutes a month, i.e. 300 minutes above the bundle:

  • odbierze.ai LITE: 19,270 + 12 × 300 × 1.49 = 24,634 PLN net
  • Smart Asystenci Starter: 5,988 + 12 × 300 × 0.78 = 8,796 PLN net

Both sums cover only the published list-price lines — subscription, setup and overage. Integrations and maintenance are quoted separately and sit on top of these figures, which is exactly why a monthly figure never answers the question in the title.

These are not equivalent products, and that is not the point — the scope of the launch, the level of support and what sits inside the starting fee differ between vendors and have to be compared separately. The point is the method: until you sum four lines over twelve months, across two scenarios for the minute count, you are comparing monthly figures that describe different things.

Note the spread on the overage rate alone: 0.73 to 1.49 PLN net per minute, roughly a factor of two. At three hundred minutes above the bundle each month, that single figure moves the annual bill by more than 2,500 PLN — and the monthly figure on the price list does not show it at all.

What really moves the bill: how often the bot hands the call over

Every handover to a person costs twice: you pay for the minutes the bot has already spent, and for the time of the employee who picks the case up. How expensive that second half gets depends on what they receive: a transcript of the conversation exists (more on that below), but until it travels to them with the call, the customer tells the whole story again. The decisive cost parameter is therefore not the price of a minute but the share of conversations the bot fails to close — and what they are handed over with.

An independent reference point for Polish exists. The Polish ASR Leaderboard (Michał Junczyk, Adam Mickiewicz University; the BIGOS dataset, NeurIPS 2024) evaluated 25 speech-recognition systems across more than four thousand recordings. The median word error rate on read speech was 14.52%; on spontaneous, conversational speech it was 32.42%. Free and commercial systems were separated by only 2.5 to 4.2 percentage points. The study is from 2024 and pre-dates the current generation of models, but the order of magnitude is where design has to start: in free conversation, roughly one word in three may be recognised as something other than what was said.

Vendor claims are measured under different conditions. ElevenLabs places Polish in the top quality tier for its Scribe v2 Realtime model, defined as an error rate of 5% or below — but that applies to clean wideband audio, not to a telephone line. A classic phone line is narrowband, sampled at 8 kHz rather than wideband, so a vendor figure describes the best case rather than a forecast for your helpline. Some SIP and mobile paths negotiate wideband codecs, so what actually reaches speech recognition is settled by a recording from your own line. The most honest definition on record is OpenAI’s: the list of languages supported by gpt-realtime is the list of languages where the word error rate is below 50%. Polish is on that list, and it is worth knowing exactly what that means.

We publish no turn-latency budgets or accuracy-versus-noise figures here. Numbers of that kind circulate on blogs with no stated measurement method, and the only credible latency is the one measured on a specific client’s line, using that client’s recordings.

Cascade or a speech-to-speech model

Architecture is a cost decision, not only a technical one. In a cascade the conversation passes through three separate components: speech recognition, a language model, and speech synthesis. In a speech-to-speech design, one model does all of it at once.

We default to the cascade, for three reasons. The Polish speech-recognition component can be swapped on its own, without rebuilding the rest, if it performs worse on the client’s recordings than on the vendor’s material. Every utterance leaves a transcript, so it is possible to show what the bot heard and what it answered on that basis. And each component can be priced and measured separately. The price for that is real: more moving parts, more contracts, and more places for something to break. A speech-to-speech model is simpler to run, but it exposes less and does not let you replace the Polish half of the puzzle.

Polish is not simply “one of the supported languages”

Voice-tool marketing talks about a hundred languages. For Polish it is worth checking the documentation, because the differences are specific and checkable:

  • Azure Voice Live — the default configuration Microsoft recommends, “automatic multilingual”, covers 15 language locales and Polish is not among them; the documentation warns that transcription quality is low for a language outside that list when no language is set explicitly. Polish is available, but through the full Azure STT locale list, not through the default. It is a quiet misconfiguration trap that only shows up in the recordings.
  • Azure Speech offers exactly three standard neural voices for pl-PL: AgnieszkaNeural, MarekNeural and ZofiaNeural. No HD variants, no multilingual voices, no speaking styles.
  • Google Cloud Text-to-Speech Chirp 3 has its HD variant generally available for pl-PL, with streaming, pace and pause control, and custom pronunciations — which matters directly for drug names, street names and surnames.
  • Deepgram lists Polish for Nova-3 (general and medical), Nova-2, Enhanced and Base, in streaming and batch. Its language page, however, covers speech recognition only — Polish support in the Aura synthesiser is not documented there, so Deepgram alone is not a complete Polish cascade.
  • Cartesia states Polish support on a dedicated language page, and ElevenLabs covers it with Flash v2.5.

On top of that comes work an English demo never surfaces and which has to be costed: declining numerals in spoken output, reading a tax identifier or an appointment reference digit by digit, spelling surnames containing ą, ę, ł, ś, ż and ź, a pronunciation lexicon for the client’s own products and streets, and text normalisation on the model side before anything reaches the synthesiser. Those are hours of work, not clicks in a console.

The lines a price list leaves out

Integrations with systems of record. The bot may not state a price, a deadline, an availability or an entitlement it has not read live from the system where that fact lives. This is not excessive caution: in Moffatt v. Air Canada (2024 BCCRT 149) the airline was held responsible for what its chatbot told a passenger, and the tribunal awarded 650.88 Canadian dollars in damages alone. Every system the bot has to read from is a separate integration to quote; the number of fields drives the mapping work inside one integration rather than the number of integrations.

Compliance. The information owed to the caller, and the requirements around recordings, are not a one-off document but part of the project. We have covered them separately — the regulatory side in our piece on AI Act readiness, and the personal-data side on the security and GDPR page. What usually has to be costed here is the wording of the opening message, the decision between recording and transcription only, retention, and getting the sub-processor contracts in order. A voice setup typically has three or four of those where a text channel has one.

Telephony. Your own Polish number is a paperwork project, not a technical one. For a pilot we usually advise against ordering one and point a SIP trunk at the number the client already has — which takes weeks out of the launch and avoids the paperwork of a number of your own. The trunk and per-minute charges remain.

Maintaining the pronunciation lexicon and the escalation thresholds. A standing line, not a rounding error.

When the honest answer is “do not buy”

If all you need is an automated receptionist that picks up out of hours, says a few scripted sentences and emails you the recording — buy a subscription from a vendor. It goes live in days, the prices are published, and nobody like us is required. It is the same principle we keep on our list of what we do not do: do not build what you can buy on subscription.

Comparing the bill against the cost of a role is tempting. The median salary for a receptionist in Poland is 5,660 PLN gross per month according to wynagrodzenia.pl, and that is the only figure we publish — we do not convert it into a total employer cost with a multiplier found online, because the multipliers in circulation do not agree even with their own arithmetic. More importantly, the comparison sets two different jobs against each other: the bot takes the repeatable calls and a person is left with the rest, so the realistic outcome is less interrupted work rather than fewer roles.

We have not delivered a voicebot yet, so we have no outcome numbers of our own — and we do not intend to borrow anyone else’s. Effectiveness claims from vendor material do exist, but they carry no stated method and are not a forecast for your project.

How to work this out for yourself

  1. Pull a three-month report from the PBX: inbound calls, unanswered calls, distribution by hour.
  2. Read the last hundred unanswered calls and split them into repeatable and unusual matters. Only the first group is in scope.
  3. Ask every vendor for four numbers in writing: the subscription, the minutes included, the overage rate and the setup fee. Separately — a quote for the integrations, with the list of systems.
  4. Sum twelve months under two scenarios: traffic inside the bundle, and traffic half again as large.
  5. Only now set the result against the number from step one. If the other side of the equation is small, the project does not make sense — and that is good news, because finding out cost one afternoon.

If you want to work out which calls are suitable for handing to an automated system in the first place, and which it must never take, we set that out on our voicebot for business page. The cost components common to all our projects are covered in our piece on what an AI implementation costs.

Frequently asked questions

What does a voicebot cost?

Published Polish price lists start at a few hundred PLN net per month and reach several thousand, but the monthly figure alone is not enough to compare offers. The bill also carries minutes above the bundle, a one-off setup fee, and a separate quote for integrating the bot with the systems it has to read. Two price lists read on 31 August 2026 come to 19,270 PLN and 5,988 PLN net over twelve months at 500 minutes a month — and most of that gap is the subscription (about 8,300 PLN), with the setup fee adding about 5,000 PLN. Both totals cover only the published list-price lines; integrations and maintenance are quoted separately and sit on top.

What is the overage rate, and why ask for it in writing?

It is the price of each minute of conversation beyond the bundle included in the subscription. In the price lists we read on 31 August 2026 it ranges from 0.73 to 1.49 PLN net per minute — roughly a factor of two. On a 500-minute bundle with real traffic of 800 minutes a month, that single rate moves the annual bill by more than 2,500 PLN. It is the line the monthly figure hides, so ask for it in writing before comparing anything else.

Will a voicebot be cheaper than hiring someone to answer the phone?

That comparison usually misleads, because it does not set the same work against itself. The median salary for a receptionist in Poland is 5,660 PLN gross per month according to wynagrodzenia.pl, and that is the only figure we publish — we do not convert it into a total employer cost using a multiplier found online. More importantly, the bot takes the repeatable calls and a person is left with the rest, so the realistic outcome is less interrupted work rather than fewer roles.

What do the setup and the integrations cost?

The setup fee in published price lists runs from zero to the low tens of thousands of PLN net, but it covers launching a ready-made platform with a conversation script — not connecting it to your systems. Integrations are quoted separately and are driven by the state of your data, because the bot may only state what it reads live from the system of record. We do not publish our own range for that part until we have seen the PBX report and the list of systems.

When is a voicebot not worth buying?

When the number of calls currently going unanswered is small, or when those calls concern unusual matters the bot would have to hand over anyway. And when all you need is an out-of-hours answering service — a subscription from a vendor covers that, and nobody like us is required. The honest order is the report from your own PBX first, then a comparison of the four lines on the bill, and only then a conversation about an implementation.