Voicebot for business — which calls it may take, and which it may not

A working voicebot can be bought on subscription, and doing so requires hiring nobody. The project starts where the bot has to be wired into the systems that hold the real data — and where somebody has to decide which calls it must not take.

What a voicebot will not solve — and why we start there

Start with the sentence that costs us work: if all you need is an answering service that picks up outside office hours, reads a few lines of script and emails you the recording, buy the subscription. Some Polish vendors publish price lists, setup takes days, and we are not needed for it. That is the same position we keep on the what we don’t do list: do not build what you can buy on a subscription.

The project starts at the second question — when the bot has to say something that is not in the script. An appointment time. An order status. Whether the customer is entitled to a refund. Every answer of that kind has to be read live from the system that knows it, not generated by a model. That is integration work, and it is what we sell: choosing a vendor on Polish-language evidence, wiring the bot to the systems of record, bounding what it may say, and putting in the disclosure and oversight the rules require.

The second limit is technical and no vendor removes it. The one independent, peer-reviewed benchmark of Polish speech recognition — the Polish ASR Leaderboard / BIGOS study, 25 systems and over 4,000 recordings — reports a median word error rate of 14.52% on read speech and 32.42% on spontaneous speech, which is the kind you hear on a phone. The study is from 2024 and does not cover the current model generation, so today’s numbers are probably better; the order of magnitude is not. And the figures vendors publish are measured on wideband audio. A classic phone line is narrowband — sampled at 8 kHz rather than wideband — and that moves those figures one way only. Not every route is narrowband: some SIP and mobile paths negotiate wideband codecs, so what speech recognition actually receives depends on the call. Which is exactly why the only benchmark that counts is recordings from the client’s own line — and everything downstream follows from it: the confirmations, the escalation thresholds, and what the bot is allowed to assert at all.

The same caution applies to claims of language support. OpenAI defines its published language list for gpt-realtime as the languages that “exceeded <50% word error rate”. Polish is on that list — and that sentence is worth reading twice before “supports Polish” is taken as an answer about quality.

What we actually do

We start with uses where a speech-recognition error costs nothing, and move down the list only once the previous step has proved its quality on real traffic:

  • Call notes into the CRM. Nobody talks to the bot at all — a person runs the call and the system transcribes it and writes the outcome into the system fields. A transcription error is seen by the salesperson before saving, so it never reaches the customer.
  • Qualifying and routing calls. The bot establishes what the call is about and transfers it to the right person or queue. It does not answer the question — a person does, only sooner and without three transfers. The cost of a mistake is one extra redirect.
  • Read-only status lookups. “Where is my delivery”, “has the invoice been paid” — the answer is a fact in a database, not a sentence a model produced. The bot reads the value and repeats it; where it is not confident it heard the reference correctly, it asks for confirmation or transfers.
  • Booking, confirming and cancelling appointments. The first step where the bot writes to a system rather than only reading from it — so the first that needs a rollback story: who undoes an entry that turns out to be wrong, how, and how the customer learns of the change.
  • Answering out-of-hours and busy-line calls. A full unsupervised conversation, which is why it is the last item on this list and not the first. We switch it on for the call types where the earlier steps showed the bot understands callers — clearly marked as a machine, with a route to a person always available.

One rule holds across the whole list: the bot does not state a price, a deadline, an availability or an entitlement it has not read live from the system that knows it. That is not theoretical caution. In Moffatt v. Air Canada (2024 BCCRT 149) the tribunal held the airline responsible for what its chatbot had said and awarded CAD 650.88 in damages. The sum is token; the principle is not.

If the same enquiries also arrive by email, chat and web form — and they usually do — the voice channel is one of several entrances to the same process. We describe the whole of it at customer service automation, which is also where the argument for starting in the back office rather than with the customer conversation lives. When the bot has to take actions across several systems instead of reading a state, the work becomes an AI agent project — this page does not repeat that comparison.

What we will not promise

  • A cold-calling voicebot. Article 398 of the Polish Electronic Communications Law prohibits using automated calling systems to send commercial information without prior consent — business recipients included — and you may not call in order to obtain that consent. The President of UKE may impose a fine of up to 3% of the previous year’s revenue or up to PLN 1,000,000, whichever is higher, and the exposure stays with the caller. Whether purely operational calls fall outside the prohibition is for the client’s lawyer to settle — we do not settle it.
  • A voicebot pretending to be human. The bot says it is a machine in the opening seconds, and a route to a person stays available throughout. We do not build scripts whose purpose is that the caller fails to notice.
  • A voicebot that assesses candidates in recruitment. Inferring emotions in the workplace is prohibited by Article 5(1)(f) of the AI Act and has been since 2 February 2025. The prohibition covers recruitment and probation, binds the deployer as well as the provider, and a candidate’s consent does not lift it.
  • A voicebot on emergency numbers. Evaluating and classifying emergency calls and prioritising dispatch is an Annex III high-risk use under the AI Act. The obligations for that category were deferred to 2 December 2027 — deferred, not cancelled — and we do not have the experience such a project would require.
  • A debt-collection voicebot as a first project. We say this plainly and on reputational grounds, not legal ones: we found no provision prohibiting automated calling for debt collection, and we will not pretend one exists. We do not, however, want the first place a company tries an automated conversation with people under pressure to be debt collection. The full set of conditions under which we advise against a project is collected at what we don’t do.

We also do not build our own speech engine and do not resell a bot subscription. The engine is bought; we answer for what it is wired to and what it is allowed to say. Deployment outcomes published by vendors we treat as marketing with no stated methodology — we do not repeat them to a client as a forecast.

Call recordings are personal data

A call recording and its transcript are as a rule ordinary personal data. The legal basis, the retention period and what the caller is told all have to be settled before go-live, not after the first month. The category is decided by what is said on the call, not by the technology. One line where that category changes does run through the technology: extracting a voiceprint in order to identify or verify the speaker is special-category biometric data under Article 9 of the GDPR and needs its own condition. By default we build on the safer side of that line — transcription with no voice profile — and we say so plainly when a client asks for something that crosses it.

The second difference from the text channel is the supplier chain. A voice integration typically adds three or four external processors — the carrier, speech recognition, the language model, speech synthesis, sometimes a separate store for recordings — where a text integration has one. Each is another link in the Article 28 chain — either a processor you contract directly, or a sub-processor you have to authorise and whose obligations are passed down that chain — and each is its own question about transfers outside the EEA. Each supplier's role is determined separately, per processing operation. How the roles are split, how that split is determined, and the questions we put to every supplier are set out on the security and GDPR page — here we add only the voice part.

Disclosing the machine — what the article requires, and what we always do

Article 50(1) of the AI Act requires that a person interacting with an AI system be informed of it clearly and at the latest at the first interaction — unless that is obvious from the circumstances and the context of use. On a phone line we do not rely on that exception: a caller has no way of knowing who will pick up, so we disclose every time. That is our delivery rule, stricter than the statutory minimum. The article applies from 2 August 2026, and it is addressed first to how the system is designed: the bot has to be built so that it discloses. Whether the disclosure actually reaches the caller then depends on how the company runs it — the script, the greeting, whatever sits in front of the bot. We build the disclosure into the script and the documentation, and we put the question of how the article allocates duties between provider and deployer in a given setup to your lawyer at the start of the project rather than settling it ourselves.

Article 50(2) adds a second duty: synthetic audio must be marked in a machine-readable way, “as far as this is technically feasible”. The transition period — until 2 December 2026 — applies to systems placed on the market before 2 August 2026, so which side of that date a given deployment falls on depends on the engine used and its provider, not on the day the integration is built. The marking duty itself is addressed to the system's provider. So there are two things to establish: what the engine's provider states, and whether that statement can be verified. We do not claim that any marking survives a telephony codec; what we deliver is a documented feasibility assessment and a record of what the system generated and when. The wider preparation — a register of AI systems, documentation, logging and human oversight — is described at the AI Act readiness audit.

Then there is the number the bot calls from or answers on. The Polish act of 28 July 2023 on combating abuse in electronic communications names caller-ID spoofing as an abuse, and since 24 September 2024 operators have been obliged to block such calls. The number is therefore a compliance artefact, not a configuration detail: a spoofed or unroutable number simply stops connecting. Separately, a Polish local number is a paperwork project — the registered address has to match the number’s region, and operators require identity and company documents. For a pilot it is usually better to point traffic at a number the client already holds.

On the Polish side the act of 3 July 2026 on artificial intelligence systems (Dz.U. 2026 poz. 1003) establishes KRiBSI — a collegiate body served by the Ministry of Digitisation and composed of representatives of the Presidents of UOKiK, KNF, KRRiT and UKE. The body exists in law and is still being stood up, so we do not claim anyone is enforcing these duties today. What we do claim is that the date the disclosure duty starts is known, and that a project going live after it has to meet it from day one.

How we work

We start with a report from the client’s own phone system: how many calls come in, at what times, how many go unanswered, and what callers ask about. We do not quote an industry missed-call figure, because every one we found traces back to vendor marketing — whereas your own report takes an afternoon to pull and is true. The same rule governs quality: recordings from the client’s line are the only benchmark we select a speech-recognition vendor on.

We choose vendors on Polish, not on the length of a language list in a brochure. One example we check every time: Azure Voice Live’s default, Microsoft-recommended “automatic multilingual” configuration covers fifteen language locales and Polish is not among them — Polish is available, but it has to be set explicitly, and Microsoft warns that transcription quality is poor when no language is set. Synthesis is similar: Azure offers exactly three standard neural voices for pl-PL, with no HD variants and no speaking styles, while Google’s Chirp 3 HD is generally available for pl-PL and supports custom pronunciations — which matters directly for product names, street names and surnames.

From there it is the same as any of our projects: one call type, the narrowest possible pilot, and the decision to widen the scope taken on data from that pilot. The full sequence is described in our delivery methodology, the whole scope of the service — from process audit to running it — at AI implementation, and the conditions a process has to meet to be worth automating at AI automation.

What it costs

On the bot side the cost has four components that vendors quote separately and that have to be summed for a comparison to mean anything: the platform subscription, a bundle of minutes with the rate for minutes beyond it, a one-off implementation fee, and the integrations with the systems of record. A fifth is running cost — scripts, the pronunciation lexicon and the escalation thresholds change whenever the offer does. The overage rate tends to be the biggest surprise, because published rates differ between vendors by roughly a factor of two.

Whether the project makes sense at all is decided by the number on the other side: how many calls go unanswered today, and how many of those concern matters the bot could handle without guessing. Where that number is small, the honest answer is “do not buy” — and we give it. We have set the whole bill out line by line, with two published price lists summed over twelve months, in our piece on what a voicebot costs. The cost components common to all our projects are covered separately, in what an AI implementation costs.

We have not delivered a voicebot yet, so we have no outcome numbers of our own — and we do not intend to borrow anyone else’s. What we can show before a project is how we work and the list of things we will not do.

Frequently asked questions

How is a voicebot different from a chatbot?

The channel, and the margin for error. A chatbot receives exactly the text the customer typed; a voicebot has to recognise speech first, and that is the step where words go missing. The one independent, peer-reviewed benchmark of Polish speech recognition (the Polish ASR Leaderboard / BIGOS study, 25 systems) reports a median word error rate of 14.52% on read speech and 32.42% on spontaneous speech. The study is from 2024 and predates the current model generation, but the order of magnitude explains why a voicebot has to confirm what it heard rather than assume it heard correctly. Telephony usually narrows it further: a classic narrowband line sampled at 8 kHz carries less than the wideband audio vendors measure on — though some SIP and mobile paths negotiate wideband codecs, so a recording from the actual line settles it.

What does a voicebot cost?

The price has four components that vendors quote separately: the platform subscription, a bundle of minutes together with the rate for minutes beyond it, a one-off implementation fee, and the integrations with the systems where the data actually lives. A fifth is running cost — scripts and the pronunciation lexicon change whenever the offer does. Some Polish vendors publish price lists, so comparison is possible, but only once every component is summed over twelve months including the overage minutes. Whether it is worth doing at all is decided by a different number: how many calls currently go unanswered. That one comes out of your own phone system in an afternoon.

Do callers have to be told they are talking to a machine?

Yes — always, in our builds. Article 50(1) of the AI Act requires that a person interacting with an AI system be informed clearly and at the latest at the first interaction, unless that is obvious from the circumstances and the context of use; it applies from 2 August 2026. On a phone line we do not rely on that exception — a caller has no way of knowing who will pick up — so we disclose every time and treat that as our own delivery rule, stricter than the statutory minimum. In practice that is a sentence in the opening seconds of the call, not a line in the terms of use. The article is addressed first to how the system is built — the bot has to be capable of disclosing — while whether the disclosure actually reaches the caller depends on the configuration at your end: the script, the greeting, and whatever sits in front of the bot. We build the disclosure into the script and the documentation; how the article allocates duties in a specific setup is a question for your lawyer.

What about call recording and the GDPR?

A recording and its transcript are as a rule ordinary personal data — the legal basis, the retention period and what the caller is told all have to be settled before go-live, not after the first month. The category is decided by what is said on the call, not by the technology. One line where that category changes does run through the technology: extracting a voiceprint in order to identify or verify the speaker is special-category biometric data under Article 9 and needs its own Article 9(2) condition — and which condition applies is for your lawyer to settle. By default we build on the safer side of that line — transcription with no voice profile. Separately, the processor chain has to be counted: a voice integration typically adds three or four external processors where a text one has one.

Will you build a voicebot that calls customers?

Not to lists that have not consented. Article 398 of the Polish Electronic Communications Law prohibits using automated calling systems to send commercial information without prior consent — business recipients included — and you may not call in order to obtain that consent. The President of UKE may impose a fine of up to 3% of the previous year’s revenue or up to PLN 1,000,000, whichever is higher, and the exposure sits with the caller, which is to say with the client. Whether purely operational calls — an appointment reminder, an order confirmation — fall outside the prohibition is a question for your lawyer, not for us.

How many calls a day go unanswered at your company?

Tell us how many there are and what they are usually about. We will say which call type we would take for a first pilot — including when the answer is “buy the subscription, you do not need us”.

Book a free consultation

Last updated: