The phone is the channel where customer patience is thinnest. A caller who waited eight minutes in a queue does not want a cheerful bot reading a script - they want the return processed, the appointment moved, the invoice explained. In 2026 a wave of AI voice platforms promises exactly that: answer every call, sound human, resolve the issue. The market has responded. AI voice agents grew from roughly $2.54 billion in 2025 to an estimated $3.51 billion in 2026, on a trajectory that forecasters put near 39 percent annual growth for the rest of the decade1. Voice AI funding jumped eightfold to $2.1 billion in a single year3.
So the tools are real and the money is real. The confusion is real too. Retell, Vapi, PolyAI, Parloa, Synthflow, NICE Cognigy, Bland AI and Salesforce Agentforce Voice all answer inbound calls, and every one of them will show you a demo where it sounds great. This guide is for the operations, service or IT leader who has to pick one and be accountable for the result. It names the tools honestly, shows where each genuinely fits, and is clear about the one thing none of them solves on its own.
That one thing is the last mile: every platform can answer and route a call, but none of them keeps how your company actually resolves a call - the answers, the exceptions, the escalation rules - and none of them owns the outcome across your real systems. That knowledge is where the value is, and it is exactly what walks out the door when your best support person leaves.
TL;DR
No platform wins every row. Retell and Vapi lead for developer-built low-latency agents, PolyAI and Parloa and NICE for large contact centres, Synthflow and Bland for no-code and volume, Agentforce Voice for Salesforce shops.
Latency is the phone-specific test. Humans swap turns in about 200 milliseconds, so sub-800 milliseconds at the 95th percentile is the floor - above it, callers interrupt and disengage.
Pricing reality is $0.07 to $0.30 per minute for build-your-own tools, and low-to-mid six figures a year for managed enterprise platforms - before integration and maintenance.
Compliance is not optional. From 2 August 2026, EU AI Act Article 50 requires you to tell callers they are speaking to an AI, and DSGVO requires consent before you record.
The durable win is a Company Brain plus an AI employee - keeping your call-handling knowledge and resolving the call across CRM, helpdesk and order systems, not just answering the phone.
What an Inbound Voice Agent Actually Has to Do
Answering the phone is the visible part. The work that determines whether a voice agent is worth deploying happens underneath the conversation, and most comparisons never test it. Before you compare vendors, be clear about the job.
- Answer instantly and sound natural - pick up on the first ring, speak with human-like timing, and never leave the dead air that makes callers say “hello? hello?”
- Understand messy, real speech - accents, background noise, half-sentences, interruptions, and callers who change their mind mid-request.
- Know your specifics - your return window, your warranty terms, which exceptions your team makes and when, and the answers that only exist in a senior agent’s head.
- Act across systems - look up the order in the ERP, check the ticket in the helpdesk, update the record in the CRM, and trigger the refund or the reschedule - not just read a value aloud.
- Escalate cleanly - recognise when it is out of its depth, hand off to a human with the full context, and never trap the caller in a loop.
- Stay compliant on the line - disclose that it is an AI, capture recording consent where required, and keep an auditable trail.
- Stay accurate as you change - when a policy or product changes, the agent’s answers change with it, without a re-build.
The 80/20 trap
Answering the call and holding a conversation is the easy 80 percent, and it is what every demo shows. The hard 20 percent - resolving the request across your systems, handling the exceptions your best people know by heart, and staying accurate as your business changes - is where value lives and where most pilots quietly die. Judge platforms on the 20 percent, not the demo.
| Layer of the job | What it covers | Who owns it today |
|---|---|---|
| The audio layer | Speech-to-text, voice, telephony, turn-taking | The voice platform |
| The conversation layer | Understanding intent, dialogue flow, tone | The voice platform |
| The knowledge layer | Your answers, exceptions and escalation rules | Usually a person, not a system |
| The action layer | Reading and writing across CRM, helpdesk, ERP | You build and maintain it |
| The compliance layer | AI disclosure, recording consent, audit trail | You, the deployer |
The audio and conversation layers are where the platforms compete. The knowledge, action and compliance layers are where deployments succeed or fail - and they are mostly left to you.
The Platforms, Honestly Compared
Each of these tools is genuinely good at something. The mistake is treating them as interchangeable. Here is where each one actually fits, without the marketing gloss.
Retell AI
- Best for - the fastest path to a managed, low-latency inbound phone agent without assembling the whole stack yourself.
- Strengths - roughly 600 milliseconds of latency with little tuning, and strong appointment and calendar handling that connects to Cal.com, Google Calendar and custom calendar APIs with real-time availability checks during the call9.
- Pricing - around $0.07 per minute managed, all-in typically $0.08 to $0.15 once providers are added8.
- Watch for - deeper integrations and complex logic still require engineering, and you own the resolution logic behind the call.
Vapi
- Best for - engineering teams that want full control over the language model, speech-to-text, text-to-speech and telephony.
- Strengths - a tuned stack reaches 500 to 700 milliseconds, and you can swap any component to optimise cost or quality7.
- Pricing - a $0.05 per minute platform fee with provider costs at cost, so all-in commonly lands at $0.10 to $0.30 per minute11.
- Watch for - flexibility is a cost: you assemble, tune and maintain more of the system, and the last-mile integration is entirely yours.
Bland AI
- Best for - high-volume calling with predictable per-minute pricing on self-hosted infrastructure.
- Strengths - bundled, all-inclusive pricing around $0.07 to $0.12 per minute and an emphasis on scale10.
- Pricing - simple per-minute at volume, which makes cost forecasting easy.
- Watch for - its centre of gravity is outbound and volume; for nuanced inbound resolution you still supply the knowledge and the integrations.
Synthflow
- Best for - no-code teams that want an agent live quickly without engineering.
- Strengths - a visual builder, about $0.09 per minute for the voice engine, and a clear path from pay-as-you-go to enterprise12.
- Pricing - usage-based to start; enterprise contracts begin around $30,000 a year scoped by volume, concurrency and integrations12.
- Watch for - no-code accelerates the easy part; exceptions, deep integration and compliance still need real ownership.
PolyAI
- Best for - large enterprises where voice volume drives most of the customer experience and quality on the line is paramount.
- Strengths - a voice-specialised platform with careful live-caller interaction design, delivered as a managed enterprise programme13.
- Pricing - custom enterprise, typically six-figure annual contracts, often cited around a $150,000 per year floor for real deployments14.
- Watch for - the price buys managed voice quality; resolution logic across your back-end systems is still a scoped integration.
Parloa
- Best for - enterprise contact centres operating under regulatory pressure and at scale, especially in Europe.
- Strengths - a voice-first agent-management platform covering the full lifecycle across design, test, scale and optimise, with 130-plus languages and customers such as Allianz, Booking.com, SAP and Swiss Life16. The Berlin-founded company raised a $350 million Series D in January 2026 at a $3 billion valuation16.
- Pricing - custom enterprise, scoped to volume and deployment.
- Watch for - built for large, complex operations; overkill for a mid-sized team that needs one workflow handled.
NICE (Cognigy)
- Best for - enterprises standardising customer experience on a full contact-centre platform.
- Strengths - NICE acquired the German conversational-AI company Cognigy in a $955 million deal that closed in September 2025, uniting Cognigy’s agentic AI with the CXone platform1920.
- Pricing - Cognigy contracts have started around $2,500 a month, with full-language enterprise deployments exceeding $300,000 a year9.
- Watch for - most valuable when you adopt the wider suite; a large commitment for a single use case.
Salesforce Agentforce Voice
- Best for - companies already standardised on Salesforce that want voice native to their CRM and data.
- Strengths - voice actions inside the Agentforce framework, grounded in Salesforce data, with a pay-per-resolution option that charges $2 when the Help Agent resolves an issue autonomously2122.
- Pricing - Flex Credits at roughly $500 per 100,000 credits, voice actions costing 30 credits each, and $2 per conversation depending on model21.
- Watch for - the value is highest inside the Salesforce ecosystem; outside it, the fit weakens.
Generic assistants (ChatGPT, Gemini, Claude) as a baseline
- Best for - understanding what the underlying models can do, and as a baseline none of the platforms should fall below.
- Strengths - excellent language understanding and generation, and increasingly capable real-time voice modes.
- Pricing - low per-token cost, but no telephony, no call routing, no CRM write-back out of the box.
- Watch for - a general assistant is not an inbound call system; it has no line, no consent capture, no company knowledge, and no ownership of the outcome.
| Platform | Primary fit | Deployment model | Pricing shape |
|---|---|---|---|
| Retell AI | Managed low-latency inbound | Developer + managed | ~$0.07-0.15/min |
| Vapi | Full-control developer stack | Developer / API | $0.05/min + providers |
| Bland AI | High-volume calling | Self-hosted infra | ~$0.07-0.12/min all-in |
| Synthflow | No-code speed | No-code builder | ~$0.09/min; ent. from $30k/yr |
| PolyAI | Voice-first enterprise CX | Managed enterprise | Custom, ~$150k/yr+ |
| Parloa | Regulated enterprise at scale | Managed enterprise | Custom enterprise |
| NICE (Cognigy) | Full contact-centre suite | Managed enterprise | ~$2.5k/mo to $300k/yr+ |
| Agentforce Voice | Salesforce-native | Platform add-on | Flex credits / $2 per resolution |
Latency and Conversational Robustness on the Phone
Voice is unforgiving in a way that chat is not. In text, a two-second pause is invisible. On a call, it is a broken conversation. This is the one dimension where phone agents differ most from their chat cousins, and where many otherwise-capable systems fall down.
Why latency is the hard constraint
- Human turn-taking is fast - people hand off conversational turns in roughly 200 milliseconds, faster than it takes to produce even a one-word reply5.
- The natural zone is 200 to 500 milliseconds - inside this range the agent feels human5.
- 500 to 800 milliseconds is noticeable but acceptable - callers tolerate it, but the magic fades5.
- Above one second, it feels broken - callers interrupt, talk over the agent, or disengage entirely6.
- Sub-800 milliseconds at p95 is the floor - not the average, the 95th percentile, because the slow calls are the ones that fail5.
Latency is a pipeline, not a number
A stitched voice pipeline spends roughly 100 to 300 milliseconds on speech-to-text, 350 to 1,000 milliseconds on language-model inference, 90 to 200 milliseconds on text-to-speech, and another 50 to 200 milliseconds on network hops between vendors4. The language model is usually the biggest and most variable slice. Every vendor boundary you add is another place latency accumulates.
| Response latency | How it feels | Caller behaviour |
|---|---|---|
| 200-500 ms | Natural, human | Normal conversation |
| 500-800 ms | Noticeable but acceptable | Slight hesitation, still engaged |
| 800-1,500 ms | Laggy | Interrupts, repeats, talks over |
| Above 1,500 ms | Broken | Disengages, asks for a human |
Robustness is more than speed
A fast agent that collapses on the first interruption is not robust. Real calls are messy, and the failure modes are specific.
- Barge-in handling - the caller interrupts and the agent must stop talking and listen, immediately.
- Noise and accents - a call from a factory floor or a moving car is not a studio recording.
- Disfluency - real people say “um”, restart sentences, and bury the request in a story.
- Numbers and spellings - order numbers, postcodes and names are where transcription errors cost the most.
- Graceful uncertainty - when the agent is not sure, it should confirm rather than guess, and escalate rather than loop.
Developer Stacks vs Managed Platforms on Latency
Developer stacks (Vapi, Retell)
- ✓ Tunable - swap models and voices to chase the lowest latency
- ✓ Transparent - you can measure each stage of the pipeline
- ✓ Fast to a prototype - a working agent in days
- ✗ You own the tuning - p95 latency in production is your problem
- ✗ Robustness is DIY - barge-in and noise handling need work
Managed platforms (PolyAI, Parloa, NICE)
- ✓ Tuned for voice quality - robustness is the product
- ✓ Built for scale - designed for high concurrency and uptime
- ✓ Support included - vendor and partner behind the deployment
- ✗ Less control - you tune within their framework
- ✗ Higher cost - six-figure commitments are common
“Agentic AI has emerged as a game-changer for customer service, paving the way for autonomous and low-effort customer experiences.”
- Daniel O’Sullivan, Senior Director Analyst at Gartner23
Want a voice agent that actually resolves the call?
Book a 30-minute call. We will map one inbound workflow end to end, from the line to your systems.

CRM, Telephony and Deployment Reality
A voice agent that cannot touch your systems is an expensive answering machine. The real difference between platforms is not whether they can read a record, but whether they can reliably write back and take action across several systems to close the call.
Telephony: how the call reaches the agent
- Bring-your-own carrier - developer platforms connect through SIP or providers like Twilio and Telnyx, giving you control and portability.
- Native contact-centre integration - enterprise platforms plug into existing contact-centre infrastructure so the agent lives alongside human agents.
- Number provisioning - some platforms provision numbers directly; others expect you to bring them.
- Warm transfer - the ability to hand a live call to a human with context is a production requirement, not a nice-to-have.
CRM and back-end: where resolution happens
- Read access - almost every platform can look up a customer or order to read a value aloud.
- Write access - far fewer can reliably update a record, create a ticket, or trigger a refund - and this is where calls actually get resolved.
- Multi-system orchestration - a real resolution often touches CRM, helpdesk and an order or ERP system in one call.
- Native vs API - Agentforce Voice is native to Salesforce; developer platforms integrate through APIs and webhooks you build and maintain.
- Idempotency and error handling - what happens when the write fails halfway is the difference between a resolved call and a double refund.
| Integration need | Developer platforms | Enterprise platforms | Salesforce-native |
|---|---|---|---|
| Telephony | SIP / BYO carrier | Native CC integration | Via Salesforce / partners |
| CRM read | API (you build) | Pre-built connectors | Native |
| CRM write-back | You build and own it | Scoped integration | Native |
| Multi-system actions | Your responsibility | Professional services | Within Salesforce mostly |
| Time to production | Days to weeks | Weeks to months | Weeks (if on Salesforce) |
The question that separates demos from deployments
Ask every vendor: “When the agent resolves this call, which systems does it write to, and what happens when one of those writes fails?” If the answer is vague, or if it turns out you are building and owning all of that yourself, you are buying an audio layer, not a resolution.
Pricing Reality: What You Actually Pay
The per-minute number on the pricing page is the smallest part of the real cost. Two platforms with the same per-minute rate can differ by an order of magnitude once you account for what you build and maintain around them.
The three cost layers
- The conversation cost - the per-minute rate for speech, model and voice. This is $0.05 to $0.30 per minute for build-your-own tools, and folded into the contract for enterprise platforms11.
- The build and integration cost - connecting to your telephony, CRM and back-end systems, and building the resolution logic. This is usually the largest line and it is rarely on the pricing page.
- The maintenance cost - keeping answers accurate as products and policies change, monitoring quality, and handling the edge cases production surfaces.
| Platform | Headline price | What is not included |
|---|---|---|
| Vapi | $0.05/min platform fee | Provider costs, all integration, tuning |
| Retell AI | ~$0.07/min | Deep integrations, resolution logic |
| Synthflow | ~$0.09/min; ent. from $30k/yr | Complex logic, back-end write-back |
| PolyAI | Custom, ~$150k/yr+ | Back-end resolution integration |
| NICE (Cognigy) | ~$2.5k/mo to $300k/yr+ | Best value only with the wider suite |
| Agentforce Voice | $2 per resolution / flex credits | Value tied to Salesforce footprint |
Outcome pricing only works if something owns the outcome
Salesforce’s $2-per-resolution model and the wider shift to outcome pricing are only meaningful when an agent actually completes an outcome end to end. A system that answers and routes but does not resolve cannot honestly be priced per resolution - because it never finishes one. Insist that “resolution” means the caller’s request is closed in your systems, not that the call was handled.
The Compliance Line Most Comparisons Skip
Voice comparisons love latency charts and hate compliance. But an inbound agent talking to EU callers has two hard legal obligations that most buyer guides ignore, and both carry real penalties.
EU AI Act Article 50: callers must be told it is AI
- The rule - providers of AI systems that interact directly with people must ensure the person is informed they are dealing with an AI, no later than the first interaction25.
- The deadline - Article 50 transparency obligations apply from 2 August 202624.
- The exception is narrow - only where it is already obvious to a reasonably observant person that they are talking to a machine25.
- The reach is global - it applies to any provider whose agent reaches EU callers, not only EU-based companies26.
- The penalty - up to 15 million euros or 3 percent of worldwide annual turnover, whichever is higher24.
In practice, compliance is simple: a short spoken disclosure at the start of the call. The failure is not technical, it is forgetting to design it in.
DSGVO and call recording: consent is not optional
- Recording needs consent - in Germany, recording a customer call generally requires the caller’s active consent, because the DSGVO and Bundesdatenschutzgesetz do not justify routine recording on their own27.
- Consent must be active - collected through a clear announcement and a confirming action, not buried in terms28.
- Transcription counts too - a transcript is still processing of personal data and needs its own lawful basis.
- The TDDDG can apply - where the voice service counts as a telecommunications service, additional rules apply.
- Use a processor, sign the paperwork - an external voice provider processing calls on your behalf requires a data-processing agreement (Auftragsverarbeitungsvertrag)27.
Where your calls are processed matters
Many voice platforms route speech through infrastructure outside the EU. For DSGVO, and for customers who ask, you need to know where audio is processed, where transcripts are stored, and who can access them. Ask for the data-flow diagram before you sign, not after your first complaint.
Inbound Voice Compliance Checklist
- A spoken AI disclosure plays at the start of every call
- Recording consent is captured actively before any recording
- Transcripts have a defined lawful basis and retention period
- You know where audio and transcripts are processed and stored
- A data-processing agreement is signed with every voice provider
- Escalation to a human is always available on request
- An audit trail records what the agent said and did
- Staff who manage the agent have basic AI-literacy training
The Gap Every Platform Shares
Line up all eight platforms and one weakness is common to every column: they run the conversation, but they do not keep how your company resolves a call, and they do not own the outcome across your systems. That is not a flaw in any one product - it is the boundary of the category.
- The knowledge is not theirs to keep - the return exceptions, the goodwill rules, the “we always do X for customer segment Y” - that lives in a senior agent’s head, not in the platform.
- The knowledge walks out the door - when your best support person leaves, the platform still answers the phone, but the judgement that resolved hard calls is gone.
- The last mile is left to you - reading a record is easy; reliably writing back across CRM, helpdesk and ERP to actually close the request is the part no platform owns.
- Every tool re-learns from zero - swap platforms and you rebuild the prompts, the flows and the integrations, because none of it was durable company memory.
- Accuracy decays - policies change, and without a living memory that updates, the agent slowly starts giving last year’s answers.
The durable asset is not the voice - it is the knowledge behind it
The voice, the model and the telephony are increasingly commodities you can swap. What compounds in value is a Company Brain: a living memory of how your company actually handles calls - the answers, the exceptions, the escalation rules - that survives staff turnover and that any voice layer can draw on. Own that, and the platform underneath becomes a replaceable part.
What closes the gap
- A Company Brain - your call-handling knowledge captured as living memory that updates from daily work and survives turnover, not a wiki that decays.
- An AI employee - not just a voice on the line, but a worker that resolves the call end to end across your real systems and escalates the rest with context.
- Real write access - the ability to act, not only to read, with proper error handling when a write fails.
- A feedback loop - every correction from your team makes the next call better, so the system improves at your company specifically.
“We believe in the future, everything will be conversational - you will talk to a homepage, you will talk to an app.”
- Malte Kosub, Co-founder and CEO of Parloa18
How to Choose: A Buyer Framework
Skip the feature matrix as your first filter. Start from your situation, then let it point you at the right two or three platforms to trial.
- Start from your call volume - a few hundred calls a month points to no-code or developer tools; hundreds of thousands points to enterprise platforms built for concurrency.
- Map your systems first - list the CRM, helpdesk and order or ERP systems a real resolution touches, because that determines integration effort more than the voice layer does.
- Define what “resolved” means - write down, for your top five call types, exactly what has to happen in your systems for the call to be done.
- Test on the hard 20 percent - trial vendors on your messiest real calls and exceptions, not the clean demo script.
- Measure p95 latency, not the average - insist on 95th-percentile latency under load, because the slow calls are the ones that fail.
- Check the compliance design - confirm AI disclosure, consent capture, data location and a signed processing agreement before you commit.
- Ask who keeps the knowledge - if the answer is “your prompts and one employee”, plan for how that knowledge becomes durable company memory.
- Price the whole thing - add build, integration and maintenance to the per-minute rate before you compare.
Vendor Trial Checklist
- It handled our five hardest real call types, not the demo
- p95 latency under load stayed under 800 milliseconds
- It wrote back correctly to every system a resolution needs
- It failed safely when a system was down or a write errored
- It disclosed it was an AI and captured consent correctly
- We know where audio and transcripts are processed
- We have a plan to keep call-handling knowledge in-house
- The all-in cost, not the per-minute rate, fits the business case
Build-Your-Own vs Managed Enterprise vs AI Employee
Build-your-own (Vapi, Retell, Synthflow)
- ✓ Fast and cheap to start - a working agent in days
- ✓ Full control - tune every part of the stack
- ✗ You own the last mile - integration and resolution are yours
- ✗ Knowledge stays fragile - it lives in prompts and people
Managed enterprise (PolyAI, Parloa, NICE)
- ✓ Voice quality and scale - robustness is the product
- ✓ Support and governance - built for regulated operations
- ✗ Six-figure commitment - heavy for a single use case
- ✗ Resolution still scoped - back-end integration is extra
How Superkind Fits
Superkind is not another voice pipeline. It builds an AI employee that owns the outcome behind the call and grounds it in a Company Brain that keeps how your company actually handles calls. For the audio layer, it can sit on top of a best-in-class voice platform - the point is what happens after the caller speaks.
- Company Brain - your call-handling knowledge - answers, exceptions, escalation rules - captured as living memory that survives when a support lead leaves.
- Resolves end to end - not just answering and routing, but taking the actions across CRM, helpdesk and order or ERP that actually close the request.
- Real write access - it acts in your systems with proper error handling, not just reads a value aloud.
- Sits on your stack - connects to your existing telephony, CRM and back-end rather than replacing them.
- Voice-platform agnostic - use the audio layer that fits; the resolution and knowledge do not move when you swap it.
- Learns from your team - every correction feeds the Company Brain, so calls get handled more like your best agent over time.
- Escalates with context - hands hard calls to a human with the full picture, never trapping the caller in a loop.
- Compliance built in - AI disclosure, consent capture and an audit trail designed into the workflow from the start.
- Outcome-based - priced against resolved calls, which only works because the AI employee actually completes them.
| Dimension | Voice agent platform | Superkind AI employee |
|---|---|---|
| Answers the call | Yes | Yes (on your chosen voice layer) |
| Resolves across systems | You build it | Owns it end to end |
| Keeps call-handling knowledge | No - lives in prompts and people | Yes - in a Company Brain |
| Survives staff turnover | Knowledge leaves with people | Knowledge stays in the company |
| Improves over time | Manual re-prompting | Feedback loop from your team |
| Switching cost | Rebuild everything | Swap the voice layer, keep the brain |
Superkind
Pros
- ✓ Owns the outcome - resolves the call, not just answers it
- ✓ Durable knowledge - a Company Brain that survives turnover
- ✓ No platform lock-in - voice layer stays swappable
- ✓ Outcome-based pricing - pay for resolved calls
- ✓ Compliance designed in - disclosure and consent from day one
Cons
- ✗ Not a self-serve tool - it needs engagement with our team
- ✗ Not for a one-off IVR - overkill if you just need a phone tree
- ✗ Needs process access - we must understand how you really handle calls
- ✗ Capacity-limited - we work with a focused number of clients
Decision Framework: Which Path Is Right for You
Different situations point to different answers. Use these signals to decide where to start.
| Your situation | What it points to | Where to start |
|---|---|---|
| You need a simple agent live this month | Speed over depth | Synthflow or managed Retell |
| You have engineers and want full control | Tunable developer stack | Vapi or Retell |
| You run a large, regulated contact centre | Managed voice quality and scale | PolyAI, Parloa or NICE |
| You are standardised on Salesforce | Native CRM integration | Agentforce Voice |
| The call must be resolved across your systems | Ownership of the outcome | An AI employee on a Company Brain |
| Your call-handling knowledge sits in one person | Key-person risk | Capture it in a Company Brain first |
Answer-and-Route vs Resolve-and-Remember
Answer-and-route (any platform)
- ✓ Deflects volume - handles the simple, repetitive calls
- ✓ Fast to deploy - value in weeks
- ✗ Stops at the hand-off - a human still finishes the work
- ✗ No memory - knowledge is not retained or reused
Resolve-and-remember (AI employee)
- ✓ Closes the call - acts across your real systems
- ✓ Keeps the knowledge - survives turnover in a Company Brain
- ✓ Compounds - improves with every corrected call
- ✗ Needs process access - more setup than a demo bot
Frequently Asked Questions
There is no single winner, because the platforms optimise for different buyers. Retell AI and Vapi lead for developer-built, low-latency phone agents. PolyAI, Parloa and NICE Cognigy lead for large enterprise contact centres. Synthflow and Bland AI sit in the no-code and high-volume middle. Salesforce Agentforce Voice fits companies already standardised on Salesforce. The right choice depends on your call volume, your existing telephony and CRM, your compliance needs, and whether you need the agent to actually resolve calls end to end rather than just answer and route them.
All-in costs for inbound calls typically land between $0.07 and $0.30 per minute once you add speech-to-text, the language model, text-to-speech and telephony. Vapi advertises a $0.05 per minute platform fee but you pay provider costs on top. Retell is around $0.07 per minute managed. Synthflow charges about $0.09 per minute for its voice engine. Enterprise platforms like PolyAI, Parloa and NICE quote custom contracts that often start in the low-to-mid six figures per year. The per-minute number is only part of the picture once you add build, integration and maintenance.
Humans hand off conversational turns in roughly 200 milliseconds, so voice agents feel natural between 200 and 500 milliseconds of response latency, acceptable between 500 and 800 milliseconds, and noticeably broken above one second. The industry treats sub-800 milliseconds at the 95th percentile as the reliability floor. Above it, callers start to interrupt, talk over the agent, or disengage. Latency is a full-pipeline problem: speech-to-text, language-model inference, text-to-speech and network hops between vendors all add up.
Yes, in the EU. Article 50 of the EU AI Act requires that people interacting with an AI system are informed they are dealing with an AI, no later than the first interaction, unless it is already obvious. The transparency obligations apply from 2 August 2026, with fines up to 15 million euros or 3 percent of global annual turnover. The rule applies to any provider whose voice agent reaches EU callers, not only EU-headquartered companies. A short spoken disclosure at the start of the call is the standard way to comply.
Recording a customer call in Germany generally requires the active consent of the caller, because the general legal bases in the DSGVO and the Bundesdatenschutzgesetz do not justify routine recording on their own. Consent is usually collected through a spoken announcement and a confirming action at the start of the call. If the voice service counts as a telecommunications service, the TDDDG adds further rules, and using an external provider requires a data-processing agreement (Auftragsverarbeitungsvertrag). Recording, transcription and storage each need a lawful basis.
A voice agent platform gives you the pipeline to answer and route calls: speech recognition, a language model, a voice, and telephony. It handles the conversation on the line. An AI employee owns the outcome behind the call - it acts across your CRM, helpdesk and order or ERP systems to actually resolve the request, follows your company-specific rules and exceptions, and keeps that knowledge in a Company Brain so it survives staff turnover. Most platforms answer the phone well but stop at the last mile where the work actually gets done.
Retell AI is usually the faster path to a managed, low-latency inbound agent, with strong appointment and calendar handling and roughly 600 milliseconds of latency without tuning. Vapi gives engineering teams full control over the language model, speech and telephony stack, which a tuned team can push to 500 to 700 milliseconds, but it expects you to assemble and maintain more of the system yourself. Choose Retell for speed to production with less engineering, and Vapi for maximum control when you have the developers to run it.
PolyAI, Parloa and NICE (which acquired Cognigy in 2025) are the platforms built for large, regulated contact centres. They offer managed deployment, extensive language coverage, agent-lifecycle tooling and deep integration with contact-centre infrastructure. Salesforce Agentforce Voice is a strong option for companies already standardised on Salesforce. These carry higher price tags, often six figures a year, but include the governance, uptime and support that large operations require.
Most can, but the depth varies. Developer platforms like Vapi and Retell connect through APIs and webhooks, which means integration is possible but you build and maintain it. Enterprise platforms ship pre-built connectors to major contact-centre and CRM systems. Salesforce Agentforce Voice is native to Salesforce. The important question is not whether it can read your CRM, but whether it can write back reliably and take actions across several systems to close the call, rather than just fetching a record to read out.
It depends on the platform. Synthflow, Bland AI and the managed side of Retell are designed to get a basic agent live with little or no code. Vapi and the deeper Retell configurations expect engineering involvement. Enterprise platforms like PolyAI, Parloa and NICE are delivered as managed programmes with vendor and integration-partner support. A no-code agent can answer calls quickly, but production-grade handling of exceptions, integrations and compliance usually needs technical ownership somewhere.
Most pilots demonstrate that the agent can answer the phone and hold a conversation, which is the easy 80 percent. They stall on the last mile: handling the exceptions your best agents know by heart, writing back to several systems reliably, meeting compliance requirements, and staying accurate as your products and policies change. When the call-handling knowledge lives only in the pilot vendor or in one employee, the agent cannot be trusted with real resolutions, so it stays a demo. Durable deployments capture that knowledge in a Company Brain and give the agent real write access.
Superkind is not a voice pipeline vendor. It builds an AI employee that resolves the call end to end and grounds it in a Company Brain that keeps how your company actually handles calls - the answers, exceptions and escalation rules - so the knowledge stays when a support lead leaves. It connects to your telephony, CRM, helpdesk and order or ERP systems and owns the outcome, not just the conversation. It can sit on top of a best-in-class voice platform for the audio layer while doing the resolution and integration work those platforms leave to you.
Related Articles
- The Best AI Customer Support Agents in 2026: Decagon, Sierra, Intercom Fin, and When to Build Your Own
- The Last-Mile Problem: Why AI Pilots Impress in the Demo and Die at Handoff
- Why 400,000 Copilot Agents Still Do Not Know Your Company
- Agent Washing: How to Tell a Real AI Employee From a Rebranded Chatbot Before You Buy
- The Integration Tax: Why AI Value Lives in the Connectors, Not the Model
Sources
- Grand View Research - AI Voice Agents Market Size And Share Report, 2026-2033
- Market.us - Voice AI Agents Market Size, Share (CAGR 34.8%)
- CloudTalk - AI Voice Agent Statistics 2026: Market Share & Use Cases
- Telnyx - Voice AI Agents Compared on Latency: Performance Benchmark
- Prodinit - Real-Time Voice AI Latency: Acceptable Ranges
- Hamming AI - Voice AI Latency: What Is Fast, What Is Slow
- Vapi AI - Speech Latency Solutions: Guide to Sub-500ms Voice AI
- Retell AI - Vapi AI Review 2026: Pricing, Features & Alternative
- Retell AI - 8 Best Voice AI Agent Companies for Contact Centers (2026)
- Builts AI - VAPI vs Bland AI vs Retell vs Synthflow (2026)
- AInora - AI Voice Agent Cost per Minute (2026)
- Synthflow - Enterprise Voice AI Pricing
- Synthflow - Honest PolyAI Review 2026: Pros, Cons, Features & Pricing
- Nurix - PolyAI Pricing 2026 Breakdown for Enterprise Voice AI
- Fin.ai - 20 Best AI Voice Agents for Phone Support Automation (2026)
- SiliconANGLE - Parloa Raises $350M to Make Enterprise CX Conversational
- Parloa - Six Months an AI Unicorn, Surpasses $50M Revenue Mark
- Dealroom - Parloa CEO Malte Kosub on Conversational Interfaces
- Aragon Research - NICE Acquires Cognigy at a 25x Premium
- NiCE - Closes Acquisition of Cognigy (Press Release)
- Salesforce - Agentforce Pricing
- Ksolves - Agentforce Pricing 2026: What Changed, What It Costs
- Gartner - Agentic AI Will Autonomously Resolve 80% of Customer Service Issues by 2029
- Stibbe - The AI Act’s Transparency Obligations: Rules, Scope and Timeline
- AI Act - Article 50: Chatbot, Deepfake and AI Content Labels
- Bratby Law - AI Act Transparency Obligations, 2 August 2026
- Dr. Datenschutz - Der Datenschutz bei der Aufzeichnung von Telefongespraechen
- Keyed - Aufzeichnung von Anrufen nach DSGVO: Was ist erlaubt?
Ready to move from answering calls to resolving them?
Book a 30-minute call with Henri. We will pick one inbound call type and map it end to end - from the line to your systems - with no commitment and no sales pitch.
Book a Demo →
