Telephony Stack for India AI Calling — Exotel vs. Twilio vs. Plivo vs. Tata Tele Compared
A complete architecture-level comparison of the four major telephony providers for Indian AI Calling deployments — audio latency, concurrency and scale, number quality and spam avoidance, DND/NCPR regulatory compliance, and cost per minute, plus a staged provider recommendation from 0 to 50,000+ calls/month and answering machine detection configuration.
⏱ 13 min read🏢 Technical Architecture & Developer Guides📅 6 July 2026
Nothing Else Matters if the Telephony Layer Drops Calls
The telephony stack is the infrastructure layer that physically connects an AI Calling Agent to the Indian PSTN. Every ASR engine selection, every TTS latency optimization, every prompt engineering refinement — none of it matters if the telephony layer introduces call drops, audio degradation, number blocking, or concurrency limits that cap the system's ability to operate at scale.
In the Indian real estate context, the telephony stack must satisfy five operational constraints simultaneously: PSTN audio quality at 8kHz/G.711 with under 150ms one-way latency, SIP trunk concurrency sufficient for outbound burst calling, a clean number pool that does not trigger carrier spam flags, regulatory compliance for commercial outbound calling, and cost-per-minute economics compatible with AI Calling's usage-based pricing model. This article compares Exotel, Twilio, Plivo, and Tata Tele Business Services across all five dimensions.
Architecture: Where Telephony Sits in the AI Calling Stack
The telephony provider is the interface between the AI Calling Agent's software layer and the physical PSTN. In an AI Calling deployment, the telephony stack handles four distinct functions.
Outbound call initiation — accepts an API call with destination number and caller ID, routes the call to PSTN, connects the buyer's phone, and notifies the AI orchestration layer via webhook when connected
Audio transport — transmits buyer audio to the AI pipeline and streams AI-generated TTS audio back to the buyer in real time via WebSocket/RTP
Call control — pause/resume, transfer to human agent, DTMF detection, call recording, voicemail detection
Number management — virtual number provisioning, caller ID assignment, DND scrubbing integration, call detail record (CDR) generation for CRM sync
The telephony layer is stateless from the AI perspective — the AI pipeline receives an audio stream and sends audio back; the telephony provider manages PSTN routing complexity. But the quality and architecture of the telephony layer determines whether the AI pipeline operates on clean, low-latency audio or degraded, high-jitter audio that cripples ASR accuracy.
The Five Evaluation Dimensions
Dimension 1: Audio Latency and Quality
One-way PSTN audio latency from buyer speech to AI pipeline receipt should be under 100ms; round-trip latency should be under 800ms total. Each additional transcoding stage in the audio path adds 15–40ms latency and introduces codec artifacts that raise ASR Word Error Rate. A telephony provider that routes Indian calls through a US PoP before reaching the AI pipeline adds 180–250ms of round-trip latency — enough to make conversation feel unnatural.
Dimension 2: Concurrency and Scale
Real estate developers running portal campaigns generate lead spikes — 200–500 new leads from a weekend Meta campaign drop require immediate outreach within minutes. The telephony stack must support concurrent outbound call bursts of 50–200 simultaneous calls without degradation. Some providers impose hard concurrency limits at the account tier level; provisioning higher concurrency requires advance notice, contracts, or enterprise tier upgrades.
Dimension 3: Number Quality and Spam Avoidance
AI Calling at scale generates call volume patterns that carrier spam detection systems flag as automated calling — high call frequency from a single number, short call durations, identical call attempt patterns. Spam-flagged numbers display as "Spam Risk" on buyer phones, triggering immediate hang-ups. Providers with limited Indian number pools cannot rotate caller IDs frequently enough to avoid spam flagging, and providers that share number pools across customers inherit spam reputation from other customers' calling patterns.
Dimension 4: Regulatory Compliance
Commercial outbound calling in India requires compliance with the Telecom Commercial Communications Customer Preference Regulations (TCCCPR) framework managed by TRAI. The telephony provider must support scrubbing of call attempts against the National Customer Preference Register (NCPR) — formerly NDNC — to avoid calling registered DND numbers. DND scrubbing compliance is a carrier-level obligation of the real estate developer deploying the AI Calling system; the telephony provider's role is to support the API hooks needed for DND-scrubbed number lists to be passed through the dialing system.
Dimension 5: Cost per Minute
At 10,000 call-minutes/month, the difference between ₹0.50/minute and ₹1.20/minute is ₹7,000/month — ₹84,000/year. At 100,000 call-minutes/month, the difference scales to ₹70,000/month. Telephony cost is the largest per-call infrastructure cost after LLM inference and must be optimized relative to call quality trade-offs.
Provider Analysis
Exotel
India's largest cloud telephony provider by Indian enterprise customer count, built exclusively for Indian PSTN connectivity with PoPs in Mumbai, Bengaluru, and Delhi NCR. Audio latency: 40–80ms one-way, best of evaluated providers for India-to-India calls. Concurrency: 50 default, 200–500 enterprise. Number pool: 10,000+ Indian virtual numbers with proactive spam-flag retirement. Native NCPR/DND scrubbing built in. Pricing: ₹0.50–₹0.70/minute. Weaknesses: limited international number support, less mature documentation than Twilio, no global PoP network. Optimal use case: any Indian domestic real estate AI Calling deployment; primary recommendation for NCR, Mumbai, Hyderabad, Bangalore, Pune markets.
Twilio
The global cloud communications leader with the most extensive developer ecosystem. Audio latency: 80–160ms one-way — routes Indian PSTN calls through Singapore PoP, adding 30–60ms vs. India-domestic routing. Concurrency: uncapped by default. Number pool: inconsistent spam reputation, customer-managed health monitoring. Streaming API (Media Streams) is the most mature bidirectional WebSocket API in the industry, with native integrations across most AI voice frameworks. No native DND/NCPR scrubbing. Pricing: ~₹1.08/minute, 40–50% higher than Exotel. Optimal use case: NRI call segments, international number provisioning, or when Twilio's ecosystem integrations justify the cost premium.
Plivo
A US-headquartered provider with strong Indian infrastructure partnerships, positioned between Twilio (premium) and Exotel (India-native). Audio latency: 60–120ms, routing through Mumbai and Chennai PoPs. Concurrency: 100 default, 500+ enterprise. Number pool: limited variety, manual quality monitoring. Streaming API less mature than Twilio but functional. Basic DND filtering, no native NCPR integration. Pricing: ~₹0.71/minute. Optimal use case: hybrid deployments where a US headquarters manages global communications infrastructure but India-origin real estate is a secondary market, or as a budget-conscious alternative to Twilio for teams already using Plivo globally.
Tata Tele Business Services (TTBS) — SmartFlo
The enterprise division of Tata Communications, providing direct PSTN access via Tata's owned network infrastructure. Audio latency: 30–60ms — lowest of evaluated providers, as Tata carries Indian PSTN traffic as a licensed carrier with zero intermediate routing hops. Concurrency: 500–5,000 with proper enterprise provisioning. Number pool: carrier-grade, no shared reputation risk. Streaming API is less developer-friendly, typically requiring SIP expertise or a SIP-to-WebSocket gateway. Full NCPR/DND compliance built in at the network layer. Pricing: ₹0.35–₹0.50/minute, lowest of evaluated providers, with volume discounts above 100,000 minutes/month. Optimal use case: large-scale developers processing over 50,000 calls/month with dedicated SIP/VoIP expertise, where cost optimization justifies integration complexity.
Comparative Summary Table
Dimension
Exotel
Twilio
Plivo
Tata Tele
Audio Latency (one-way)
40–80ms
80–160ms
60–120ms
30–60ms
Default Concurrency
50 calls
Uncapped
100 calls
500+ calls
Max Concurrency
500 (enterprise)
Uncapped
500
5,000+
Number Pool Quality
Large, monitored
Moderate, unmonitored
Limited
Carrier-grade
Streaming API Maturity
Good
Best
Adequate
Complex (SIP)
Native DND Scrubbing
Yes
No
No
Network-level
Developer Experience
Good
Best
Good
Enterprise-sales
Cost/Minute (₹)
0.50–0.70
1.00–1.20
0.71–0.90
0.35–0.50
Setup Time
Hours
Minutes
Hours
Days–Weeks
Best For
Mid-large India RE
NRI/Global/Enterprise
Hybrid/Global
Very large scale
The Recommended Architecture by Scale
Stage 1: 0–5,000 calls/month (New Deployment)
Primary: Exotel. API-key setup in minutes, native Indian infrastructure, DND scrubbing built in, developer-friendly documentation, lowest complexity for the first deployment. Cost: ₹2,500–₹3,500/month at 5,000 calls × avg 1 minute.
Primary: Exotel (retain for domestic Hindi-Hinglish calls). Secondary: Twilio (add for NRI callers, international number provisioning, or when using Twilio ecosystem tools).
def select_telephony_provider(lead: Lead) -> str:
# NRI leads → Twilio (international number, international routing)
if lead.country_code not in ["+91"]:
return "twilio"
# Premium English-first leads → Twilio (better English TTS/ASR ecosystem)
if lead.segment == "luxury" and lead.language_preference == "english":
return "twilio"
# Default Indian domestic → Exotel
return "exotel"
Primary: Tata Tele (cost optimization and maximum concurrency). Secondary: Exotel (overflow, rapid provisioning, developer tools). Tertiary: Twilio (NRI/international).
💡
At 100,000 calls/month × 1.2 min avg: Tata Tele at ₹0.40/min = ₹48,000/month vs. Exotel at ₹0.60/min = ₹72,000/month. A ₹24,000/month saving (₹2.88 lakh/year) justifies the SIP integration engineering investment.
Critical Integration: Voicemail Detection
One frequently overlooked telephony-layer configuration for Indian real estate AI Calling is answering machine detection (AMD). Buyer calls in India go to voicemail when the buyer does not answer — the AI must detect this within 3–4 seconds and terminate the call rather than running a qualification script against a voicemail system.
Exotel AMD — built in via a machine_detection parameter in the call initiation API, returning a machine or human decision within 3 seconds
Twilio AMD — a MachineDetection=Enable parameter in the call API, 3–5 second detection time, available on all account tiers
Plivo/Tata AMD — available but less reliable for Indian carrier voicemail patterns, requiring custom validation
Without AMD configured, the AI will attempt to qualify voicemail systems, consuming LLM tokens and call-minutes, and generate false "no answer" dispositions in the CRM.
Frequently Asked Questions
Number spam flagging occurs when a single number makes over 200–300 calls/day with short average call duration (under 30 seconds). Mitigation requires three actions: number rotation — provision a pool of 5–10 numbers and distribute calls across the pool so no single number exceeds 50–80 calls/day; caller ID personalization — use project-specific or developer-branded caller IDs rather than generic VoIP ranges; and call quality improvement — short calls (hang-ups) are the primary spam signal, so improving qualification rate reduces short calls. Exotel's number health monitoring proactively flags numbers reaching spam threshold and can trigger automatic rotation; implement the same logic manually for Twilio/Plivo.
Full round-trip budget for a conversational response: PSTN + telephony transport (20–80ms), ASR transcription (150–250ms), NLU entity extraction (10–30ms), LLM inference (300–600ms), TTS first byte (80–200ms), audio stream to buyer (20–80ms) — 580ms to 1,240ms total. LLM inference is the dominant latency contributor at 40–50% of total. Telephony transport is the smallest controllable variable — at under 10% of total latency, upgrading from Twilio (160ms) to Exotel (80ms) saves 80ms, meaningful but not the primary optimization lever. Prioritize LLM inference optimization before telephony provider switching for latency reduction.
Commercial outbound calling services that operate as third-party calling service providers are subject to the Department of Telecommunications' Other Service Provider (OSP) registration framework. If you are deploying AI Calling as a platform sold to real estate developer clients — a B2B SaaS provider making calls on clients' behalf — OSP registration may be required depending on your call routing architecture. If the real estate developer deploys AI Calling for their own lead pipeline (self-use), OSP registration typically does not apply. Consult telecommunications regulatory counsel for your specific deployment structure — this is a structuring question, not a technical one, and the answer depends on whether calls originate from the developer's registered entity or a third-party service provider's infrastructure.
Multi-provider routing is standard practice at moderate-to-large scale and the integration complexity is manageable — the orchestration layer simply selects a provider per call based on lead attributes (NRI status, segment, current provider capacity) and each provider's WebSocket audio interface is normalized to the same internal audio pipeline format. The main complexity is maintaining consistent call disposition logging and CDR reconciliation across providers with different webhook payload shapes, which a thin adapter layer per provider resolves. Below roughly 5,000 calls/month, single-provider (Exotel) is simpler and sufficient; multi-provider routing becomes worthwhile once NRI volume, cost optimization, or redundancy requirements justify it.
Skipping AMD wastes both LLM inference cost and call-minute cost on every voicemail-routed call — a full qualification attempt against a voicemail greeting can run 30–90 seconds of billed call time and several LLM turns before the AI recognizes there's no real conversation happening. At typical Indian mobile voicemail/no-answer rates of 15–25% of dial attempts, deployments without AMD configured can waste 10–20% of total telephony and inference spend on non-conversations. AMD detection adds only 3–5 seconds of call setup time and reliably eliminates this waste, making it one of the highest-ROI configuration changes available at the telephony layer.
Disclaimer: Telephony pricing, concurrency limits, API features, and regulatory interpretations referenced in this article reflect market conditions and provider documentation as of Q2 2026. Provider specifications, pricing tiers, and regulatory frameworks evolve continuously. Actual per-minute costs depend on volume commitments, contract terms, and routing configurations. All financial estimates are provided for planning purposes only — verify current pricing directly with each provider before production deployment decisions. Regulatory compliance requirements must be validated with qualified legal counsel for your specific deployment jurisdiction and use case.