Overview
A fintech client wanted an AI agent that could make outbound phone calls — to new leads and to existing customers — across 20 countries in South Asia, Africa, Europe and Latin America, in each person's own language, including English and Urdu. The agent walks people through onboarding and activation, answers their questions from an approved knowledge base, handles objections, and stays inside financial-compliance rules the whole way. Operators run everything from a web dashboard.
I designed and built the platform largely on my own, working directly with the client on requirements. The client is confidential under NDA, so this page describes the engineering, not the business.
What I built
- A real-time voice pipeline. Streaming speech-to-text, a language model with tools, and streaming text-to-speech, orchestrated with LiveKit Agents. Replies start in about a second, with noise isolation, proper handling for when the caller interrupts, and tuning specifically for Urdu speech.
- A dialler that behaves. Campaign pacing, each country's legal calling hours, retry policies, do-not-call lists honoured on every channel, and automatic failover between telephony carriers.
- Carrier routing with local caller ID. Per-country caller-ID pools across SIP trunks from more than one carrier, so an agent calling Lagos shows a number that makes sense in Lagos.
- CRM integration. A scheduled sync of around 13,000 contacts, funnel-stage tracking, and campaigns that keep themselves filled with the right people as contacts move through the funnel.
- A knowledge base with an approval gate. Retrieval-augmented generation over topic-scoped content. Nothing reaches a caller until an administrator approves it, and a read-only search API lets a second product reuse the same knowledge.
- Compliance and after-call work. Guardrails on what the agent may say, call recording to a choice of S3-compatible storage, transcript translation, and sentiment and outcome analysis on every call.
- An operator dashboard. Next.js: live call queue, campaigns, contacts, funnel, analytics, diagnostics, settings, user login and a full audit log.
The interesting engineering
Latency is measured, not guessed. Every hop — endpointing, transcription, the model's first token, first audio — is timed on live calls, not in a benchmark. One of the cheaper wins was priming the language model while the phone is still ringing, so the first reply doesn't pay a cold-start penalty. I built a public latency budget calculator out of the lessons.
A SIP relay for a carrier that only trusts one IP. One carrier would only accept traffic from a single fixed IP address — incompatible with a modern autoscaled voice platform whose media servers move around. Rather than drop the carrier, I set up a Kamailio SIP relay on AWS with a static address: the carrier sees one stable peer, and the platform behind it stays elastic.
Providers are plug-ins. Speech, language model, voice and telephony all sit behind small adapters. Swapping a carrier or a speech provider is configuration plus a thin adapter, not a rewrite — which matters when rates and route quality shift country by country.
Compliance is enforced by the system, not by operators. Calling hours, contact limits and retry caps live in the dialler itself, and an opt-out on any channel is recorded and honoured immediately and permanently. An operator can't override them by accident.
Tested like it matters. More than 1,000 automated tests and CI on every change — the kind of safety net you want before letting software dial real people.
Stack
Python, FastAPI, LiveKit Agents (WebRTC / SIP), OpenAI and Anthropic models, Deepgram, Cartesia, Krisp, PostgreSQL with SQLAlchemy and Alembic, Qdrant with Voyage AI embeddings, Next.js / React / TypeScript with TanStack Query, Telnyx and IDT SIP trunking, Kamailio, AWS Lightsail, Cloudflare R2, Railway, Docker and GitHub Actions.