AI · Case study
CortexCraft Voice Agent
Production platform for building and running Gemini Live phone agents. It gives you a dashboard to configure them, Twilio and Plivo call handling, RAG document grounding, saved transcripts, and an external REST API so other products can drive calls.
- Role
- Software Developer, CortexCraft.ai
- Team
- 5 contributors
- Timeline
- May 2026 to present
- Commits
- 589 of 626 (SaaS platform) + 67 of 75 (core engine)
- Stack
- Python · FastAPI · Gemini Live · Twilio · Plivo · WebSocket · RAG · React · Docker
Responsibilities
- Campaign runner and outbound call dispatch
- Gemini Live session handling in the core engine
- Post-call transcript extraction and customer webhooks
- Call-pipeline reliability: routing, socket reconnects, TLS
- 01
Dashboard
frontend/app (Next.js)
- 02
Campaign API
campaigns/routes.py
- 03
Worker tick
campaigns/runner_service.py
- 04
Gemini Live call
core engine: gemini_live/client.py
- 05
Phone network
Twilio or Plivo
- 06
Post-call extraction
calls/extraction_service.py
- 07
Customer webhook
webhooks/emit.py
DataPostgreSQLRedis
Situation
Every client wanted the same thing with a different script: an AI that answers or places phone calls, knows their documents, and hands a transcript back to whatever system they already run. Rebuilding that stack per client would have meant maintaining the same telephony and streaming bugs in several places at once.
Task
Build one platform where an agent is configuration rather than code: created from a dashboard or an API, grounded in uploaded documents, reachable over more than one telephony provider, and drivable by external systems.
Action
I built a FastAPI + React platform where each agent is a JSON config plus a folder of RAG documents. A prompt generator turns that config into the full system prompt and the inbound/outbound greetings. Calls stream over a WebSocket to Gemini Live, with separate Twilio and Plivo handlers behind one interface and a browser channel for testing without burning call minutes; transcripts are written back into each agent’s conversation store. An external REST API exposes agent creation and call placement, so other products consume the platform instead of forking it. Most of the real work was production reliability: fixing outbound-call routing, adding reconnect handling for Gemini 1006 socket drops, tightening goodbye detection so calls end cleanly, and tracing a Docker TLS failure to NAT hairpinning on a self-hosted domain.
# One Gemini Live session per phone call: audio in, audio out, text alongside.
while True:
config = self._live_config(include_language_code=include_language_code)
connected = False
try:
async with self.client.aio.live.connect(model=self.model, config=config) as session:
connected = True
await _maybe_call(session_open_callback)
event_queue: asyncio.Queue = asyncio.Queue()
tasks = [
asyncio.create_task(self._send_audio_loop(session, audio_input_queue, event_queue)),
asyncio.create_task(self._send_text_loop(session, text_input_queue, event_queue)),
asyncio.create_task(self._receive_loop(session, event_queue, audio_output_callback, audio_interrupt_callback)),
]
# ... stream events until the call ends, then cancel the tasks
except Exception as exc:
if not connected and include_language_code and _language_code_rejected(exc):
# The 3.1 Live preview may refuse SpeechConfig.language_code:
# retry once without it so the PSTN call still connects.
include_language_code = False
continue
raiseResult
One platform now backs several shipped products instead of several codebases. The multi-tenant lead dialer, KisanVoice’s phone surveys and client-specific agents all run on it, so a fix to the call pipeline reaches every one of them at once.
589
commits authored
1 API
drives every product
Live
in client deployments
Looking back
What I would do differently
- Move agent document search from Postgres full-text to embeddings, so agents match meaning rather than exact words.
- Write an incident report for every production fix, not only the NAT hairpinning one.