All work

AI · Case study

CortexCraft Voice Agent

Production platform for building and running Gemini Live phone agents. It gives you a dashboard to configure them, Twilio and Plivo call handling, RAG document grounding, saved transcripts, and an external REST API so other products can drive calls.

Role
Software Developer, CortexCraft.ai
Team
5 contributors
Timeline
May 2026 to present
Commits
589 of 626 (SaaS platform) + 67 of 75 (core engine)
Stack
Python · FastAPI · Gemini Live · Twilio · Plivo · WebSocket · RAG · React · Docker

Responsibilities

  • Campaign runner and outbound call dispatch
  • Gemini Live session handling in the core engine
  • Post-call transcript extraction and customer webhooks
  • Call-pipeline reliability: routing, socket reconnects, TLS
How it worksAn outbound call, end to end
  1. 01

    Dashboard

    frontend/app (Next.js)

  2. 02

    Campaign API

    campaigns/routes.py

  3. 03

    Worker tick

    campaigns/runner_service.py

  4. 04

    Gemini Live call

    core engine: gemini_live/client.py

  5. 05

    Phone network

    Twilio or Plivo

  6. 06

    Post-call extraction

    calls/extraction_service.py

  7. 07

    Customer webhook

    webhooks/emit.py

DataPostgreSQLRedis

Situation

Every client wanted the same thing with a different script: an AI that answers or places phone calls, knows their documents, and hands a transcript back to whatever system they already run. Rebuilding that stack per client would have meant maintaining the same telephony and streaming bugs in several places at once.

Task

Build one platform where an agent is configuration rather than code: created from a dashboard or an API, grounded in uploaded documents, reachable over more than one telephony provider, and drivable by external systems.

Action

I built a FastAPI + React platform where each agent is a JSON config plus a folder of RAG documents. A prompt generator turns that config into the full system prompt and the inbound/outbound greetings. Calls stream over a WebSocket to Gemini Live, with separate Twilio and Plivo handlers behind one interface and a browser channel for testing without burning call minutes; transcripts are written back into each agent’s conversation store. An external REST API exposes agent creation and call placement, so other products consume the platform instead of forking it. Most of the real work was production reliability: fixing outbound-call routing, adding reconnect handling for Gemini 1006 socket drops, tightening goodbye detection so calls end cleanly, and tracing a Docker TLS failure to NAT hairpinning on a self-hosted domain.

voice-agent-core-engine/backend/src/gemini_live/client.pypython
# One Gemini Live session per phone call: audio in, audio out, text alongside.
while True:
    config = self._live_config(include_language_code=include_language_code)
    connected = False
    try:
        async with self.client.aio.live.connect(model=self.model, config=config) as session:
            connected = True
            await _maybe_call(session_open_callback)
            event_queue: asyncio.Queue = asyncio.Queue()
            tasks = [
                asyncio.create_task(self._send_audio_loop(session, audio_input_queue, event_queue)),
                asyncio.create_task(self._send_text_loop(session, text_input_queue, event_queue)),
                asyncio.create_task(self._receive_loop(session, event_queue, audio_output_callback, audio_interrupt_callback)),
            ]
            # ... stream events until the call ends, then cancel the tasks
    except Exception as exc:
        if not connected and include_language_code and _language_code_rejected(exc):
            # The 3.1 Live preview may refuse SpeechConfig.language_code:
            # retry once without it so the PSTN call still connects.
            include_language_code = False
            continue
        raise

Result

One platform now backs several shipped products instead of several codebases. The multi-tenant lead dialer, KisanVoice’s phone surveys and client-specific agents all run on it, so a fix to the call pipeline reaches every one of them at once.

589

commits authored

1 API

drives every product

Live

in client deployments

Looking back

What I would do differently

  • Move agent document search from Postgres full-text to embeddings, so agents match meaning rather than exact words.
  • Write an incident report for every production fix, not only the NAT hairpinning one.