# +callforagents: a real phone line for every AI agent

> **ixigo Hackweek 2026 · built by the Alonso Team: Ernesto, Jose Miguel, Fran and Carlos (sqaas)**
>
> Live at **[callforagents.com](https://callforagents.com)** · dashboard at **app.callforagents.com** · MCP server at **mcp.callforagents.com**

AI agents can already browse, write code and pay. Most of the world's errands still go through a phone call: the dentist in Madrid, the clinic in Pune, the courier in Singapore, the landlord in Dubai. Many of those businesses have no API, no booking page and no email anyone reads. They do have a phone that rings.

**+callforagents lets any AI agent make that call.** One sentence in Claude, ChatGPT, Cursor or Claude Code ("call this restaurant and book a table for four on Saturday at nine") and a real phone rings. The person picking up sees a number they recognise, often the user's own mobile. They hear a natural, local voice that says up front it's an AI assistant. A few seconds after hanging up, the agent gets back a verified outcome, a transcript, the recording and the exact cost.

The agent can also **answer**. Missed calls on the user's mobile forward to us. The agent picks up as the user's assistant, takes the message, books what it can and sends the summary.

---

## Contents

1. [At a glance](#1-at-a-glance)
2. [Who it helps](#2-who-it-helps)
3. [The experience](#3-the-experience)
4. [System architecture](#4-system-architecture)
5. [The life of a call](#5-the-life-of-a-call)
6. [The voice engine](#6-the-voice-engine)
7. [India: local caller ID and audio that stays in India](#7-india-local-caller-id-and-audio-that-stays-in-india)
8. [Pick-up: the agent answers your missed calls](#8-pick-up-the-agent-answers-your-missed-calls)
9. [Money: a prepaid ledger that can't double-spend](#9-money-a-prepaid-ledger-that-cant-double-spend)
10. [Trust and safety, built into every layer](#10-trust-and-safety-built-into-every-layer)
11. [Identity, caller ID and the line wallet](#11-identity-caller-id-and-the-line-wallet)
12. [A first-class MCP server and ChatGPT app](#12-a-first-class-mcp-server-and-chatgpt-app)
13. [Real time everywhere](#13-real-time-everywhere)
14. [After the call: verified outcomes, search and evidence](#14-after-the-call-verified-outcomes-search-and-evidence)
15. [Mission Control: product analytics and the admin API](#15-mission-control-product-analytics-and-the-admin-api)
16. [How we engineer quality](#16-how-we-engineer-quality)
17. [The numbers](#17-the-numbers)
18. [Repository map](#18-repository-map)
19. [The team](#19-the-team)

---

## 1. At a glance

| | |
|---|---|
| **What** | Outbound and inbound phone calls (and SMS) for any AI agent, through one MCP server, a REST API and a ChatGPT app |
| **Where we call** | Spain, the rest of the EU, the UK, the US, Canada, India, Singapore and the UAE |
| **Caller ID** | The user's own verified mobile, a shared local line, or a dedicated number, chosen per country by policy |
| **Voices** | GPT-Live, Gemini Live, Grok Voice, Deepgram Flux with a choice of LLM, and ElevenLabs agents, picked per language and per call. Castilian Spanish, British English and Hinglish/Indian English voices |
| **Inbound** | Carrier call forwarding from the user's own mobile to our local lines; the agent answers by name |
| **Results** | Verified outcome (confirmed only when the *other side* says so), summary, highlights tied to the exact audio, timed transcript, recording, line-item cost |
| **Money** | Prepaid, per-second billing; a Durable Object ledger per account with reserve, extend and settle; Stripe top-ups in local currency |
| **Safety** | AI disclosure on every call, objective and destination screening, quiet hours, opt-out list, live safety monitor, masked data, audit trail |
| **Stack** | 100% Cloudflare: Workers, Durable Objects (SQLite), D1 with FTS5, R2, KV, Queues, Analytics Engine, rate limiters. LiveKit Cloud for India media |
| **Built in** | About two weeks: 176 commits, ~33,000 lines of TypeScript, 51 database migrations, all in production |

---

## 2. Who it helps

```mermaid
mindmap
  root((+callforagents))
    Agent builders and power users
      Call these 10 clinics, find a Tuesday slot
      Claude · ChatGPT · Cursor · Claude Code · custom agents
      Pay per second, no setup
    Busy professionals
      Answer my phone when I can’t
      Keep their own number
      Summary after every missed call
    People calling across borders
      A Spanish number for Spain
      An Indian number for India
      A voice that sounds local
    Small teams
      One account, many agents
      Spend limits and per-agent permissions
```

**Why this matters in India and the markets we serve.** US-first agent-calling products call every country from a US number, and that call gets ignored or blocked. Voice platforms expect a developer to build an assistant and provision numbers before anything happens. +callforagents works the other way round. A general-purpose agent just asks for a call, and we handle the rest:

- the right local caller ID
- a voice that fits the place
- the law on caller ID and AI disclosure
- what to do if nobody answers
- the bill

Testers in India have been making real calls to friends and family through our Indian line: house-party check-ins, quick confirmations, messages passed on in Hinglish. They hear a local voice from a number they recognise.

---

## 3. The experience

```mermaid
sequenceDiagram
  autonumber
  actor U as User
  participant A as Their AI agent (Claude, ChatGPT…)
  participant P as +callforagents
  actor R as Restaurant
  U->>A: "Book a table for 4 on Saturday at 9 at La Tasca"
  A->>P: find_business("La Tasca", near "Chamberí, Madrid")
  P-->>A: phone, address, opening hours
  A->>P: prepare_call (free quote)
  P-->>A: caller ID shown, price/min, max minutes, disclosure line
  A->>U: "I'll call from your number, about $0.15/min. OK?"
  U->>A: "Go"
  A->>P: place_call(objective, user_confirmed)
  P->>R: ☎ rings from the user's own Spanish mobile
  R-->>P: "La Tasca, ¿dígame?"
  P->>R: "Hola, soy Nora, la asistente de IA de Ernesto…"
  Note over P,R: natural Castilian conversation, barge-in, tools
  P-->>A: outcome: confirmed by the restaurant ✔ · transcript · recording · $0.21
  A->>U: "Booked for 4, Saturday 21:00, under Ernesto"
```

Everything a person would expect from a good assistant is built in:

- **Finds the number** (`find_business`, via Google Maps with Search as fallback).
- **Quotes before dialling** (`prepare_call`): the caller ID, the price, the longest call the balance allows, and the exact disclosure sentence.
- **Asks you mid-call.** The voice agent can put one question back to the agent that placed the call ("is Thursday at 4 OK?") and relay the answer (`live_questions`, `answer_call_question`).
- **Lets you take over.** `jump_in` rings you and joins you into the live call.
- **Retries smartly** (`retry_call`), keeps **contacts** (`save_contact`, `import_contacts`), sends **SMS** from your own number, and **searches** every call you've made by what was said in it.
- **Shows the call inside ChatGPT** as a live widget, with a player and synced transcript.

---

## 4. System architecture

```mermaid
flowchart TB
  subgraph Clients["Where people and agents are"]
    direction LR
    MCPC["Claude · ChatGPT · Cursor<br/>Claude Code · custom agents"]
    WEB["Dashboard (React, EN/ES)<br/>app.callforagents.com"]
    SITE["Public site<br/>callforagents.com"]
    ADM["Mission Control + Admin API"]
  end

  subgraph CF["Cloudflare edge"]
    direction TB
    PLAT["<b>cfa-platform</b> Worker<br/>MCP (2026-07-28 + legacy) · OAuth 2.1 · REST<br/>Stripe webhooks · inbound routing · policy"]
    subgraph DOs["Durable Objects (SQLite)"]
      LED[("LedgerDO<br/>one per account<br/>holds · entries")]
      HUB[("AccountHub<br/>one per account<br/>WebSockets · SSE · budgets")]
      CALL[("CallDO<br/>one per call<br/>bridge · VAD · trace · cost")]
    end
    ENG["<b>voice-playground</b> Worker<br/>call engine + benchmarking lab"]
    D1[("D1<br/>accounts · phones · lines · calls<br/>contacts · FTS5 call search · rules")]
    R2[("R2<br/>recordings · snapshots<br/>support files · contact photos")]
    Q[["Queues<br/>post-call review · webhooks<br/>support email (+ DLQ)"]]
    AE[("Analytics Engine<br/>usage events")]
    KV[("KV<br/>OAuth grants · edge budgets")]
  end

  subgraph Voice["Voice models"]
    direction LR
    OAI["GPT-Live<br/>(WS + SIP)"]
    GEM["Gemini Live"]
    XAI["Grok Voice"]
    DG["Deepgram Flux<br/>+ LLM"]
    EL["ElevenLabs<br/>Agents"]
  end

  subgraph Tel["Telephony"]
    direction LR
    TW["Twilio<br/>Media Streams · SIP"]
    LK["LiveKit Cloud (India)<br/>cfa-bridge agent"]
    PL["Plivo<br/>Indian SIP trunk"]
    VO["Vonage<br/>SMS"]
  end

  PHONE(("📱 Phones<br/>worldwide"))
  STRIPE["Stripe<br/>Checkout · Tax"]

  MCPC -- "Streamable HTTP" --> PLAT
  WEB -- "HTTPS + WebSocket" --> PLAT
  ADM --> PLAT
  PLAT <--> LED
  PLAT <--> HUB
  PLAT --> D1
  PLAT -- "service binding (RPC)" --> ENG
  ENG --> CALL
  CALL -- "push call changes" --> HUB
  CALL --> R2
  PLAT --> Q
  PLAT --> AE
  PLAT --> KV
  CALL <--> OAI & GEM & XAI & DG & EL
  CALL <-- "media WS" --> TW
  CALL <-- "media WS" --> LK
  LK <-- SIP --> PL
  TW <--> PHONE
  PL <--> PHONE
  PLAT --> VO --> PHONE
  STRIPE -- "signed webhooks" --> PLAT
```

**Design principles that made it possible to ship this much, this fast:**

| Principle | How it shows up |
|---|---|
| **One owner per piece of state** | One Durable Object per **call** (its audio, timing and trace), one per **account** for money (strictly serial: no locks, no double-spend), and one per account for **real-time** fan-out. |
| **The engine is a product of its own** | The call engine began as a benchmarking playground for realtime voice models. The platform drives it over a typed service binding, so every engine improvement reaches production calls at once. |
| **Policy as data** | Prices, destinations, calling hours, per-number limits and free credit live in a rules table that owners edit from Mission Control. Changes are live within a minute, with no deploy. Each country's caller-ID law is a routing table, not a code path. |
| **Everything idempotent** | Ledger entries have unique refs (`stripe:{session}`, `call:{id}`). Settlement is guarded by a single conditional `UPDATE`, so the hub, `get_call` and the sweeper can race safely. Queue jobs can be redelivered without harm. |
| **Privacy by construction** | Phone numbers are masked by default in every internal view. Call content opens only when the user shares it, and every look is shown to them. Analytics rows never contain a phone number. |

---

## 5. The life of a call

```mermaid
sequenceDiagram
  autonumber
  participant A as AI agent
  participant P as Platform (MCP)
  participant S as Safety + routing
  participant L as LedgerDO
  participant C as CallDO
  participant K as Carrier
  participant Ph as Phone
  participant Q as Queue
  A->>P: place_call {to, objective, on_behalf_of}
  P->>S: screen objective (rules + LLM classifier) + destination (country, quiet hours, premium, opt-out)
  S-->>P: allow · confirm · block, plus prompt rules
  P->>P: chooseLine (own · shared · dedicated) by country policy
  P->>L: reserve(call, first window)
  P->>C: start(prompt from objective, voice, language, caller ID, time limit)
  C->>K: dial (Media Streams, SIP direct, or LiveKit)
  P-->>A: {call_id, status: queued, call_url}
  K->>Ph: ring
  Ph-->>K: answer
  K->>C: media stream starts
  C->>C: model starts and greets with the AI disclosure
  loop while talking
    C->>L: extend(call) (refused means a graceful goodbye)
    C-->>P: live status + transcript (hub → dashboard, SSE, webhooks)
  end
  C->>C: hang-up (goodbye watcher, end_call, silence watchdog or limit)
  C->>L: settle(actual cost), release the rest of the hold
  P->>Q: review job
  Q->>C: Gemini listens to the recording + transcript
  C-->>P: verified outcome, highlights with audio timings
  P-->>A: get_call: outcome · transcript · recording · cost
```

Every step leaves a trace. The CallDO keeps a timed event log of the whole call (carrier events, model events, VAD turns, tool calls, hang-up reason), which powers support, quality analysis and the latency charts.

---

## 6. The voice engine

The engine is the heart of the product. It turns a phone line into a natural, low-latency conversation, whichever model is speaking.

```mermaid
flowchart LR
  subgraph In["Caller audio (8 kHz μ-law)"]
    MS["Media Streams / LiveKit frames"]
  end
  subgraph Bridge["CallSession (one per call)"]
    direction TB
    CV["Caller VAD<br/>(energy, adaptive floor)"]
    GAP["Gap filler<br/>(silence for quiet trunks)"]
    REC["Dual-channel recorder<br/>(playback clock)"]
    AV["Agent VAD<br/>(what the phone actually heard)"]
    BG["Background ambience mixer<br/>(office · café · city…)"]
    BYE["Goodbye watcher +<br/>silence watchdog"]
    TOOLS["Tools: end_call · web_search<br/>get_current_time · ask_user_agent<br/>agent-defined live_tools"]
  end
  subgraph Prov["LiveProvider interface"]
    direction TB
    P1["openai-live"]
    P2["gemini-live"]
    P3["xai-live"]
    P4["deepgram-agent"]
    P5["elevenlabs"]
  end
  OUT["Phone"]
  MS --> CV --> GAP --> Prov
  Prov --> BG --> OUT
  Prov --> REC
  MS --> REC
  REC --> AV
  CV -- "barge-in: clear queued audio" --> Prov
  Prov --> TOOLS
  BYE -- "hang up" --> OUT
```

**What makes it feel human:**

- **One interface, five model families.** Each provider (GPT-Live, Gemini Live, Grok Voice, Deepgram Flux with a chosen LLM, ElevenLabs Agents) sits behind the same `LiveProvider` contract: `connect`, `greet`, `sendAudio`, `interrupted`, transcripts, tools, usage. Picking the best engine per language is a setting, not a rewrite. Admins can override the engine per call for side-by-side tests.
- **Barge-in that actually works on phones.** Our own caller VAD detects the other person speaking over the agent. The engine then clears the audio Twilio or LiveKit has queued, so the agent stops mid-word, as a person would.
- **A playback clock, not a send clock.** Agent audio is timestamped by when the phone *plays* it, not when we sent it. Latency metrics, transcript timings and the recording all reflect what the person really heard.
- **Natural endings.** Realtime models often say "bye" and never call `end_call`. A goodbye watcher in English and Spanish recognises when both sides have closed and hangs up cleanly. A silence watchdog ends calls that went quiet.
- **Localisation, never disguise.** 26 curated voices tagged by locale (`es-ES`, `en-GB`, `hi-IN`), scenario prompts in Castilian Spanish, British English and Hinglish, and optional soft background ambience. The AI disclosure is said in every voice, on every call.
- **Agent-defined tools during the call.** The calling agent can hand the voice agent up to four of its own tools (`live_tools`) and a topic it may ask about (`live_questions`). Each question is treated as untrusted, limited to that topic, capped in number, refused if it asks for secrets, and relayed as information only.
- **Pre-warming.** While the phone rings, the engine opens the model session or fetches signed URLs, so the first words come right after pickup.

---

## 7. India: local caller ID and audio that stays in India

Calling India well needs an Indian number, an Indian voice and media that doesn't detour through another continent. We built a dedicated path for it.

```mermaid
sequenceDiagram
  autonumber
  participant C as CallDO (engine)
  participant LK as LiveKit Cloud (India)
  participant B as cfa-bridge agent (Node)
  participant PL as Plivo SIP trunk
  participant Ph as 📱 Jio / Airtel / Vi phone
  C->>LK: CreateRoom + dispatch cfa-bridge
  C->>LK: CreateSIPParticipant (Indian caller ID, Krisp on)
  LK->>PL: SIP INVITE
  PL->>Ph: ring
  Note over C: prewarm only, the model waits for pickup
  Ph-->>PL: answer
  PL-->>LK: 200 OK
  LK-->>B: sip.callStatus = active
  B->>C: open media WebSocket (Media Streams protocol)
  C->>C: start model and greet ("नमस्ते, Anika here…")
  loop 20 ms frames
    Ph->>B: caller audio (Opus → PCM)
    B->>C: μ-law frames
    C->>B: agent speech bursts
    B->>B: FIFO pump, 20 ms copied frames, ≤400 ms queued
    B->>Ph: smooth playout
  end
  Ph-->>B: hang up
  B->>C: stop + per-call audio stats
```

**The engineering behind the India line:**

- **Same protocol, new transport.** The `cfa-bridge` LiveKit agent speaks the Twilio Media Streams protocol to the engine. The whole voice engine runs unchanged on the Indian path: every provider, VAD, recorder, tool and post-call review.
- **Answer-gated start.** The model starts only once LiveKit reports the callee answered. The greeting, with its AI disclosure, is the first thing the person hears.
- **Smooth audio at any burst size.** Some engines send several seconds of speech at once. The bridge keeps its own FIFO and feeds LiveKit exactly one 20 ms frame at a time, keeping the queue under 400 ms. That gives smooth playout, instant barge-in and no overflow.
- **Per-call audio telemetry.** Engine frames in, frames delivered, capture errors, slow or stuck captures, barge-in clears and caller frames are logged for every call, so audio quality is measured, not guessed.
- **Krisp noise cancellation** on the callee leg, and ElevenLabs sessions configured for native 8 kHz μ-law in both directions (no resampling on the hot path).
- **Hinglish done right.** Indian voices (Anika, Anjali, Raju, Viraj…), a Hinglish prompt style, and English calls to India defaulting to an Indian English voice.

---

## 8. Pick-up: the agent answers your missed calls

```mermaid
flowchart LR
  X(("Caller")) -->|calls| M["User's mobile"]
  M -->|"busy · no answer · unreachable<br/>(carrier forwarding)"| D["Our local line"]
  D --> R{"routeInbound"}
  R -->|"Diversion / History-Info header"| U1["User found by forwarded number"]
  R -->|"dedicated line"| U2["User found by number"]
  R -->|"caller is a known contact"| U3["User found by contact"]
  R -->|"shared line, unknown"| ASK["Ask who they're calling"]
  U1 & U2 & U3 --> AG["Agent answers by name:<br/>'Hi, you've reached Maria's AI assistant'"]
  AG --> SUM["Summary + recording<br/>→ dashboard · webhook · the user's agent"]
```

- **Activation in one tap.** We generate the GSM forwarding codes (`**61*<number>**20#`, `**67*…#`, `**62*…#`, `**004*…#`) and `tel:` links that open the dialler prefilled. Country profiles cover where operators expect forwarding to be set in Settings instead.
- **Robust caller detection.** We parse SIP `Diversion` and `History-Info` (with RFC 4458 causes) and carrier "forwarded from" fields across Twilio, Telnyx and Vonage.
- **The user's own instructions and webhook.** Pick-up follows each user's instructions (what to say, what to collect) and posts the finished call, with the conversation, to their agent.

---

## 9. Money: a prepaid ledger that can't double-spend

```mermaid
stateDiagram-v2
  [*] --> Reserved: reserve(call, first window)\nbefore dialling
  Reserved --> Talking: answered
  Talking --> Talking: extend(window)\nevery minute
  Talking --> Ending: extend refused\n(balance) → polite goodbye
  Talking --> Settled: settle(actual cost)
  Ending --> Settled: settle(actual cost)
  Reserved --> Released: not answered\n(free)
  Settled --> [*]: charge entry + hold released
  Released --> [*]
```

- **One LedgerDO per account.** Every balance change for an account is serial by construction, so concurrent calls can't overspend. There are no distributed locks, and the balance updates in real time.
- **Integer micro-dollars.** 1 USD = 1,000,000 µ$, so there's no floating-point drift. Top-ups in EUR, INR, SGD or AED mint credits at the pack's fixed USD value, so the ledger never holds FX.
- **Per-second billing, after a 10-second minimum**, from a simple two-tier rate card (standard / premium destinations). Unanswered calls are free.
- **Stripe as money-in only.** Signed webhooks credit the ledger idempotently. Refunds return the right share of credits, and chargebacks freeze the account automatically. Sales are recorded with Stripe's fee, the card country and the presentment currency.
- **Credit with expiry**: trial credit for a confirmed number (once per number, ever), vouchers and gift links, and an agent-connection reward, each tracked to its source.

---

## 10. Trust and safety, built into every layer

```mermaid
flowchart TB
  O["Objective + destination"] --> S1["Objective screen<br/>rules + LLM classifier<br/>(personal · service · commercial)"]
  O --> S2["Destination screen<br/>country · emergency · premium · short codes<br/>quiet hours at the destination · opt-out list · block list"]
  S1 & S2 --> V{"verdict"}
  V -->|block| B["Refused, with the reason"]
  V -->|confirm| C["needs_confirmation:<br/>first call to a new number,<br/>personal mobiles, relationship asked"]
  V -->|allow| P["Prompt rules appended:<br/>AI disclosure · no sales · no card numbers or codes"]
  P --> L["Live safety monitor<br/>opt-out → platform-wide list<br/>repeated abuse → hang up<br/>codes/cards read out → flagged"]
  L --> R["Post-call review + audit trail"]
```

- **AI disclosure, always.** Every call opens with "…<name>'s AI assistant", in every voice and language. That's the EU's direction of travel and the right default everywhere.
- **Not for spam, by design.** No sales, surveys, debt collection or political calls. First calls to new numbers and calls to personal mobiles need the user's explicit OK, carried as `user_confirmed`.
- **Quiet hours at the destination**, computed from the callee's timezone.
- **Rate limits at every layer**: Cloudflare rate limiters on sign-up, MCP and anonymous traffic; exact durable counters in a guard Durable Object for global budgets; per-number attempt and per-day limits, with higher limits for trusted accounts.
- **Secrets never cross the call.** The voice agent never asks for or repeats card numbers, one-time codes or passwords, and a live monitor flags them if the other side reads them out.
- **Privacy-first operations.** Internal views mask phone numbers. Unmasking is explicit and audited. Call content opens to the team only through a support request the user shares (time-boxed and revocable) or the user's opt-in to "Help improve". Each look is logged and shown to the user.

---

## 11. Identity, caller ID and the line wallet

```mermaid
flowchart LR
  subgraph Wallet["Line wallet (one account)"]
    OWN["<b>own</b><br/>user's verified mobile<br/>reaches every country we call"]
    SH["<b>shared</b><br/>our local pool per country<br/>calls its own country"]
    DED["<b>dedicated</b><br/>rented for one account"]
  end
  REQ["place_call → country"] --> POL{"Country policy +<br/>user's own_scope"}
  POL --> OWN
  POL --> SH
  POL --> DED
  OWN -. "carrier refuses" .-> RETRY["retry once<br/>from a shared local line"]
```

- **Sign-in by phone.** We place a verification call to the user, who types the code shown on screen, or the carrier confirms the number. That one step proves ownership and registers the number as a caller ID with the carrier, so calls can show the user's own mobile from day one.
- **Per-agent permissions.** Each connected agent (via OAuth) or API key has its own grant. Using the owner's personal number is deny-by-default for agents until the owner allows it.
- **User-chosen reach.** The user decides where their own number shows: everywhere, only their country, only where we have no local line, or picked countries. Shared local lines cover the rest.

---

## 12. A first-class MCP server and ChatGPT app

- **Both MCP eras in one stateless server.** It supports the **2026-07-28** protocol (no handshake, `server/discover`, `resultType`, modern error codes and HTTP statuses) and the legacy `initialize` handshake (2025-11-25, 2025-06-18, 2025-03-26), over Streamable HTTP, with no sessions to manage.
- **OAuth 2.1** with PKCE, RFC 9728 protected-resource metadata and dynamic client registration, so connecting from Claude or ChatGPT is one click. Headless agents use hashed API keys (`cfa_live_…`).
- **26 carefully designed tools**, each with JSON Schema input and output, read-only, destructive and open-world annotations, and descriptions written for models. Taken together, the tool set teaches the agent the right flow (profile → quote → confirm → call → result).
- **ChatGPT app.** Calls render as an interactive widget inside ChatGPT, with live status, a recording player and a synced transcript. It's packaged with a plugin manifest and skills.
- **Signed webhooks and live streams.** `webhook_url` plus HMAC-SHA256 signatures (`CFA-Signature t=…,v1=…`) for `call.started | status | transcript | ended | reviewed`. Agents can also follow a live call over Server-Sent Events.

---

## 13. Real time everywhere

```mermaid
flowchart LR
  CALL["CallDO<br/>status · transcript lines"] -- "push on every change" --> HUB["AccountHub DO<br/>(one per account)"]
  HUB -- "hibernating WebSocket" --> DASH["Dashboard<br/>ringing → answered → ended"]
  HUB -- "SSE" --> AG["Agents following a call"]
  HUB -- "signed webhooks (queued)" --> WH["Agent backends"]
  PLAT["Platform events"] -- "admin_events + __admin hub" --> MC["Mission Control live log"]
```

- **Hibernating WebSockets.** Idle dashboards cost nothing, and the object wakes only to deliver an event. The engine pushes every change; polling is a safety net that speeds up when pushes stop.
- **Live transcript** to the dashboard, to agents over SSE and to webhooks, as the call happens.

---

## 14. After the call: verified outcomes, search and evidence

- **Verified, not claimed.** Gemini listens to the recording with the transcript and returns a summary, an outcome status and per-objective verdicts. An objective counts as **confirmed only when the other party said so**, never because the AI says it succeeded.
- **Evidence you can play.** Highlights (confirmations, decisions, constraints, follow-ups) cite transcript segments with audio-aligned timings, so the dashboard plays the exact clip. Every quote is checked against real segment ids, so nothing is invented.
- **Full-text search over every call.** An FTS5 index in D1 covers objectives, summaries and both sides of every conversation, with speaker and time. Prefix and accent-insensitive search means "cita|appointment" finds either language.
- **Dual-channel recordings in our own R2.** They're independent of any provider and kept on a fixed retention policy.

---

## 15. Mission Control: product analytics and the admin API

The owners' control room is a live, dark-themed dashboard for running the platform.

- **Product metrics:** DAU, WAU and MAU, stickiness, new, returning, resurrected and churned users, weekly retention cohorts, activation (time to first call, time to first payment), accounts and numbers by country code, agents by client, payers, ARPU and ARPPU.
- **Growth and sales:** daily trends against the previous period, the signup funnel, and Stripe payments, refunds and chargebacks with fees and card countries.
- **A persistent live event log:** sign-ups, numbers confirmed, agents connected, payments, calls, safety flags, help requests and vouchers. It's kept for 90 days and streams live, and clicking a line opens that user's 360 view.
- **Operations:** live calls on a radar, traffic over 48 hours, engine latency and margins per model, the Twilio queue, and rules, blocks and holds editable without a deploy.
- **Read-only admin API** (`cfa_admin_…` keys): the same numbers for scripts and agents, with a cursor-based event feed (`/admin/events?after=`). Phone numbers are always masked, call content never opens through a key, and lookups are audited.
- **Usage analytics on Workers Analytics Engine**, sampled and queried with SQL templates, never containing a phone number.

---

## 16. How we engineer quality

```mermaid
flowchart LR
  PG["Voice playground<br/>same engine, every model"] --> SIM["Simulated callers<br/>(scripted TTS, same timing)"]
  PG --> REAL["Real phone calls"]
  SIM & REAL --> MET["Metrics per call<br/>first-audio latency · response p50/p90<br/>turns · interruptions · cost"]
  MET --> ACC["Accent scoring<br/>(native Castilian / British / Indian)"]
  MET --> PICK["Best engine + voice per language"]
  PICK --> PROD["Production defaults"]
  REAL --> FOR["Audio forensics<br/>dual-channel spectrograms · frame stats"]
  FOR --> PROD
```

- **The playground is the lab.** The same engine runs against simulated callers and real phones, measuring latency, turn-taking, cost and accent nativeness, so engine and voice choices per language are backed by data.
- **Audio forensics as a habit.** Every call has a dual-channel recording, a timed trace and per-call bridge statistics. Spectrograms of what the phone actually heard let us tune audio to the frame.
- **Tests for the logic that matters**: rate card and margins, Stripe signature and refund maths, caller-ID policy, forwarding codes, Diversion parsing, MCP schema conformance, safety screens, contacts, analytics maths (DAU/WAU/MAU, cohorts) and admin key scopes.
- **Continuous delivery.** Every merged change goes to production on Cloudflare in seconds. Database changes are versioned migrations (51 so far).

---

## 17. The numbers

| | |
|---|---|
| Commits in about two weeks | **176** |
| Lines of TypeScript | **~33,000** |
| D1 migrations | **51** |
| MCP tools | **26** |
| Voice model families | **5** (plus OpenAI over SIP) |
| Curated voices | **26** across `es-ES`, `en-GB` and `hi-IN` |
| Durable Object classes | **4** (CallDO, Registry, LedgerDO, AccountHub) |
| Countries called | **ES · EU · UK · US · CA · IN · SG · AE** |
| Dashboard languages | **English and Spanish** |
| Time from push to production | **seconds** |

---

## 18. Repository map

| Path | What |
|---|---|
| `apps/platform` | Platform Worker: MCP server, OAuth, REST API, Stripe, ledger, safety, routing, Pick-up, search, Mission Control, and the dashboard (`app/`, React) |
| `apps/voice` | Call engine and playground: `src/` is the shared engine (bridge, providers, VAD, recorder, prompts, tools), `worker/` the Cloudflare Worker (CallDO, LiveKit, OpenAI SIP, review) |
| `apps/pickup-agent` | `cfa-bridge`: the LiveKit agent for India (outbound and Pick-up) |
| `apps/site` | Public site (EN/ES) |
| `apps/chatgpt-plugin` | ChatGPT plugin manifest, MCP config and skills |
| `packages/ai`, `packages/analytics` | Shared AI keys and providers; consent-first product analytics |
| `docs/hackweek` | Product, research, decisions and build notes (20 documents) |
| `launch` | Launch film: scenes, voice-over, music and render scripts |

---

## 19. The team

**Alonso Team, from sqaas:**

- **Ernesto**
- **Jose Miguel**
- **Fran**
- **Carlos**

We set out to give every AI agent something that sounds simple and is hard to do well: a real, trustworthy phone line that works in the countries where people need it most, sounds local and is honest about being an AI. Behind one sentence typed into an agent sit carrier routing, local caller-ID law, realtime voice models, frame-accurate audio, a serial money ledger, verified outcomes and privacy-first operations.

**One sentence in, one real call out. Thank you for reading.** ☎️
