Skip to main content
This integration combines Vobiz telephony with the Pipecat voice agent framework to build intelligent AI-powered phone calls.

Overview

What you’ll build: An outbound calling system with real-time AI conversations powered by OpenAI (STT → LLM → TTS), automatic call recording, and bidirectional audio streaming. Three moving parts do all the work:

Call flow

ActorsYouserver.pyVobizbot.pyCustomer
Phase 1 · Placing the call
1
You→server.py
Trigger the call.
2
server.py→Vobiz
Server calls the Vobiz Call API, filling answer_url from PUBLIC_URL.
3
Vobiz→server.py
Call accepted.
Phase 2 · Answer & instructions
4
Vobiz→Customer
Vobiz dials out over PSTN — the phone rings, then the customer answers.
5
Vobiz→server.py
On answer, Vobiz asks your server what to do.
6
server.py→Vobiz
Returns VobizXML: record the session, and stream the caller’s audio to your bot.
Phase 3 · The live conversation
7
Vobiz→bot.py
WebSocket upgrade on /ws, then a single start event carrying the IDs and the negotiated audio format.
parse_vobiz_start() reads this before the transport is built. mediaFormat is authoritative — prefer it over your own contentType.
8
Vobiz↔bot.pybidirectional · ~20 ms frames
Audio flows both ways as base64 media events.
VobizFrameSerializer→Silero VAD→OpenAI STT→GPT→OpenAI TTS→back to Vobiz
Phase 4 · Teardown & recording
9
Customerhangs up →Vobiz→bot.py
With auto_hang_up=True the serializer also issues a REST hangup, so a bot-initiated end tears the call down cleanly.
10
Vobiz→server.py
Recording callback fires; the helper downloads the file.
Steps 1–3 are only for outbound calls. For inbound calls the flow starts at step 6 — Vobiz calls your /answer URL directly. See Receiving inbound calls.

Features

AI Voice Conversations

Natural conversations powered by OpenAI GPT + TTS/STT

Outbound Calling

Trigger calls via REST API from anywhere

Automatic Recording

All conversations automatically recorded and saved

Real-time Streaming

Bidirectional audio via WebSockets

Prerequisites

Vobiz Account with Auth ID and Auth Token → Sign up
OpenAI API Key for LLM, STT, and TTS → Get API key
Python 3.11+ installed on your system
ngrok for local development → Download ngrok
Python 3.11 is a hard floor. pipecat-ai 1.x declares requires-python >=3.11. On Python 3.10, pip does not fail — it silently backtracks and installs an ancient 0.0.x Pipecat, and then every import in bot.py breaks with confusing ModuleNotFoundErrors. Check with python --version before installing.

Version compatibility

This integration targets the Pipecat 1.x API. Versions are pinned in requirements.txt:
Pipecat 1.9 and 1.10 are published but are not covered by this pin. Widen the upper bound only after re-testing — the serializer contract has changed once already inside 1.x (FrameSerializer.setup() took a StartFrame before 1.3 and a FrameProcessorSetup after).
Pipecat 1.x deprecated pipecat.pipeline.task in favour of pipecat.pipeline.worker (PipelineTask → PipelineWorker). The old import path still works for all of 1.x and is what this repo uses, but it is scheduled for removal in Pipecat 2.0 — which is why the pin is <2.

Installation

1

Clone the repository

2

Install dependencies

This installs FastAPI, Pipecat, OpenAI SDK, and other required packages.
3

Configure environment

The repo ships an env.example file. Copy it and fill in your values:
.env
RequiredOptional
VOBIZ_L16_ENDIAN only matters for audio/x-l16. RFC 2586 specifies big-endian (network byte order), but some Vobiz accounts transport L16 as little-endian. If L16 media frames arrive at the expected size yet STT returns no transcripts, flip this to le. Leave the default μ-law encoding and you never hit this.

Usage

1

Start the server

The server runs on http://0.0.0.0:7860.
2

Start ngrok

In a new terminal, expose your local server:
Copy the ngrok URL from the output (for example, https://abc123.ngrok-free.app).
Important: Update PUBLIC_URL in your .env file with this ngrok URL, then restart the server.
3

Make a call

There are two ways to trigger an outbound call - pick whichever fits your stack.
The repo exposes POST /start on the local server. It wraps the Vobiz Call API and auto-fills answer_url from your PUBLIC_URL, plus uses VOBIZ_PHONE_NUMBER as from when set.
The field is phone_number, not to. server.py returns 400 Missing 'phone_number' in the request body for anything else.
Response:
What happens next:
  • Phone rings at the phone_number you passed
  • When answered, Vobiz requests XML from your server’s /answer endpoint
  • Server returns <Record> + <Stream> pointing at wss://…/ws
  • Vobiz opens the WebSocket and sends a start event, then media frames
  • AI assistant speaks and listens (STT → LLM → TTS)
  • On hangup, Vobiz posts the recording URL to /recording-ready, which downloads it to recordings/

How the flow works

Everything above is driven by one XML document and one WebSocket protocol. This section is what you need if you are adapting the repo rather than running it as-is.

The answer XML

/answer returns this (with {PUBLIC_URL} and the wire format substituted in):
Why each attribute matters:
audioTrack="both" is incompatible with bidirectional="true". The media server accepts the XML, sends a start event, and then never sends audio — a silent failure that looks like a broken bot. Use audioTrack="inbound" for any bidirectional stream. To capture both legs, use two separate one-way <Stream> elements.

The WebSocket protocol

Vobiz speaks JSON text frames. VobizFrameSerializer handles all of this for you; the shapes are here so you can debug what you see on the wire.
bot.py reads this with parse_vobiz_start() before building the transport, because it carries both the IDs needed for REST hangup and the negotiated audio format.
mediaFormat is authoritative. It reflects what the media server actually negotiated. If it differs from the contentType you asked for in your <Stream> XML, trust this event — the XML is a request, the start event is the agreement. The serializer adopts these values automatically and logs a warning on mismatch.
At 8 kHz μ-law this arrives roughly every 20 ms as 160-byte payloads. The serializer base64-decodes, converts μ-law → linear PCM, and resamples to the pipeline rate; outbound TTS makes the same trip in reverse.
With auto_hang_up=True the serializer also issues a REST hangup using call_id + your auth credentials, so a bot-initiated end tears the call down properly rather than leaving it hanging.

Sample rates

The Pipecat pipeline

bot.py builds a standard cascaded pipeline. Two details are specific to Pipecat 1.x and telephony:
bot.py
Two mistakes that produce a silent bot:
  • Passing vad_analyzer= to FastAPIWebsocketParams. In Pipecat 1.x the transport no longer has that field, and Pydantic silently discards it — you get no error and no turn detection. It belongs on LLMUserAggregatorParams.
  • Leaving add_wav_header at its default. A WAV header prepended to every frame is interpreted as audio by the media server, and the caller hears noise.
Pipeline order matters — transport.output() sits before context_aggregator.assistant() so what the bot actually said is what gets recorded into context:

Receiving inbound calls

Configure your Vobiz number to handle incoming calls with your Pipecat agent.
1

Open Applications

Log in to the Vobiz Console and navigate to the Applications section in the sidebar.Open Applications
2

Create an application

Click Create New Application and give it a name (for example, “Pipecat Agent”).Create an application
3

Configure URLs

Set the Answer URL to your ngrok URL (for example, https://.../answer) and select POST method. You can use the same URL for Hangup or leave it blank.Configure the Answer URL
4

Assign phone number

Go to Phone Numbers, select your number, and assign it to the application you just created.Attach your phone number
Success: Calls to your Vobiz number will now be handled by your local Pipecat server!

Quick reference

Server endpoints

The call registry is an in-memory dict. It resets on restart and does not work across multiple workers — use Redis or a database before running more than one process in production.

Project files

VobizFrameSerializer and parse_vobiz_start are not in this repo — they ship in the separate pipecat-vobiz package, which installs into the pipecat.serializers namespace. That is why bot.py imports them from pipecat.serializers.vobiz rather than from a local file.

Customizing the bot

Edit bot.py to customize your AI assistant:
Change Bot Personality
Change TTS Voice
The bot as shipped waits for the caller to speak first — on_client_connected only logs. To have it greet the caller, queue an LLMRunFrame in that handler so the LLM produces its opening line from the system prompt.

Troubleshooting

Almost always Python 3.10. pipecat-ai 1.x requires 3.11+, and on 3.10 pip installs a 0.0.x release instead of failing. Run python --version, then pip show pipecat-ai — if the version starts with 0.0., recreate your virtualenv on 3.11+.
Check VAD wiring. In Pipecat 1.x, vad_analyzer belongs on LLMUserAggregatorParams, not on FastAPIWebsocketParams. The transport no longer has that field, and Pydantic discards unknown fields silently, so the mistake produces no error at all — just a bot that never detects the end of a turn.
Set add_wav_header=False on FastAPIWebsocketParams. A WAV header on every frame is played as audio by the media server.
You almost certainly have audioTrack="both" together with bidirectional="true". That combination is not supported and fails silently. Use audioTrack="inbound".
Byte order. Set VOBIZ_L16_ENDIAN=le. The docs and RFC 2586 specify big-endian, but some accounts transport L16 little-endian. Using the default audio/x-mulaw avoids the problem entirely.
Two common causes: the sample rate is 24000 (not reliably supported — use 8000 or 16000), or PUBLIC_URL is stale. After restarting ngrok the URL changes; update PUBLIC_URL and restart the server, since the answer XML is built from it.
If PUBLIC_URL is unset the server falls back to the request Host header, which is localhost in local runs — Vobiz cannot route to that. The server prints a warning when this happens. Set PUBLIC_URL to your public HTTPS URL.
Known quirk. The answer XML requests fileFormat="wav", so Vobiz serves a .wav URL — but download_recording.py unconditionally appends .mp3 to the saved filename. The bytes are a RIFF/WAVE container (8 kHz, 16-bit, stereo); only the extension is wrong. Rename it to .wav, or change the fileFormat in the <Record> element to match the extension you want.
Integration complete!You can now make AI-powered phone calls with Vobiz and Pipecat.

Next steps

  • Customize your AI assistant’s personality in bot.py
  • Deploy to production (AWS/GCP/Heroku) instead of ngrok
  • Add custom business logic and integrations

Resources

Vobiz Documentation External Resources

Build it with an AI agent

Clone, configure, and run the Vobiz-X-Pipecat repo - your first AI voice agent in ~5 minutes.

Open in Cursor