Skip to main content

Building a VoiceAI Agent on DialStack with OpenAI

Answer calls to a DialStack number with an OpenAI Realtime voice agent. A small relay server sits between the two: DialStack streams the call's audio to it over the attach WebSocket, and the relay forwards it to the OpenAI Realtime WebSocket. No OpenAI SIP setup is involved.

This is the BYO Receptionist pattern with OpenAI as the AI.

1. Set up OpenAI​

  1. Create an API key on the OpenAI platform for a project with access to the Realtime API. The relay sends it as Authorization: Bearer <key>.
  2. Pick the agent's model, voice and instructions. Unlike a hosted agent, a Realtime session is configured by the client, so these live in the relay's environment (below), not in an OpenAI dashboard.

2. Set up DialStack​

Create a Voice App whose webhook URL points at your relay, and route a number or dial-plan node to it. See Voice Apps and the Routing section of the BYO guide.

3. Write the relay​

The relay does three things: it answers the webhook with an attach action, accepts DialStack's media WebSocket, and pipes audio between that socket and OpenAI. Both sides speak G.711 μ-law at 8 kHz, so the audio passes through unchanged.

JavaScript
import http from 'node:http';
import express from 'express';
import WebSocket, { WebSocketServer } from 'ws';
import { DialStack, MediaStream } from '@dialstack/sdk-server';

const dialstack = new DialStack(process.env.DIALSTACK_API_KEY);
const app = express();

// 1. A call reached the Voice App: attach it to our media endpoint.
app.post('/webhook', express.raw({ type: 'application/json' }), async (req, res) => {
const event = DialStack.webhooks.constructEvent(
req.body,
req.header('x-dialstack-signature'),
process.env.VOICE_APP_WEBHOOK_SECRET
);
if (event.event === 'call.received') {
await dialstack.calls.update(
event.call_id,
{ actions: [{ type: 'attach', url: 'wss://relay.example.com/media' }] },
{ dialstackAccount: event.account_id }
);
}
res.status(200).end();
});

// 2. DialStack opens the media WebSocket: bridge it to OpenAI Realtime.
const server = http.createServer(app);
new WebSocketServer({ server, path: '/media' }).on('connection', (ws) => {
const call = new MediaStream(ws);
const openai = new WebSocket('wss://api.openai.com/v1/realtime?model=gpt-realtime-2.1', {
headers: { Authorization: `Bearer ${process.env.OPENAI_API_KEY}` },
});

openai.on('open', () => {
openai.send(
JSON.stringify({
type: 'session.update',
session: {
type: 'realtime',
instructions: 'You are a friendly phone receptionist. Keep answers short.',
audio: {
input: {
format: { type: 'audio/pcmu' },
turn_detection: { type: 'server_vad', interrupt_response: true },
},
output: { format: { type: 'audio/pcmu' }, voice: 'marin' },
},
},
})
);
openai.send(JSON.stringify({ type: 'response.create' })); // greet the caller
});

// Agent audio goes out in 20 ms frames (160 bytes of μ-law).
let queued = Buffer.alloc(0);
const pacer = setInterval(() => {
if (queued.length < 160) return;
call.sendAudio(queued.subarray(0, 160).toString('base64'));
queued = queued.subarray(160);
}, 20);

call.on('audio', (frame) => {
if (openai.readyState !== WebSocket.OPEN) return;
openai.send(JSON.stringify({ type: 'input_audio_buffer.append', audio: frame.payload }));
});

openai.on('message', (data) => {
const msg = JSON.parse(data.toString());
if (msg.type === 'response.output_audio.delta') {
queued = Buffer.concat([queued, Buffer.from(msg.delta, 'base64')]);
} else if (msg.type === 'input_audio_buffer.speech_started') {
queued = Buffer.alloc(0); // the caller barged in: drop what the agent hadn't said yet
}
});

const hangUp = () => {
clearInterval(pacer);
openai.close();
call.close();
};
call.on('close', hangUp);
openai.on('close', hangUp);
});

server.listen(8080);

The voice-ai-agent example in the DialStack SDK is the same relay, with error handling, drift-free pacing and logging. It also covers the other vendors in this section. To run it:

Bash
git clone https://github.com/dialstack/dialstack-sdk.git
cd dialstack-sdk
npm install && npm run build # the example links the local server package
cd examples/voice-ai-agent
npm install
cp .env.example .env

Set these in .env:

VariableValue
PUBLIC_URLThe public HTTPS URL of the relay (for example an ngrok URL).
DIALSTACK_API_KEYYour DialStack secret key.
VOICE_APP_WEBHOOK_SECRETThe signing secret of the Voice App.
OPENAI_API_KEYThe OpenAI API key.
OPENAI_MODELOptional. The Realtime model. Defaults to gpt-realtime-2.1.
OPENAI_VOICEOptional. The agent's voice. Defaults to marin.
OPENAI_INSTRUCTIONSOptional. The agent's instructions (its system prompt).

Then start it and call the number:

Bash
npm run dev -- --provider openai

The agent speaks first: once OpenAI confirms the session configuration, the relay asks for a response, so the caller hears a greeting without having to say anything.

Audio formats​

DialStack's media WebSocket carries μ-law at 8 kHz. OpenAI Realtime accepts and produces G.711 μ-law directly (audio/pcmu), so the relay sets that format for both input and output in its session.update and passes the audio through untouched in both directions. No decoding and no resampling.

Barge-in and hang-up​

  • The relay turns on OpenAI's server-side voice activity detection with interrupt_response, so when the caller starts speaking OpenAI cancels the response in progress. On input_audio_buffer.speech_started the relay also drops the agent audio it has queued but not yet played, so the caller isn't talked over.
  • When the caller hangs up, the relay closes the OpenAI socket. When the OpenAI session ends, the relay closes the DialStack media socket. That ends the attach, and DialStack runs the next action, if you chained one after it.

Transfer to a human​

When the agent decides the caller needs a person, send the call a transfer. It replaces the running attach: DialStack closes the media socket, the relay closes the OpenAI session, and the caller is connected to the target.

JavaScript
await dialstack.calls.update(
callId,
{ actions: [{ type: 'transfer', target: '100' }] },
{ dialstackAccount: accountId }
);

See Replacing Actions.