Building Full-Duplex Voice Agents with GPT-Live-1 in the API
To build natural, low-latency voice experiences, developers can now leverage gpt-live-1 in the api. Released on September 10, 2026, this native, full-duplex voice model collapses listening and speaking into a single, cohesive engine. Unlike traditional voice pipelines that chain speech-to-text, LLM inference, and text-to-speech together, this model processes incoming and outgoing audio simultaneously, allowing agents to handle interruptions, understand backchannels, and navigate pauses in real time.
In this tutorial, you will learn how to initialize a WebRTC-based voice session using the new Live Sessions API, configure Responses delegation to offload deep reasoning to a backend model, and manage the session lifecycle using Node.js and Express. We will cover the core architecture, pricing mechanics, step-by-step implementation, and troubleshooting strategies.
What is GPT-Live-1 in the API?
As documented in the OpenAI Announcement, GPT-Live-1 represents a paradigm shift in how conversational voice applications are built. Rather than treating voice as an asynchronous, turn-based text exchange, the gpt-live-1 model operates continuously. It evaluates incoming audio frames multiple times per second to decide whether to speak, listen, pause, or yield to an interruption. According to the Unite.AI coverage, this architecture reduces turn-taking latency to an average of 0.798 seconds, compared to 1.41 seconds for legacy turn-based models like GPT-Realtime-2.1.
This capability is made possible by separating the conversational interface from the reasoning backend. The live model acts as a highly responsive, low-latency front-end layer that manages the immediate flow of speech. Complex tasks, database queries, and tool executions are delegated to a secondary backend model, such as gpt-5.6-terra or gpt-5.6-luna. This design keeps the voice conversation fluid and uninterrupted, even when the backend is performing heavy computations or waiting on external APIs.
The pricing structure has also been simplified. Instead of charging for input and output audio tokens, OpenAI bills the front-end voice layer at a flat $0.05 per minute ($3.00 per hour), calculated down to the second. Any backend reasoning, tool usage, or search operations are billed separately under their standard token rates, as detailed in the OrcaRouter Analysis.
Understanding the Split Architecture and Delegation Modes
When integrating gpt-live-1 in the api, you must choose between two delegation modes to handle complex user requests:
- Responses Delegation: The
gpt-live-1model automatically coordinates with a managed OpenAI backend model (e.g.,gpt-5.6-terra). When a user asks a question requiring deep reasoning or tool execution, the voice layer silently routes the request to the backend model, receives the text response, and synthesizes it back into the audio stream. - Client Delegation: Your application server intercepts delegation events via a sideband WebSocket connection. This allows your custom orchestration layer, agent framework, or third-party models to execute workflows and return the final text or instructions back to the active voice session.
For most standard implementations, Responses delegation is preferred because OpenAI manages the context compaction, session state, and model routing automatically. This tutorial focuses on setting up Responses delegation using the gpt-5.6-terra model as the backend reasoner.
Prerequisites and Setup
Before writing code, ensure your environment meets the following requirements:
- Node.js: Version
22.6or later is required to support modern fetch APIs and ES modules natively. - Dependencies: You will need the
openaiandexpresspackages. - API Key: A valid OpenAI API key configured in your environment variables (
OPENAI_API_KEY). Note that the Free tier cannot access this model; you must be on a paid usage tier.
Initialize a new Node.js project and install the dependencies:
npm init -y
npm install openai expressEnsure your package.json includes "type": "module" to enable ES module imports (import/export syntax).
Step-by-Step WebRTC Session Implementation
To establish a voice session from a browser, we use WebRTC. The client browser captures microphone input, generates a Session Description Protocol (SDP) offer, and sends it to our Node.js server. Our server then forwards this offer to OpenAI's Live Sessions endpoint, receives an SDP answer, and returns it to the client to establish a peer-to-peer connection.
1. The Node.js Express Server
Create a file named server.js. This server exposes a POST /api/session endpoint that handles the SDP exchange with OpenAI's /v1/live/sessions endpoint.
import express from 'express';
import { OpenAI } from 'openai';
const app = express();
app.use(express.json());
app.use(express.static('public')); // Serve frontend files
const openai = new OpenAI({
apiKey: process.env.OPENAI_API_KEY,
});
app.post('/api/session', async (req, res) => {
try {
const { sdp } = req.body;
if (!sdp) {
return res.status(400).json({ error: 'SDP offer is required' });
}
// Send the SDP offer to the OpenAI Live Sessions endpoint
const response = await fetch('https://api.openai.com/v1/live/sessions', {
method: 'POST',
headers: {
'Authorization': `Bearer ${process.env.OPENAI_API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
model: 'gpt-live-1',
transport: {
type: 'webrtc',
sdp: sdp,
},
session: {
instructions: 'You are a warm, helpful customer assistant. Keep your spoken responses brief and natural.',
voice: 'marin',
delegation: {
type: 'responses',
responses: {
model: 'gpt-5.6-terra',
instructions: 'You are the backend reasoning agent. Process the user request, execute tools if needed, and provide a clear answer.',
},
},
},
}),
});
if (!response.ok) {
const errorText = await response.text();
console.error('OpenAI API Error:', errorText);
return res.status(response.status).send(errorText);
}
const data = await response.json();
// Return the session ID and the remote SDP answer to the client
res.json({
id: data.id,
sdp: data.transport.sdp,
});
} catch (error) {
console.error('Failed to create live session:', error);
res.status(500).json({ error: 'Internal Server Error' });
}
});
const PORT = process.env.PORT || 3000;
app.listen(PORT, () => {
console.log(`Server listening on port ${PORT}`);
});2. The Client-Side WebRTC Handshake
Create a directory named public and place an index.html file inside it. This file requests microphone permissions, sets up the WebRTC peer connection, and sends the SDP offer to our Express server.
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<title>GPT-Live-1 Voice Client</title>
</head>
<body>
<h1>GPT-Live-1 Voice Agent</h1>
<button id="start-btn">Start Conversation</button>
<button id="stop-btn" disabled>Stop</button>
<p id="status">Status: Disconnected</p>
<script>
let peerConnection;
let localStream;
const startBtn = document.getElementById('start-btn');
const stopBtn = document.getElementById('stop-btn');
const statusText = document.getElementById('status');
startBtn.onclick = async () => {
statusText.textContent = 'Status: Requesting microphone...';
try {
localStream = await navigator.mediaDevices.getUserMedia({ audio: true });
peerConnection = new RTCPeerConnection();
// Add local audio tracks to the connection
localStream.getTracks().forEach(track => {
peerConnection.addTrack(track, localStream);
});
// Set up remote audio playback
peerConnection.ontrack = (event) => {
const remoteAudio = document.createElement('audio');
remoteAudio.srcObject = event.streams[0];
remoteAudio.autoplay = true;
document.body.appendChild(remoteAudio);
};
// Create WebRTC data channel for control events
const dataChannel = peerConnection.createDataChannel('events');
dataChannel.onopen = () => {
statusText.textContent = 'Status: Connected to session!';
};
// Create SDP Offer
const offer = await peerConnection.createOffer();
await peerConnection.setLocalDescription(offer);
statusText.textContent = 'Status: Exchanging SDP with server...';
const response = await fetch('/api/session', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ sdp: offer.sdp })
});
if (!response.ok) {
throw new Error('Failed to negotiate WebRTC session');
}
const data = await response.json();
// Apply the remote SDP answer from OpenAI
await peerConnection.setRemoteDescription(new RTCSessionDescription({
type: 'answer',
sdp: data.sdp
}));
startBtn.disabled = true;
stopBtn.disabled = false;
} catch (err) {
console.error(err);
statusText.textContent = `Status: Error - ${err.message}`;
}
};
stopBtn.onclick = () => {
if (peerConnection) peerConnection.close();
if (localStream) localStream.getTracks().forEach(track => track.stop());
startBtn.disabled = false;
stopBtn.disabled = true;
statusText.textContent = 'Status: Disconnected';
};
</script>
</body>
</html>A Runnable Simulation of Session Configuration
To verify that your payload structure matches the schema required for gpt-live-1 in the api, you can use the following runnable script. This script validates the configuration object offline using built-in Node.js assertions, ensuring your integration logic is structurally sound before making actual network requests.
import assert from 'assert';
function validateSessionConfig(config) {
if (config.model !== 'gpt-live-1') {
throw new Error('Invalid model ID. Must be "gpt-live-1".');
}
if (!config.transport || config.transport.type !== 'webrtc') {
throw new Error('Invalid transport. Only "webrtc" is supported for client-side browser sessions.');
}
if (!config.session || !config.session.delegation) {
throw new Error('Session delegation configuration is missing.');
}
if (config.session.delegation.type !== 'responses') {
throw new Error('This setup requires "responses" delegation.');
}
return true;
}
// Sample configuration matching our Express server setup
const samplePayload = {
model: 'gpt-live-1',
transport: {
type: 'webrtc',
sdp: 'v=0\no=- 123456 2 IN IP4 127.0.0.1\ns=-\nt=0 0\na=group:BUNDLE 0\nm=audio 9 UDP/TLS/RTP/SAVPF 111\nc=IN IP4 127.0.0.1'
},
session: {
instructions: 'Keep voice responses under 10 words.',
voice: 'marin',
delegation: {
type: 'responses',
responses: {
model: 'gpt-5.6-terra',
instructions: 'Handle background reasoning tasks.'
}
}
}
};
try {
const isValid = validateSessionConfig(samplePayload);
console.log(JSON.stringify({
status: 'success',
message: 'Configuration payload is valid for gpt-live-1 in the api.',
validationResult: isValid
}, null, 2));
process.exit(0);
} catch (error) {
console.error('Validation failed:', error.message);
process.exit(1);
}Common Errors and Troubleshooting
When migrating to or implementing the gpt-live-1 model, developers frequently encounter several distinct failure modes. The table below outlines these issues and how to resolve them.
| Error / Symptom | Root Cause | Resolution |
|---|---|---|
404 Not Found on session creation |
Attempting to use legacy Realtime WebSocket endpoints or standard Chat Completions. | Ensure you are targeting POST /v1/live/sessions and using the model ID gpt-live-1. |
| Audio drops or fails to connect | WebRTC ICE candidate negotiation failure or blocked ports. | Ensure your application server and client can negotiate media transport. In restrictive networks, configure a TURN server. |
| Extremely high billing on short calls | Misunderstanding the 15-second initialization charge. | Note that the 15-second initialization fee is credited back against actual session duration once the call starts. Avoid creating sessions that are immediately discarded without media flow. |
| Model ignores complex system instructions | Attempting to cram detailed business rules into the front-end session.instructions. |
Keep the front-end instructions limited to speaking style and tone. Move all complex business logic, tool definitions, and workflows to the backend model prompt (e.g., inside delegation.responses.instructions). |
Next Steps and Migration
If you are migrating from the legacy Realtime API (such as gpt-realtime-2.1), be aware that there is no direct drop-in path. The Realtime API relied on a persistent WebSocket connection for both control events and raw audio streams, using token-based billing. GPT-Live-1 uses WebRTC for client audio transport and bills based on session duration, requiring a complete rewrite of your media connection layer.
Additionally, because GPT-Live-1 does not emit an output-audio-done event, you must track playback queues on the client side if you need to trigger UI updates when the agent finishes speaking. To extend this application, explore using the sideband WebSocket endpoint (wss://api.openai.com/v1/live/sessions/{session_id}/attach) to monitor events and inject real-time text corrections or steering commands during an active voice session.
Frequently asked questions
Can I use GPT-Live-1 with standard Chat Completions?
No. GPT-Live-1 is not supported through standard Chat Completions, Responses, or legacy Realtime endpoints. It requires the dedicated Live Sessions endpoint at /v1/live/sessions.
How does billing work for GPT-Live-1 voice sessions?
Voice sessions cost a flat $0.05 per minute and are billed per second. A 15-second initialization charge is billed up front but credited back against actual session duration once the session begins.
What is the difference between Responses and Client delegation?
Responses delegation automatically routes complex reasoning and tools to a managed OpenAI backend model. Client delegation intercepts events, allowing your server to run custom workflows before returning results to the session.
