Scaling WebRTC for Real-Time Voice AI: Inside OpenAI's Split Architecture
To achieve low-latency voice AI at a scale of over 900 million weekly active users, OpenAI bypassed traditional media termination models in favor of a custom "split relay plus transceiver" architecture. This innovative approach to scaling WebRTC separates stateless edge routing from stateful backend protocol handling, resolving Kubernetes port exhaustion while maintaining standard client compatibility.
In May 2026, OpenAI members of technical staff Yi Zhang and William McDonald published an architectural outline detailing how the company adapted its infrastructure to support real-time audio for products like ChatGPT voice and the OpenAI Realtime API. This infrastructure also powers the voice capabilities of GPT-6 Astra, which was introduced to combine advanced reasoning, computer use, and stronger judgment for complex tasks across code, applications, and research within Codex and ChatGPT Work. While initial industry speculation focused on video and WebSocket scaling, OpenAI's published technical details confirm that the core engineering breakthrough lies in scaling WebRTC audio streams to eliminate conversational latency.
The Port Exhaustion Bottleneck in Scaling WebRTC
Traditional WebRTC implementations rely on a "one-port-per-session" media termination model. In this conventional setup, every active media session requires a unique public UDP port on the host server to handle incoming and outgoing packets. While this model works well for small-scale deployments or dedicated hardware, it introduces severe operational complexity when deployed at scale within cloud-native environments like Kubernetes.
According to analysis published on Quantum Zeitgeist, maintaining a natural conversational pace—where awkward pauses, clipped interruptions, or delayed barge-ins are minimized—requires sub-second network latency. Traditional WebRTC media servers struggle to maintain this pace under massive concurrent loads due to infrastructure limitations. Managing large public port ranges in Kubernetes is operationally brittle. Platform teams must handle complex port planning, navigate uneven port utilization across nodes, and manage highly fragile rollout patterns where restarting a pod can disrupt thousands of active ports.
To address this, OpenAI evaluated and rejected several alternative architectures:
- Direct Per-Session UDP Exposure: While this preserves the standard WebRTC model, it pushes unsustainable complexity into the Kubernetes infrastructure layer. Managing tens of thousands of open UDP ports per node complicates security policies and load-balancer configurations.
- TURN-Style Relays: These relays act as intermediaries to bypass NAT restrictions, but they introduce a heavy, stateful layer directly into the media path. For OpenAI's workloads, which are predominantly 1:1 sessions between a user and a model, TURN relays solve a much wider routing problem than necessary and add unnecessary latency.
- Selective Forwarding Units (SFUs): SFUs are the industry standard for multi-party video conferencing, where they route multiple media streams among participants. However, because OpenAI's voice sessions are 1:1 interactions between a single user and an AI model, treating the AI model as a conferencing participant via an SFU introduces redundant overhead.
Inside the Split Relay Plus Transceiver Architecture
To bypass these limitations, OpenAI's engineering team decoupled the WebRTC stack into two distinct, specialized layers: a lightweight, stateless relay layer at the edge and a stateful transceiver layer in the backend. This "split relay plus transceiver" model is detailed in OpenAI's engineering notes on delivering low-latency voice AI at scale.
The Lightweight Relay Layer is deployed close to end-users at the network edge. Its sole responsibility is to accept incoming UDP packets and statelessly forward them to the backend. Because it maintains no session state, the relay layer remains incredibly fast, simple, and resilient to traffic spikes. It also minimizes public UDP exposure by acting as a single, clean entry point for client media streams.
The Dedicated Transceiver Layer resides deep within OpenAI's backend infrastructure. This layer owns the entire stateful WebRTC lifecycle, including:
- Interactive Connectivity Establishment (ICE) negotiation
- Datagram Transport Layer Security (DTLS) handshakes
- Secure Real-time Transport Protocol (SRTP) decryption and encryption
- Session lifecycle and codec negotiation
By centralizing the complex, stateful protocol machinery in the transceiver layer, OpenAI avoids duplicating this logic across various backend services or pushing custom routing logic to the client. As Eran Stiller noted on InfoQ, the core value of this pattern is the decomposition itself: preserving standard protocol behavior at the edge while scaling the system via a thin, stateless routing layer.
The client application continues to interact with a standard WebRTC endpoint, unaware of the internal split. This allows developers to build full-duplex voice applications using standard APIs, as outlined in our guide on building full-duplex voice agents with GPT-Live-1 in the API.
Client-Side Implementation and Code Example
Because the split relay plus transceiver architecture preserves standard WebRTC protocols on the client side, developers do not need to implement custom routing or proprietary SDKs to interface with the Realtime API or GPT-6 Astra. The client initiates standard ICE and DTLS negotiations, which are transparently proxied by the edge relay to the stateful transceiver backend.
The following example demonstrates a standard client-side WebRTC connection setup using native browser APIs. This setup establishes a peer connection, configures local audio tracks, and prepares to receive the low-latency audio stream from the transceiver backend:
async function initializeVoiceSession(apiEndpoint, sessionToken) {
const pc = new RTCPeerConnection({
iceServers: [{ urls: 'stun:stun.l.google.com:19302' }]
});
// Handle incoming audio track from the backend transceiver
pc.ontrack = (event) => {
if (event.track.kind === 'audio') {
const audioElement = document.createElement('audio');
audioElement.srcObject = event.streams[0];
audioElement.autoplay = true;
document.body.appendChild(audioElement);
}
};
// Acquire local microphone access
const localStream = await navigator.mediaDevices.getUserMedia({ audio: true });
localStream.getTracks().forEach(track => pc.addTrack(track, localStream));
// Create offer and set local description
const offer = await pc.createOffer();
await pc.setLocalDescription(offer);
// Send SDP offer to the backend (via the stateless edge relay)
const response = await fetch(apiEndpoint, {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': `Bearer ${sessionToken}`
},
body: JSON.stringify({ sdp: pc.localDescription.sdp })
});
const { sdp: answerSdp } = await response.json();
await pc.setRemoteDescription(new RTCSessionDescription({ type: 'answer', sdp: answerSdp }));
}This standard implementation ensures cross-platform compatibility across both desktop browsers and mobile operating systems, utilizing built-in hardware echo cancellation and jitter buffering. For more advanced implementations, developers can refer to our tutorial on building natural voice experiences with GPT-Live-1.
Architectural Trade-Offs and Recommendations for Platform Engineers
While the split relay plus transceiver architecture solves the massive scaling challenges of real-time voice AI, it introduces specific trade-offs that platform engineers must evaluate before adopting this pattern:
Increased Internal Network Latency: By separating the edge relay from the stateful transceiver, packets must make an extra internal network hop. OpenAI mitigates this by optimizing the internal routing path and placing edge relays close to users. However, for smaller organizations without a global edge network, the latency of this extra hop could outweigh the benefits of port simplification.
Operational Complexity of Two Layers: Instead of managing a single fleet of WebRTC media servers, platform teams must now operate, monitor, and deploy two distinct services. The stateless edge relays must be highly available and distributed globally, while the backend transceivers require robust state-synchronization mechanisms to handle failovers without dropping active voice sessions.
For teams building interactive media systems or real-time voice agents, the recommended path forward is to defer this level of decomposition until port exhaustion or Kubernetes routing limits become an active bottleneck. For applications operating at moderate scale, standard WebRTC architectures or managed SFU services remain highly effective. However, as real-time AI agents become a standard product pattern, adopting a split architecture will become a critical blueprint for scaling to millions of concurrent sessions.
Frequently asked questions
Why did OpenAI reject traditional TURN relays for voice AI?
TURN relays introduce a heavy, stateful intermediary into the media path. For predominantly 1:1 model-to-user sessions, they add unnecessary latency and solve a broader routing problem than required.
How does the split relay plus transceiver architecture prevent port exhaustion?
It moves stateful WebRTC operations to a dedicated backend transceiver layer. The edge relay remains stateless, allowing traffic to be routed without requiring a unique public UDP port per active user session on the host.
Does OpenAI's WebRTC architecture require custom client-side SDKs?
No. The split architecture is designed to preserve standard WebRTC behavior for client applications. Standard browser and mobile APIs can interact directly with the endpoint without custom routing logic.
When should a platform team adopt a split WebRTC architecture?
Teams should adopt this pattern when scaling 1:1 media sessions to a level where managing large public UDP port ranges in Kubernetes becomes an operational bottleneck or causes uneven resource utilization.
Sources
- OpenAI’s 4 Steps To Low-Latency Voice AI At Global Scale — Quantum Zeitgeist
- How OpenAI delivers low-latency voice AI at scale | OpenAI — How OpenAI delivers low-latency voice AI at scale | OpenAI
- OpenAI Outlines WebRTC Architecture for Low-Latency Voice AI at Scale — InfoQ
