What Is WebRTC? Real-Time Audio and Video in 2026

What Is WebRTC? Real-Time Audio and Video in 2026
Blog 11 min read

What WebRTC is, how it works, why SFU and TURN matter, what changed in 2026, and how to host it in Europe.

WebRTC is the browser stack behind Google Meet, Zoom’s web client, and most productized voice rooms. Cameras and microphones work without plugins, without Flash, and without asking users to install a desktop app first. For SaaS teams that sell meetings, telehealth consults, or live collaboration, that native path is the difference between a feature that ships and a support queue that never ends.

This article explains what WebRTC actually does on the wire, why peer-to-peer stops scaling past a handful of people, how STUN and TURN decide whether a call connects, and how to host an SFU on a European Cloud VPS without guessing at UDP and traffic. If you are still choosing between media and messaging stacks, read WebRTC vs WebSocket. For long-lived application data without cameras, see What is WebSocket.

What WebRTC is and why it matters

WebRTC (Web Real-Time Communication) is a set of browser APIs and network protocols for interactive audio, video, and data. The browser captures media, negotiates a session, encrypts every packet, and adapts to loss and jitter. Your product owns the UX and the backend; the browser owns the hard parts of capture and playback.

Before WebRTC, web video meant plugins, proprietary clients, or “please download our app.” Meet-class products killed that pattern: open a link, grant mic/camera once, join. That expectation now applies to any SaaS that promises “talk in the browser.” If your stack cannot deliver that path, competitors who can will feel faster even when your feature list is longer.

WebRTC is not a single server product. It is a client capability plus a negotiation model. Production systems almost always add signaling, an SFU for groups, and TURN for hostile networks. Ignoring those pieces is how demos work on office Wi‑Fi and fail on mobile data.

How WebRTC works in plain language

Four ideas cover most of what operators need to reason about failures.

getUserMedia

getUserMedia asks the browser for microphone and camera streams (and optionally screen share). Permissions, device labels, and echo cancellation live here. If capture fails, no amount of SFU tuning helps — fix permissions and device selection first.

RTCPeerConnection

RTCPeerConnection is the session object. It holds local and remote media tracks, encryption state, and ICE connectivity. One connection can carry several tracks; in SFU designs each participant typically has a connection toward the media server rather than a full mesh to every peer.

SDP

Session Description Protocol (SDP) is the text “offer/answer” that lists codecs, directions, and media lines. Clients exchange SDP through your signaling channel — HTTPS, WebSocket, or another control plane you control. WebRTC does not define signaling transport; that is intentional and easy to get wrong if you treat SDP as optional metadata.

ICE

Interactive Connectivity Establishment (ICE) finds a working path between endpoints. Candidates may be host (local), server-reflexive (via STUN), or relay (via TURN). Trickle ICE sends candidates as they appear instead of waiting for a full gather — enable it, or users sit on “Connecting…” longer than necessary.

End-to-end, the flow is: capture → create offer → send SDP over signaling → set remote answer → gather/trickle ICE → DTLS handshake → SRTP media. When something breaks, decide whether you failed at capture, signaling, ICE, or media quality. Those are different tickets.

Peer-to-peer reality and why groups need an SFU

Two browsers on friendly networks can send media peer-to-peer. That mesh works for 1:1 calls and tiny rooms. Past a few participants, every client uploads a copy of its stream to every other client. Upload bandwidth and CPU collapse first; support tickets follow.

An SFU (Selective Forwarding Unit) sits in the middle. Clients upload once; the SFU forwards encrypted packets to subscribers without fully decoding and re-encoding every frame (unlike a classic MCU mix). CPU still matters — crypto, packet routing, and simulcast selection are real work — but the topology matches how products actually grow.

  • Mesh — fine for 1:1 or very small rooms; each client uploads N−1 streams
  • SFU — default for products: LiveKit, mediasoup, Janus, Jitsi Videobridge, and similar
  • MCU — mixes into one layout; heavy CPU; use when you must (legacy SIP bridges, fixed composites)

Most roadmaps should assume SFU + TURN from day one. Retrofitting TURN after launch week is more expensive than provisioning it with the first node.

STUN vs TURN: when calls stick on “Connecting”

STUN helps a client discover its public address as seen from the internet. That is enough when both sides can open a UDP path. Many corporate networks, carrier-grade NAT, and hotel Wi‑Fi block that path. Then ICE needs a relay.

TURN relays media through a server you control. Relayed minutes cost bandwidth (often roughly double the one-way path) and need open UDP/TCP listeners with a correct public IP advertisement. Coturn is common; some SFUs embed TURN. Time-limited credentials are mandatory — never ship a static relay password to the browser.

Calls stuck on connecting are usually ICE failures, not “the codec is wrong.” Checklist when that happens:

  • Is UDP media reachable on the public IP, not only on a private Docker network?
  • Does Coturn (or equivalent) advertise the right external-ip?
  • Can a phone on mobile data complete Trickle ICE with a relay candidate?
  • Is signaling actually delivering SDP and candidates both ways?

Lab laptops on open Wi‑Fi hide these failures. Always test from a restrictive network before you call the stack production-ready.

DataChannel versus media

WebRTC also offers DataChannels for arbitrary messages over the same ICE/DTLS session. They are useful for game state, file chunks, or control messages tightly coupled to a call. They are not a free replacement for application WebSockets: congestion control, reliability modes, and ops tooling differ. For chat, presence, and product events outside a call, a normal WebSocket (or HTTPS) control plane is usually clearer. Keep DataChannel for call-adjacent data; keep product messaging on the patterns in What is WebSocket.

2026: simulcast, SVC, AV1/Opus, WebTransport and MoQ

Simulcast and scalable video coding (SVC) let the SFU forward a lower layer when a subscriber is weak, without re-encoding the whole room. That is still the practical way to keep group calls usable on mixed networks.

Codecs keep moving: Opus remains the audio workhorse; AV1 and modern video profiles show up more often in browsers for quality-per-bit. Your SFU and client SDKs must agree on what is offered in SDP — “AV1 everywhere” is not a hosting decision alone.

WebTransport (QUIC/HTTP3) and Media over QUIC (MoQ) matter for low-latency client–server fan-out and broadcast-style delivery. They do not retire WebRTC for two-way meetings and voice rooms. Plan hosting so UDP to your European node works for ICE/TURN today; treat MoQ as an additional path for one-to-many live when your product needs it, not as a reason to delete the SFU.

Languages and SDKs

WebRTC is a W3C/IETF standard, not a single language runtime. Client and server stacks differ:

  • JavaScript / TypeScript — browser Web API (getUserMedia, RTCPeerConnection). Default for web products; documented on MDN and webrtc.org.
  • Node.js — mediasoup and similar libraries for SFU workers; signaling often shares the same Node process as your app.
  • Go — Pion WebRTC and LiveKit’s SFU; popular for high-pps media servers.
  • Python — aiortc for bots, AI participants, and tooling (not the usual choice for a multi-tenant SFU).
  • Java / Kotlin and Swift — native Android and iOS clients using platform WebRTC builds.
  • C++ — Google’s reference libwebrtc; embedded and high-performance media paths.
  • Rust — webrtc-rs for safe systems services.
  • C# / .NET — desktop and server wrappers around native WebRTC.

Beyond meetings, WebRTC also shows up in niche streaming clients — for example NVIDIA Isaac Sim’s browser streaming path uses WebRTC for interactive sim views (Isaac Sim WebRTC client guide). The protocol is the same; the product is not a Zoom clone.

Install and operating systems

Clients: Chrome, Firefox, Safari, and Edge ship WebRTC on Windows, macOS, Linux, Android, and iOS. End users do not install a plugin. Mobile apps link a native WebRTC SDK instead of relying only on the in-app browser.

Servers: Production SFU and Coturn deployments are overwhelmingly Linux (Ubuntu/Debian). Windows Server can host some stacks, but Docker images, UDP tuning docs, and community runbooks assume Linux. Typical first steps on Ubuntu:

sudo apt update
sudo apt install -y coturn
# SFU example: follow your vendor’s Docker guide (LiveKit, mediasoup workers, Janus)
# Open UDP media range + 3478/443 for TURN on the public interface

LiveKit-style nodes commonly need TCP/UDP 443, TCP 80 for certificates, a TCP fallback port, UDP 3478, and a wide UDP media range on the host network — not only inside a private Docker bridge. Prefer host networking for media containers.

Latency numbers worth believing

Do not trust a “faster protocol” claim from a raw MB/s bench that never opens a camera. Use industry latency buckets instead. OpenVidu’s overview of low-latency live streaming places videoconferencing and cloud gaming in the sub‑100 ms interactive tier (conversations degrade past roughly ~150 ms), while Low-Latency HLS/DASH typically lands around 2–5 seconds, and classic HLS/DASH with multi-second segments often sits in the tens of seconds once playlist buffering is included (WebRTC vs HLS and DASH, Part 1). WebRTC wins when a human must react in the loop; HLS/DASH win when you need CDN-scale one-way delivery and can tolerate delay.

Security defaults and the gaps you still own

Media on WebRTC is encrypted by design: DTLS for key exchange and SRTP for the media path — browsers do not offer an “unencrypted RTP” switch. That foundation is summarized well in WebRTC Security: How Safe Is It? (DEV). Gaps are usually in your implementation:

  • Signaling over plain ws:// instead of wss:// (MITM on SDP/ICE)
  • P2P host candidates leaking IP addresses when anonymity matters — SFU topologies hide peer IPs from each other
  • Weak app auth: encrypted media does not stop someone who stole a room token

Require WSS, short-lived TURN credentials, and server-side authorization for publish/subscribe. Camera/mic access still goes through browser permission prompts — that layer is not optional.

Self-host versus managed RTC

Managed RTC SaaS buys time: SDKs, global edges, less pager duty. Self-hosting buys control: data residency, predictable unit economics at scale, and the ability to keep media inside your VPC or region story. Many teams start managed, then move hot rooms or EU-only tenants onto their own SFU when invoices or compliance force the issue.

Self-host cost is not only the VPS line item. You own Coturn credentials, certificate renewal, UDP firewall rules, version upgrades, and load tests. If nobody on the team can read ICE stats, stay managed longer. If you already run Kubernetes or disciplined Linux fleets in Europe, an SFU node is a normal service — size it like a network appliance, not like a WordPress box.

Hosting WebRTC in Europe: latency, UDP, and traffic

Interactive media is sensitive to RTT and loss. Hosting in Europe keeps EU users on a shorter path than bouncing media through another continent “because the API gateway lives there.” Put the SFU and TURN close to participants; keep chat databases wherever your compliance model requires, but do not force every RTP packet through the wrong region.

WebRTC prefers UDP. A pretty dashboard with closed UDP ranges is still a broken call. Budget for:

  • Open UDP media ports on the public interface (ranges vary by SFU; examples include 10000–20000 or 50000–60000)
  • TURN on UDP/TCP and often TLS on 443 for locked-down clients
  • Sustained egress — HD rooms and relayed TURN minutes eat monthly traffic allotments faster than static sites

On EuroVDC Cloud Server plans, practical baselines for a European node:

  • Cloud VPS Medium Plus — 4 vCPU / 8 GB — labs, 1:1 calls, small rooms, Coturn + light SFU; catalog around €66.99/mo (ex-VAT; confirm on the product page)
  • Cloud VPS XL Plus — 8 vCPU / 24 GB — production SFU + TURN for moderate concurrency; catalog around €169.99/mo
  • Cloud VPS XXL — 12 vCPU / 32 GB — denser rooms or simulcast-heavy loads

Prefer compute-heavy shapes over tiny RAM-starved instances. When a single VPS saturates packets-per-second or the traffic quota, split TURN from the SFU, then consider dedicated capacity for sustained multi-gigabit fan-out. Soft path check from our Sofia datacenter: use the Looking Glass before you promise regional latency in sales decks.

Operator checklist

  1. Public IPv4 (and IPv6 if you advertise it), DNS for app + TURN hostnames, trusted TLS — browsers reject casual self-signed setups on public sites
  2. Open the UDP media range end-to-end; verify from mobile data, not only the office LAN
  3. Time-limited TURN credentials; rotate and scope by room or user
  4. Advertise the correct public IP to Coturn / TURN (external-ip); host networking for Docker SFUs beats nested NAT
  5. Raise UDP buffers if you see unexplained loss under load:
net.core.rmem_max = 26214400
net.core.wmem_max = 26214400
net.core.rmem_default = 26214400
net.core.wmem_default = 26214400
  1. Enable Trickle ICE; collect ICE stats (RTT, candidate type host / srflx / relay)
  2. Offer simulcast or SVC so the SFU can downshift without re-encoding
  3. Load-test reconnect storms and TURN-heavy clients; size traffic before launch week

Typical port surface for a LiveKit-style node (adjust to your stack): TCP/UDP 443, TCP 80 for certificates, SFU TCP fallback if used, UDP 3478 for TURN, plus the UDP media range on the public IP.

Related guides

Order capacity on Cloud VPS, open the UDP path, deploy SFU + TURN, and test from a restrictive network. For application messaging without media, continue with What is WebSocket. For a side-by-side decision matrix and hybrid pattern, read WebRTC vs WebSocket.

References

webrtc real-time sfu turn video audio

EuroVDC

Discover the Power of Cloud

KVM virtualization · Instant scaling · 24/7 support

Rent a Cloud Server

Did you find this content useful?

– People found it useful

Share on Social Media

What Is WebRTC? Real-Time Audio and Video in 2026