← back to posts

$ posts/deploying-webrtc-from-the-ground-up.md

---
date:  2026-09-23
tag:   Technical
read:  12 min
---

Deploying a real-time WebRTC app from the ground up: DNS, TLS, one VM, and a deploy button

You can’t host a WebRTC app on Cloud Run, Vercel, or Railway: their load balancers only route HTTP over TCP, and real-time audio travels over UDP, so the media path simply dies. Deploying aloud meant leaving managed platforms for a bare VM, and assembling, from the ground up, the networking knowledge that managed platforms let you skip. This is that knowledge, organized the way I wish I’d found it.

Aloud started on managed containers, and the move to real time (WebSocket to WebRTC, TCP to UDP) is what forced the migration, and it turned out to be the best infrastructure education I didn’t sign up for. Suddenly I was the one answering for DNS records, TLS certificates, firewall rules, and what actually happens between git push and a live site.

Owning a domain, actually

The problem: a browser has to find your one machine among billions using nothing but a name, and that name-to-machine mapping has to stay under your control so the name can follow the app when the machine changes. The domain name system is that mapping, and “owning” a domain is an annually renewable lease on a piece of it.

The lease runs through a delegation chain. ICANN sits at the top; each top-level domain is operated by a registry (for example, Verisign operates .com, PIR operates .org); you lease from an accredited registrar. The registry doesn’t store your actual records. It stores a pointer to your nameservers, the servers that answer “what’s the IP for this name?” Those hold the truth: the A record maps your name to an IPv4 address (for example, work-aloud.com to your VM’s static IP), CNAME aliases one name to another, MX names your mail servers.

Here’s the chain doing its job, the first time a resolver anywhere in the world looks up my domain. The recursive resolver is the middleman doing the legwork (your ISP’s, or a public one like 1.1.1.1); everything it learns is cached per TTL, so repeat lookups skip straight to the end:

browser ─▶ recursive resolver (1.1.1.1)
              │ "work-aloud.com?"
              ▼
           root server ──────▶ "ask .com's servers"      (runs under ICANN's root zone)
              ▼
           .com TLD server ──▶ "ask this domain's        (the registry; holds only the
                                nameservers: Cloudflare"  NS pointer, not my records)
              ▼
           Cloudflare nameserver ▶ "A record:            (authoritative; holds the
                                    <the VM's IP>"        actual records I set)
              ▼
           browser connects to the VM

Notice who’s missing: the registrar. It lives in the write path, not the read path. Registering the domain is what got the “ask Cloudflare’s nameservers” pointer into the .com registry, and changing nameservers goes through the registrar too, but no DNS query ever touches it. The reason the two roles blur is that most registrars are also a DNS host: the service that runs nameservers on your behalf, keeping them answering queries around the world and giving you the dashboard where you edit your records. The nameservers are the machines; the DNS host is the company operating them. That’s my setup exactly: Cloudflare the registrar wrote the pointer into the registry, and Cloudflare the DNS host operates the nameservers answering step three. Two hats, one invoice.

Once you see that chain, the operational rule writes itself. DNS is the linchpin of everything downstream: move your app to a new VM without updating the A record, and users and the certificate authority alike keep knocking on the old IP. Always DNS first, then move. The record’s TTL is how long resolvers may cache the old answer (TTL 300 means up to 5 minutes of stragglers), so lower it a day before a planned migration.

How I applied this in aloud: the domain lives at Cloudflare (~$10/yr, DNS hosting free), but with the proxy toggled off, “grey-cloud” DNS-only mode. Cloudflare’s proxy gives you CDN caching and DDoS protection by routing traffic through their network, but it only forwards HTTP. Turn it on and the WebRTC UDP path silently dies, the exact failure that rules out serverless hosts, now hiding inside a DNS setting.

TLS without tears

The problem: plain HTTP is readable and modifiable by everything between the user and your server, from coffee-shop wifi to any proxy on the path. For a voice app it isn’t even a judgment call, because browsers refuse to grant microphone access to insecure pages. No HTTPS, no product.

The solution is a dependency chain, best learned backwards from the thing you want. HTTPS is just HTTP inside a TLS connection (encrypted, and authenticated so you know you’re talking to the real server). TLS needs a certificate, a file signed by a Certificate Authority attesting “this key-holder controls this domain”. Browsers ship with a list of CAs they trust. And ACME is the protocol that automates getting a cert from Let’s Encrypt, the free nonprofit CA.

In aloud’s stack that entire chain is handled by Caddy, the reverse proxy: the one process that faces the internet and forwards each request to the right container behind it. It earns its place in two phases. Phase one, once: Caddy asks Let’s Encrypt for a cert, Let’s Encrypt replies “prove it” with a challenge token, and because the DNS A record points at this VM and Caddy is what’s listening there, Caddy serves the token, passes the challenge, and stores the cert in a Docker volume (renewing every ~60 days, silently). Phase two, every connection: the browser opens https://work-aloud.com, Caddy presents the cert, the browser checks the CA’s signature against its built-in trust list, and an encrypted channel comes up. Caddy terminates TLS, decrypting at the front door and forwarding plain HTTP to the containers behind it.

The second encryption channel

Here’s the part that’s genuinely non-obvious, and the problem this section solves: the audio never goes through Caddy, so the TLS setup above protects none of it. Caddy only carries the signaling, the SDP offer/answer handshake (Session Description Protocol: “here’s what I can do and where to reach me,” answered in kind) where the two peers exchange capabilities and network candidates. The media then flows directly between browser and backend over UDP, and it needs its own encryption:

browser ──HTTPS (TCP 443)──▶ Caddy ──HTTP──▶ FastAPI      signaling: SDP offer/answer
browser ◀════════ SRTP over UDP (keys from DTLS) ════════▶ FastAPI      media: bypasses Caddy

UDP can’t use TLS (TLS assumes TCP’s ordered, reliable stream), so WebRTC uses two cousins. DTLS runs the key-agreement handshake over UDP, and SRTP encrypts each audio packet with the agreed keys. But DTLS certs are self-signed, no CA involved, so where does trust come from? Each side’s SDP includes a fingerprint (hash) of its DTLS cert, and that SDP traveled over the CA-backed HTTPS channel. The two systems chain: Let’s Encrypt secures the signaling, and the signaling vouches for the self-signed certs that secure the audio. You configure the first; WebRTC handles the second automatically.

The VM and its firewall

The problem: an internet-facing machine starts getting scanned within minutes of receiving a public IP, and every open port is attack surface you now personally answer for, the thing managed platforms were quietly handling. So the posture is default-deny: block everything, and make each opening argue for itself.

The box is deliberately boring: one e2-small GCE instance (2 vCPU, 2 GB RAM plus 2 GB swap), Debian, a reserved static IP, which is the first billable resource and the value that goes in the A record. Even the zone is a product decision: us-west1 (Oregon), close to the primary user, protecting the app’s 3-second latency budget. Exactly three openings:

tcp:80,443   ← the web (80 only so Caddy can redirect and pass cert challenges)
udp:1-65535  ← WebRTC media (ephemeral ports; only the media listener uses them)
tcp:22       ← SSH, but ONLY from Google's IAP range; port 22 never faces the internet

Two details that took me a minute. The firewall is stateful, so outbound API calls (Deepgram, Gemini, Cartesia) need no inbound rules; replies to connections the VM initiated ride back in automatically. And SSH-over-IAP means “log in through Google’s identity-aware tunnel,” which turns “who can shell into prod” into an IAM question instead of a network question.

On the VM, four containers run from one compose file: Caddy on 80/443, the Next.js frontend on 3000, the FastAPI backend on 7860 plus the UDP media listener, and Postgres on 5432 bound to loopback only (listen_addresses=127.0.0.1; reaching it means the IAP tunnel first, then a local connection). One WebRTC-specific wrinkle: all four run with Docker’s host networking instead of the default bridge, because WebRTC binds unpredictable ephemeral UDP ports and bridge networking can’t forward ports it doesn’t know about in advance. The firewall stays the real boundary; only what it allows is reachable anyway.

The first deploy ritual

Nothing glamorous, everything instructive: SSH on, install Docker and git (the only host installs; everything else lives in images), authenticate to GitHub with a personal access token, pull the repo, fill in .env and chmod 600 it, then docker compose -f docker-compose.prod.yml up -d --build. The production compose file differs from dev in exactly two ways: it adds the Caddy service, and it drops the localhost port mappings so nothing but Caddy and the WebRTC listener face outward.

Knowing it’s up: logs and alerts

The problem: once the app runs on a VM you own, nobody is watching it but you, and you are asleep for a third of every day. Observability on a self-managed box has three separate jobs: get the logs somewhere durable and queryable, notice when the site is down, and notice when the disk is filling.

Logs first. Google’s Ops Agent on the VM tails every container’s stdout and ships it to Cloud Logging, with our structured JSON parsed into queryable fields. That last part is the reason to configure it properly: the lazier route (Docker’s gcplogs driver) ships each JSON log line as one opaque string, and you lose the ability to query by field. With parsing in place, debugging a session is one query filtered by its session_id, and exceptions with stack traces group automatically in Error Reporting. One caveat learned by reading the fine print: single-line errors without a traceback never reach Error Reporting, so counting those (like latency-budget breaches) happens in log queries instead.

Alerting is two Cloud Monitoring policies emailing me: an uptime check hitting https://work-aloud.com/healthz from multiple regions every 5 minutes, and a disk alert at 80% used. The uptime check got an accidental fire drill: it shipped with a mangled path, genuinely failed from every region, and the email arrived within minutes. An alert you’ve seen fire is worth ten you haven’t.

Shipping after day one: delivery, not deployment

Every PR runs CI on GitHub Actions: a pytest suite pinning the app’s behavioral contracts (pipeline assembly, signaling API, prompt rules, latency thresholds), ruff, a frontend build check, and a Claude review on the diff. Merging to main means the artifact is deployable. Actually deploying is a separate, manually triggered workflow.

That’s continuous delivery, deliberately not continuous deployment, and the reasoning is stage-appropriate rather than dogmatic. This early, I want to personally use a build and try to break it before promoting it, because dogfooding a voice UX surfaces improvement ideas no test suite will. Auto-deploy is earned later, when the suite has enough history to be trusted alone.

The deploy button is where the git model pays off. main is the workbench; prod is not a branch anyone develops on but a bookmark meaning “this exact commit is verified live.” The workflow takes a commit (blank means tip of main), refuses anything not on main’s history (so nothing that skipped PR review and CI can ever reach production), opens an IAP tunnel to the VM as a deliberately minimal service account (it can open the tunnel and see the VM, nothing else), and hard-resets the VM’s checkout to that commit. Reset, not pull, on purpose: a pull can only move forward, while reset makes the working copy exactly the chosen commit in either direction, which is what makes rollback the same button with an older SHA pasted in. Then it rebuilds containers, polls /healthz until the site answers, and only after verification force-moves the prod pointer and pushes an immutable deploy-YYYYMMDD-HHMMSS tag. Those tags accumulate into a permanent rollback menu (Actions logs expire after ~90 days, tags never do), and a failed deploy leaves neither a moved pointer nor a tag, so history only ever contains verified releases. Two honest caveats: database migrations don’t reverse themselves when you roll back, and in-flight voice sessions die when the backend restarts, true of every deploy.

What I’m deliberately not doing yet

Running this costs about $14-15/month for the VM, disk, and IP, plus ~$10/yr for the domain; the per-use AI providers dominate the bill at any real usage. The gaps I’m consciously accepting at demo scale, in the order I’d fix them: no database backups (pgdata lives on the VM disk; a managed Postgres is the first upgrade once user data matters), secrets in a chmod 600 env file (Secret Manager is the upgrade), no TURN relay server (users on UDP-blocking networks simply can’t connect; renting one is the fix, and the prerequisite for any multi-region story), and one VM as a single point of failure (real scale-out needs session affinity, since each voice session is an in-memory pipeline, or a managed WebRTC transport as the buy-not-build alternative). Writing the list down matters more than fixing it: each item is a decision with a trigger, not a surprise waiting.

The compressed version

Serverless routes HTTP, not UDP; real time means owning a box. DNS is the linchpin: records first, moves second, and mind the proxy toggle. HTTPS needs TLS, TLS needs a certificate, a certificate needs a CA, and ACME automates the whole chain through the reverse proxy. WebRTC media bypasses that proxy entirely and secures itself with DTLS-SRTP, vouched for over the channel you did secure. Open exactly the ports you can defend, tunnel SSH through identity instead of exposing it, ship logs somewhere durable and make sure at least one alert has actually fired. And make deploys a button that only moves the verified-live bookmark after the health check passes, so rollback is the same button pointed backwards.

EOF · back to posts