Aloud is a voice-first work assistant: you tap one button and your work becomes a conversation with someone who handles the tiny details. Reading becomes an audiobook with an expert inside; it can summarize, recite, or answer questions about what you would otherwise be reading. Writing becomes letting your train of thought run out loud while an assistant puts it down coherently: say “write that up” and a structured document appears on screen while the conversation keeps moving, ready for you to go over a specific line or phrasing, as hands-on or as high level as you want to be. So you can work while out and about. It’s live at work-aloud.com.
What’s inside
- Architecture in one paragraph: the whole system in one breath, with the pipeline diagram.
- The vocabulary: six words (frame, turn, step, session, hot path, barge-in) that give a microscopic view of the pipeline, each built on the last.
- The cascade pipeline: why a chain of specialized models (Deepgram Flux to Gemini to Cartesia) beats one speech-to-speech model, and how Pipecat runs the conveyor belt.
- The network problem: sub-second audio means WebRTC over UDP, what that rules out, and why hosting moved to a bare VM.
- The agent loop at speech latency: a hand-rolled ~200-line loop, what counts as the hot path, and the async tricks that keep tools off it.
- The memory layer: automated context engineering, with deterministic machinery first and LLM judgment only where it earns it.
Two pieces got their own ground-up deep dives instead of sections here: the cloud infrastructure (DNS, TLS, the VM, the deploy button) and auth (sessions, JWTs, verification seams).
Architecture in one paragraph
Your browser captures mic audio and streams it over WebRTC to a FastAPI backend on a single Google Compute Engine VM, where a Pipecat pipeline cascades it through Deepgram Flux for speech-to-text and turn detection, a custom agent loop around Gemini 2.5 Flash for the thinking and tool calls, and Cartesia Sonic-3 for text-to-speech. First audio typically lands 1-1.5 seconds after you stop talking, streaming back over the same WebRTC connection. Conversation state lives as text on the server; artifacts and transcripts land in Postgres on the same VM; a Next.js frontend, Caddy for HTTPS, and a Firebase auth wall complete the box, all in Docker.
The rest of this post walks the latency-critical parts: the pipeline, the network, the agent loop, and the memory design.
browser mic ─WebRTC→ Deepgram Flux ─→ agent loop (Gemini 2.5 Flash) ─→ Cartesia Sonic-3 ─→ speaker
(STT + native (thinking off, tool calling, (~150ms to
end-of-turn) streams text as it thinks) first audio)
The vocabulary
Six words carry the rest of this post, and each one is built from the ones before it.
A frame is the atom. Everything that moves through the system is a typed message called a frame: a 20ms chunk of mic audio, one transcribed word, one LLM token, the event “user started speaking”. The pipeline is frames flowing through processors.
A turn is one exchange built from many frames: you speak, the system decides your thought is finished (end-of-turn detection), and the agent responds until it’s done or you cut it off.
A step is one LLM call inside the agent’s half of a turn. A response can take several steps: the model speaks, calls tools, reads the results, speaks again. A step that comes back as pure text, no tool calls, is what ends the turn.
A session is one continuous conversation, from tapping Talk to hanging up: many turns, one dedicated pipeline instance, one row in the database.
The hot path is the span from your end-of-speech to the agent’s first audio. It’s the only latency a human actually feels, and the entire design bends around it.
Barge-in is you interrupting the agent while it’s speaking. The pipeline has to go silent immediately, and then reconcile the fact that the agent said less than it generated.
The cascade pipeline
Voice agents come in two shapes: one multimodal speech-to-speech model (OpenAI Realtime, Gemini Live), or a cascade of specialized stages. Aloud is a cascade, and the deciding reason is that conversation state stays as text on the server. Text state is what makes providers swappable and a real memory layer possible. Audio-native models keep their state inside the model, where you can’t engineer it.
Pipecat is an open-source Python framework for building real-time voice and multimodal agents: the streaming problems every voice app shares (ordering, backpressure, provider adapters, interruption handling) come already solved at the framework level, so the app supplies only its own logic. In this stack it’s the chassis, a conveyor belt for frames with each stage an async processor connected to its neighbors by queues. On a barge-in, an interruption frame races down the belt flushing every queue: Gemini stream dropped, TTS stopped, queued audio discarded, silence in ~100ms. Every provider hides behind a factory function in one file, so when a better STT or a cheaper LLM ships, the swap is one function.
Each stage was picked for being fastest in its class at its one job. Gemini Flash runs with thinking pinned off on every call (the raw API defaults it on); the runner-up LLM is Claude Sonnet, and the swap is one factory. A sanitizer sits in front of Cartesia stripping markdown and identifiers from the token stream, since a voice model reads asterisks out loud if you let it.
The LLM stage is the one part that has been rebuilt repeatedly, each time for a reason: first a single call with the entire context stuffed in, then a real agent loop with multi-step tool calls and reasoning, then a growing tool set, then custom context-window management, with the memory layer as the next step. That arc is the agent loop section below.
The network problem
Humans read a response gap over ~800ms as unnatural, so the transport can’t waste any budget. The target user is on a phone, walking, which raises the bar further. WebSockets ride on TCP, which retransmits lost packets. That’s fatal for live audio or any live streaming, where a late packet is worse than a missing one. WebRTC rides on UDP and bundles everything that makes UDP sound smooth: the Opus codec (~10x compression with no audible loss), a jitter buffer to reorder uneven arrivals, packet-loss concealment to synthesize over gaps, adaptive bitrate for flaky networks. It also brings the browser’s echo cancellation, which barge-in quietly depends on: without it, the agent hears its own voice through your mic and interrupts itself. Even the VM’s location is a latency decision (us-west1, close to the primary user, protecting the 3-second budget).
Latency then gets managed like a product requirement: per-stage budgets, structured per-turn breakdowns in the logs, WARN above 1s per stage, ERROR above 3s from end-of-speech to first audio. Typical measured: 1-1.5 seconds, with the first streamed sentence going to TTS while Gemini is still generating the rest.
The WebRTC choice is also why the hosting changed. Aloud started on managed containers, and going real-time (TCP to UDP, WebSocket to WebRTC) forced the move to a personal VM, because serverless load balancers (Cloud Run, Vercel, Railway) only route HTTP and the media path simply dies. Everything that decision dragged in, from DNS and TLS certificates to the default-deny firewall, the first SSH deploy, and why releases are a button I press rather than a push hook, is its own ground-up post: Deploying a real-time WebRTC app from the ground up.
The agent loop at speech latency
An agent loop is just a while-loop around an LLM call: build messages, call the model with tools, stream text out, execute tool calls, append results, repeat until a text-only (no tool call) step ends the turn. Pipecat ships a built-in one-lap version. Aloud’s document workspace (create, list, read, and edit artifacts as agent tools) needed real control over how many laps run, what enters context, and whether speech and tools interleave, so I replaced that single pipeline stage with my own ~200-line loop. Same frames in, same frames out; the rest of the pipeline never noticed.
The loop’s mechanics are all bounded. Tool calls within a step execute concurrently, and a 10-second timeout just becomes a tool result the model can react to. A turn caps at 5 steps plus one forced wrap-up call with tool selection forbidden, because a “wrap up now” instruction alone can still come back as a tool call, and then your cap has no teeth. Even the session greeting is one tool-forbidden step with no user turn; without a named entry point for it, the session opens in silence.
The organizing rule falls straight out of the vocabulary: the hot path contains exactly one LLM call, the one that produces speech. Everything else is deterministic code or happens where nobody is waiting:
- Speak while working. The model’s first streamed sentence goes to TTS immediately (“let me write that up…”), tools execute while it’s talking, and results feed the next utterance. First-word latency becomes independent of tool-chain length.
- The user’s speaking time is free compute. While they talk, the system can pre-emptively fetch: a memory retrieval is an embedding lookup on the partial transcript (tens of milliseconds, no LLM call), with results injected before the turn’s main call.
- Writes never block audio. Transcripts and per-turn usage metrics are recorded by observers tapped onto the pipeline, off the belt. Even token counting is a local character heuristic, because a provider
count_tokensAPI is a 50-200ms network call you’d be paying on a log line. - Guarantees live in code, not prompts. “Say something before slow tools” fails; Flash routinely emits a tool call as its entire first chunk. The loop itself speaks a canned filler line (“let me look at that…”) if nothing has been said when a tool round starts.
The rewrite rested on two assumptions that the documentation couldn’t verify, so before any loop code I ran two spikes, agile jargon for a small throwaway experiment built to answer one risky question before the design commits to the answer. The first asked whether a foreign processor could sit in Pipecat’s LLM slot without breaking TTS streaming and barge-in. It could, but the interruption mechanism turned out to be a method hook rather than a frame you can watch for. The second asked whether Gemini’s raw streaming API behaves the way the loop needs, and it changed two decisions: function calls arrive as a single atomic chunk rather than token by token, which is why the filler line can fire the instant a call arrives, and the raw API defaults thinking to on (Pipecat’s wrapper had been turning it off silently), so my client pins it off on every call.
Coding harnesses like Claude Code face the same discipline with a different denominator. They optimize cost-per-task with prompt caching, parallel tool calls, and context compaction (a cached prefix also cuts time-to-first-token, which is why caching matters even when money doesn’t). A voice harness optimizes time-to-first-audio, so it trades differently: a canned filler line over a smarter silence, a step cap that forces a wrap-up call, and a barge-in policy that cancels in-flight reads but lets writes finish, because a half-done write is worse than a moot one.
Model selection runs backwards here too. The usual playbook, starts with the best model in every seat, establishes baseline quality, then downgrades seats one by one to find the cheapest configuration that still clears the quality bar. A voice agent inverts the axis: I started with the fastest capable model, Gemini Flash, accepting that fastest usually means less intelligent. Once the loop is stable and the latency budget is instrumented end to end, the plan is to walk the ladder the other way: swap in increasingly intelligent models and find the smartest one that still reliably lands inside the target latency.
Two context rules from the rewrite that I now consider non-negotiable for any tool-calling agent: every tool round appends the assistant side (its text and tool calls) before the results, since orphan tool results are unsendable on both Claude and Gemini, and every issued tool call gets a terminal result, even cancelled ones. And after a barge-in, the context must record what the user actually heard (the sentences that reached the speaker), not what the model generated, or the agent believes it said things nobody heard.
All of this is instrumented without touching the hot path. Every write (transcripts, per-step timings, LLM usage, one trace row per model call) goes through a shared background batch writer: the hot path enqueues, batches flush about once a second, and a failed write logs and drops rather than ever blocking audio. The result surfaces in an admin dashboard with p50/p95 latency, budget-breach counts, estimated spend, and a per-session per-turn drill-down that shows usage but never content. Teardown gets the same care as the hot path: the client polls a liveness endpoint every 5 seconds so a crashed backend drops the UI to a “connection lost” state instead of a frozen one, a graceful shutdown says goodbye over the data channel before cancelling pipelines, and sessions orphaned by a process death are swept closed at next boot with their usage inferred.
The memory layer
The biggest open engineering problem, deliberately deferred but designed for: cross-session memory in the MemGPT tradition, meaning session summaries (topics, tone, themes) and cross-session patterns (recurring topics, sentiment trends, notable shifts). The design principle I’ve committed to is that most context management is deterministic machinery, not agent cognition. Token counting, eviction triggers, summarization scheduling, and retrieval run as plain code outside the agent, and only a thin layer surfaces as agent-visible tools (“remember this,” “search my memory”). LLM judgment is reserved for the steps that need it, writing the summaries and spotting the trends, and those run as background jobs on cheap models, where latency doesn’t matter. The vectors behind “search my memory” will live in the same Postgres via the pgvector extension: semantic search becomes an ORDER BY over rows that obey the same user-scoping as everything else, with no second datastore to sync. At personal-app scale, an exact scan over a few thousand vectors is ~1ms with perfect recall, so the fancy index waits until the data justifies it.
The loop was built against a one-function seam, build_messages(), precisely so a custom context window with full control (sections, budgets, compression) can slot in behind it without touching loop control flow. That was the real architecture lesson of the whole project: build the seam, defer the feature.
EOF · back to posts