Dual Subtitle Player for English Lectures
Bilingual captions that keep English lectures learnable
English, not difficulty, blocked the lectures
An English-speaking instructor taught quant trading and the students were Korean, so the lectures needed Korean subtitles.
There were three requirements. First, the source and the translation have to be on the same screen. Second, it has to work the same way outside the course VOD, on YouTube videos and live Q&A sessions that have no caption track at all. Third, watching the same lecture several times should run the translation only once.
The first version was a single VOD player for the course. The second requirement meant it could not stay tied to one player: it had to be a subtitle layer that sits on top of anything.

One Worker, everything else client side
One Durable Object per live session, one socket, hibernation off
The upstream STT socket has to stay open between client audio frames. A Worker dies with the request, so there is nowhere to hold that state, and a DO fits the slot exactly. Giving up hibernation means paying for as long as the session is alive. In exchange the object is created with newUniqueId() so nothing is persisted, and when stop arrives it waits at most 4 seconds for translations in flight before cutting the connection.
STT behind an adapter interface, swapped with the STT_PROVIDER variable
Real-time STT changes price and quality every quarter. ElevenLabs Scribe v2 Realtime runs by default, and putting Deepgram and a mock behind the same connect/onFragment/onError/onClose contract means swapping any of them touches nothing else in the pipeline. The whole pipeline had to run e2e without keys, so the mock adapter was not optional, it was required.
The extension calls the engine from the background service worker, not the content script
A content script's fetch goes out with the page origin. A request from https youtube.com to http://127.0.0.1:8787 is blocked by Chrome's local network access rules. The service worker goes out on the extension's host permission, so that wall is not there. It meant writing a message round-trip wrapper (engineFetch) that mimics the fetch signature exactly, and it only supports requests whose body is a string.

Read the captions, otherwise listen
The behavior forks in two. When a caption track exists (Tier 1) the whole track is pulled up front, translated in one pass, and cached, and during playback requestAnimationFrame only lines the cues up with the clock. That is why the added latency is zero. When there is no track (Tier 2), the audio becomes the source. The extension takes the stream from tabCapture and an AudioWorklet in an offscreen document cuts it into 16 kHz PCM16 frames of 100 ms, and the macOS app cuts the system audio ScreenCaptureKit hands it into the same shape. On the receiving end the Durable Object splits on sentence boundaries, translates, and sends the latency back with each one. This piece covers only that live path and the design of the translation cache key. Range serving for the private R2 media and the study features in the web player are each worth their own write-up, so I am leaving them out here.

Committing a sentence without waiting for silence
The first time you wire up streaming STT you usually do this: throw the partials away, collect only what the recognizer marks final (committed), group those into sentences, then translate. It is what the docs say and it looks safe. But Scribe's commit strategy is VAD. The commit arrives when the speaker pauses. On a lecture or a talk, where nobody pauses, five to ten seconds of text lands at once, and until then the screen holds either nothing or a dim intermediate line. As a subtitle it is unusable.
I picked out of the partial stream the sentences that are already settled in everything but name, and emitted those first. completeSentences looks for the last sentence-ending punctuation that already has the next word attached behind it. Everything up to that point is cut off and promoted immediately as isFinal: true, speechFinal: true, and only the remaining tail goes out as a partial. When the real committed text arrives later, the prefix already emitted (emittedPrefix) is stripped off and only the rest flows through, so nothing is duplicated. On top of that the segmenter cuts again on four conditions (sentence-ending punctuation, 0.6 s of silence, 80 characters, 1.5 s since the start), and the clock that condition 4 needs is fed by the DO extrapolating the last fragment's timestamp on a 250 ms setInterval.
A partial keeps wobbling in the middle of a sentence, but if the next word is already attached behind a sentence, the recognizer has no reason to touch that sentence again. That one observation deletes the entire wait for VAD silence. Cutting on sentences is a requirement too, not a preference: English into Japanese or Korean flips the word order, so translating fragments produces nonsense. On a 20-second streaming measurement with real ElevenLabs and OpenRouter wired in, the time from sentence commit to translation on screen came out at p50 1101 ms, p95 1132 ms.
A cache key that is a fingerprint of the content, recomputed and checked by the server
The first cut of a translation cache gives you keys like cues/<videoId> or <videoId>-en-ko.
I defined the key as a hash of identity plus content. hash = sha256(platform, videoId, srcLang, dstLang, model, contentHash), where contentHash is a hash of the source cues' count + start times and source text concatenated.
Because the key is the fingerprint of the content, no client can compute any key other than the one for the captions it actually holds.
What a user can do
Soft rate limit, deferred contract tests
Rate limiting is a fixed window in KV. KV is eventually consistent, so this is not an exact quota, it is a soft ceiling that stops a runaway client. The code comment says as much. If I did it again I would move it to Cloudflare's Rate Limiting binding, or to Durable Objects sharded by IP hash.
Live subtitles are pinned to the wall clock. They are drawn at the moment they arrive rather than against the video timeline, so if you rewind, the subtitles do not follow. The rendering models of Tier 1 and Tier 2 are fundamentally different, and merging them means a live session would have to store its cues anchored to media time. That runs straight into the decision not to store anything (no audio, no transcript), so I deferred it.
The YouTube caption fetch path is a four-step fallback: the InnerTube ANDROID client, via the engine, a direct fetch of the watch page, and intercepting the player response. Today almost everything ends at step one, but that is a fact verified in August 2026 and it can break at any time. There is not one line of contract test for any of the paths, so if one dies quietly the fallback covers it up. At minimum there should have been a smoke test that hits each path on its own.













