Skip to content

Interested in AI, automation, blockchain, web and apps

Seoul, KR--:-- GMT
Let’s Talk

Work/AI/KO

Dual Subtitle Player for English Lectures

Dual Subtitle Player for English Lectures

Bilingual captions that keep English lectures learnable

English, not difficulty, blocked the lectures

An English-speaking instructor taught quant trading and the students were Korean, so the lectures needed Korean subtitles.

There were three requirements. First, the source and the translation have to be on the same screen. Second, it has to work the same way outside the course VOD, on YouTube videos and live Q&A sessions that have no caption track at all. Third, watching the same lecture several times should run the translation only once.

The first version was a single VOD player for the course. The second requirement meant it could not stay tied to one player: it had to be a subtitle layer that sits on top of anything.

System boundary and external dependencies
System boundary and external dependencies

One Worker, everything else client side

One Durable Object per live session, one socket, hibernation off

The upstream STT socket has to stay open between client audio frames. A Worker dies with the request, so there is nowhere to hold that state, and a DO fits the slot exactly. Giving up hibernation means paying for as long as the session is alive. In exchange the object is created with newUniqueId() so nothing is persisted, and when stop arrives it waits at most 4 seconds for translations in flight before cutting the connection.

STT behind an adapter interface, swapped with the STT_PROVIDER variable

Real-time STT changes price and quality every quarter. ElevenLabs Scribe v2 Realtime runs by default, and putting Deepgram and a mock behind the same connect/onFragment/onError/onClose contract means swapping any of them touches nothing else in the pipeline. The whole pipeline had to run e2e without keys, so the mock adapter was not optional, it was required.

The extension calls the engine from the background service worker, not the content script

A content script's fetch goes out with the page origin. A request from https youtube.com to http://127.0.0.1:8787 is blocked by Chrome's local network access rules. The service worker goes out on the extension's host permission, so that wall is not there. It meant writing a message round-trip wrapper (engineFetch) that mimics the fetch signature exactly, and it only supports requests whose body is a string.

Deployment and infrastructure
Deployment and infrastructure

Read the captions, otherwise listen

The behavior forks in two. When a caption track exists (Tier 1) the whole track is pulled up front, translated in one pass, and cached, and during playback requestAnimationFrame only lines the cues up with the clock. That is why the added latency is zero. When there is no track (Tier 2), the audio becomes the source. The extension takes the stream from tabCapture and an AudioWorklet in an offscreen document cuts it into 16 kHz PCM16 frames of 100 ms, and the macOS app cuts the system audio ScreenCaptureKit hands it into the same shape. On the receiving end the Durable Object splits on sentence boundaries, translates, and sends the latency back with each one. This piece covers only that live path and the design of the translation cache key. Range serving for the private R2 media and the study features in the web player are each worth their own write-up, so I am leaving them out here.

Core data model
Core data model

Committing a sentence without waiting for silence

The first time you wire up streaming STT you usually do this: throw the partials away, collect only what the recognizer marks final (committed), group those into sentences, then translate. It is what the docs say and it looks safe. But Scribe's commit strategy is VAD. The commit arrives when the speaker pauses. On a lecture or a talk, where nobody pauses, five to ten seconds of text lands at once, and until then the screen holds either nothing or a dim intermediate line. As a subtitle it is unusable.

I picked out of the partial stream the sentences that are already settled in everything but name, and emitted those first. completeSentences looks for the last sentence-ending punctuation that already has the next word attached behind it. Everything up to that point is cut off and promoted immediately as isFinal: true, speechFinal: true, and only the remaining tail goes out as a partial. When the real committed text arrives later, the prefix already emitted (emittedPrefix) is stripped off and only the rest flows through, so nothing is duplicated. On top of that the segmenter cuts again on four conditions (sentence-ending punctuation, 0.6 s of silence, 80 characters, 1.5 s since the start), and the clock that condition 4 needs is fed by the DO extrapolating the last fragment's timestamp on a 250 ms setInterval.

A partial keeps wobbling in the middle of a sentence, but if the next word is already attached behind a sentence, the recognizer has no reason to touch that sentence again. That one observation deletes the entire wait for VAD silence. Cutting on sentences is a requirement too, not a preference: English into Japanese or Korean flips the word order, so translating fragments produces nonsense. On a 20-second streaming measurement with real ElevenLabs and OpenRouter wired in, the time from sentence commit to translation on screen came out at p50 1101 ms, p95 1132 ms.

A cache key that is a fingerprint of the content, recomputed and checked by the server

The first cut of a translation cache gives you keys like cues/<videoId> or <videoId>-en-ko.

I defined the key as a hash of identity plus content. hash = sha256(platform, videoId, srcLang, dstLang, model, contentHash), where contentHash is a hash of the source cues' count + start times and source text concatenated.

Because the key is the fingerprint of the content, no client can compute any key other than the one for the captions it actually holds.

What a user can do

Install the extension and set the cue engine address and language pair once
Install the extension and set the cue engine address and language pair once

Read how the two paths work on the landing page, then move to the extension or the web player
Read how the two paths work on the landing page, then move to the extension or the web player

Adjust the overlay type and color and check it right in the preview
Adjust the overlay type and color and check it right in the preview

Open a YouTube video and two lines, source and translation, attach on their own
Open a YouTube video and two lines, source and translation, attach on their own

Open the same video again and it comes straight from cache with no translation
Open the same video again and it comes straight from cache with no translation

Turn captions off and on for just this tab from the popup
Turn captions off and on for just this tab from the popup

Capture the sound of a tab with no caption track and build subtitles live
Capture the sound of a tab with no caption track and build subtitles live

End the live session and still get the last sentence out
End the live session and still get the last sentence out

Run text track captions on non-YouTube sites through the same path
Run text track captions on non-YouTube sites through the same path

Pick a lecture in the web player and scrub back to any point
Pick a lecture in the web player and scrub back to any point

Change caption languages, their order, type size, and playback speed in the player
Change caption languages, their order, type size, and playback speed in the player

The menu bar app listens to system audio and puts subtitles over any app
The menu bar app listens to system audio and puts subtitles over any app

Set the engine address, API keys, and target language from the menu bar
Set the engine address, API keys, and target language from the menu bar

1 / 1

Soft rate limit, deferred contract tests

Rate limiting is a fixed window in KV. KV is eventually consistent, so this is not an exact quota, it is a soft ceiling that stops a runaway client. The code comment says as much. If I did it again I would move it to Cloudflare's Rate Limiting binding, or to Durable Objects sharded by IP hash.

Live subtitles are pinned to the wall clock. They are drawn at the moment they arrive rather than against the video timeline, so if you rewind, the subtitles do not follow. The rendering models of Tier 1 and Tier 2 are fundamentally different, and merging them means a live session would have to store its cues anchored to media time. That runs straight into the decision not to store anything (no audio, no transcript), so I deferred it.

The YouTube caption fetch path is a four-step fallback: the InnerTube ANDROID client, via the engine, a direct fetch of the watch page, and intercepting the player response. Today almost everything ends at step one, but that is a fact verified in August 2026 and it can break at any time. There is not one line of contract test for any of the paths, so if one dies quietly the fallback covers it up. At minimum there should have been a smoke test that hits each path on its own.

Read next

Link-Based USDT Transfers Without a Wallet

KaiaPay — 2025