Skip to content

Interested in AI, automation, blockchain, web and apps

Seoul, KR--:-- GMT
Let’s Talk

Work/AI/KO

Search Tool for Scanned Textbooks and Uncaptioned Lectures

Search Tool for Scanned Textbooks and Uncaptioned Lectures

Searching scanned textbooks and subtitle free lectures, with citations

Unsearchable scans, subtitle free lectures

The course material to handle split two ways. One half was scanned PDF textbooks with no searchable text, the other was English lecture videos with no subtitles.

The ask came down to four things. Only accounts on a specified email domain get in. Material a professor distributes and material a student uploads never mix. Every LLM answer carries a page or timestamp citation. And the whole thing runs without a server that stays on. That last condition effectively picked the stack.

System boundary and external dependencies
System boundary and external dependencies

Indexing and chat, no servers

Everything inside one Cloudflare account: Workers, D1, R2, Queues, Workflows, Vectorize, Workers AI, Containers

There was no budget for an always on server, and when auth, storage, queues, and vector search all attach as bindings, there is almost no infrastructure code left to write. Workers call each other over service bindings, so internal traffic never needed a public endpoint either. The OCR container (study-ocr) needs a Workers Paid plan before you can push an image, and Worker memory is capped at 128MB.

Login belongs to Cloudflare Access, not the app. The app only verifies the JWT.

One Allow policy on the permitted email domain is the entire access control story. apps/api/src/auth.ts verifies Cf-Access-Jwt-Assertion against JWKS (cached 10 minutes) and looks the user up by the email claim, creating a student row on first sight. I never wrote a password field or a session table. Access cannot be imitated locally, so AUTH_DEV_BYPASS=1 plus a study_dev_user cookie swaps in a seeded account.

Indexing moves out of the request: one queue message, then a Workflow with many steps

OCR is budgeted at 90 minutes and subtitle translation at 60. Neither fits inside a request response lifetime. Workflow steps let me set timeout and retry per step, and when a cues-en.json already exists the run skips STT and redoes translation only. There can be only one live queue consumer, so PR previews do not deploy study-ingest.

Chat is an Agents SDK Durable Object, one instance per (user, scope) pair

Keeping conversation history in DO state removes the session table, and the WebSocket can stream status → context → delta → answer straight through. A source view, a course, a notebook, and everything are each a different room. The DO name is the isolation boundary, so the moment the client gets to pick it, the client can walk into someone else's room.

EmbeddingGemma on Workers AI for embeddings, OpenRouter for answer generation

Embedding runs thousands of times during indexing, so it has to stay inside the account to be cheap and fast. The answer model was the opposite. I wanted to swap it often, so the chat input accepts a model string you can change on the spot. With no key locally, a mock response summarizing the retrieved chunks streams over the same protocol, and when Vectorize is unavailable retrieval drops down to D1 keyword search.

Deployment and infrastructure
Deployment and infrastructure

Shared permission rules, API and chat

The front door is a single worker, study-web. Workers Assets serves the SPA, and only /api/* is intercepted first and proxied to study-api over a service binding. From the browser's point of view the static files, the API, and the WebSocket all live on one origin.

study-api is one Hono app: CRUD over D1, R2 multipart upload, queue publishing, FSRS grading, and the agent proxy. Files are never handed out as R2 links. The API takes the Range request and answers it itself, and the PDF viewer's page jumps and the video seeking ride on top of that.

When an upload finishes, one study-ingest queue message goes out, and the consumer starts BookIngestWorkflow for a book or VideoIngestWorkflow for a video. Books stream the PDF into the OCR container and get NDJSON back. Videos get English subtitles from ElevenLabs, then a Korean translation, and write cues.json plus two VTT files to R2. Both pipelines end the same way: chunks embedded into Vectorize, rows into D1. Progress is overwritten onto sources.ingest_stage and ingest_progress, which the screen polls every 2.5 seconds.

When a question arrives, the study-agent Durable Object reads the user and the scope back out of its own name, builds the list of source ids that user may read, passes it into the Vectorize filter, and keeps only the top 8 chunks as context. That "may this person read it" judgement lives in exactly one place, packages/access, and the API and the agent call the same function.

Core data model
Core data model

Why the server rewrites the chat room name

The browser picks the room name, something like /agents/notebook-agent/notebook-123, connects, and permissions get checked after the answer is generated. That gives one room per notebook, so everyone who can see that notebook shares one DO's history, and editing the address drops you into someone else's room. The permission check also exists in two copies, one in the API and one in the agent, and the two drift apart over time.

The browser only states a scope (all / course:<id> / notebook:<id> / source:<id>). The API proxy checks with canUseScope that the caller may use that scope, then rewrites just the name segment of the URL path to <userId>~<scope> before forwarding to the agent. The DO splits this.name on ~ to recover the user id and the scope, and before retrieval it builds an allow list with scopeSourceIds() and puts it into a Vectorize source_id $in filter. The read rule SQL exists once, as READABLE_WHERE in packages/access/index.ts, and the API and the agent share it.

Session state is physically split per (user, scope), so histories cannot mix, and the user id appears nowhere in the URL, so forging the address only ever leads back to your own room. Because one function decides access, the mismatch where the API refuses but chat still answers cannot arise by construction.

Translating 2,000 subtitle lines without losing one line

Either throw the whole transcript at the model in one go, or call it one line at a time. All at once, the model merges or splits short sentences on its own, the line count stops matching, and the timestamp alignment shifts wholesale. One line at a time means 2,000 calls with no surrounding context, so terminology changes from sentence to sentence.

I send batches of 20 lines and carry the previous 5 EN/KO pairs along as context. The response is forced to {"ko":[...]} JSON, and the parser strips <think> blocks, code fences, and leading numbers, then truncates when the list is too long and pads with the original English when it is too short so the count matches (parseKoList). If a batch still fails it splits in half and retries recursively, and the model falls back Gemini 2.5 Pro → Flash → DeepSeek. If it fails down to a single line, that line stays English and the run moves on.

Whichever path a failure takes, the invariant holds: the number of lines coming out equals the number going in. Cue count is preserved, so start/end in cues.json survive intact, and the t_start of the roughly 45 second chunks built on top of them is exact. Clicking a citation chip jumps the video to the moment that line is actually spoken because of that invariant. The unit of failure also narrows from a whole video to one line.

What a user can do

Sign in with a school email for the first time and land on home
Sign in with a school email for the first time and land on home

Enter an invite code and join a course
Enter an invite code and join a course

Create a course and set up its weeks
Create a course and set up its weeks

Upload a PDF textbook and turn it into searchable material
Upload a PDF textbook and turn it into searchable material

Paste a lecture video link and get dual subtitle material
Paste a lecture video link and get dual subtitle material

Ask a source a question and click a citation to jump to the original
Ask a source a question and click a citation to jump to the original

Make review cards, a whole source or only the pages I annotated
Make review cards, a whole source or only the pages I annotated

Grade the cards that are due today
Grade the cards that are due today

Bundle sources into a notebook and ask across them
Bundle sources into a notebook and ask across them

Request to share my material with a course and get it approved
Request to share my material with a course and get it approved

Attach material to a week and schedule when it opens
Attach material to a week and schedule when it opens

Read the course analytics and download the CSV
Read the course analytics and download the CSV

An admin adjusts roles and quota and reads the audit log
An admin adjusts roles and quota and reads the audit log

1 / 1

Figure captioning still returns zero

The biggest hole is figures. The OCR container extracts figures alongside page text and ships them as base64 inside the NDJSON, and the worker throws those lines away whole, without JSON.parse or atob. That was a CPU budget decision. The result is that figures.json is always an empty array and the caption-figures step finishes with zero items. VLM captioning and figure citations are all there in the code, and none of it actually runs. If I built it again, the container would PUT figures straight to R2 and pass only the keys to the worker. Trying to push bytes through the worker was the wrong choice from the start in this architecture.

Video audio also goes to ElevenLabs from inside worker memory, so a two hour lecture is roughly the ceiling. The educator analytics are aggregated on request rather than in real time, and with no KV binding there is no cache either, so every load recomputes. That has to move to precomputed aggregates before hundreds of people are on it at once. Some leftovers are still sitting there unused and undeleted: the jobs table, sources.area, the old review_state. Testing leans on Playwright e2e, and packages/access/*.test.ts, which pnpm test:unit points at, does not exist yet. Logic collected in one place, like the permission rules, is exactly what should have had unit tests first.

Read next

Timed Reveal Travel Log Map Promotion Site

RIIZE — 2026