Skip to content

Interested in AI, automation, blockchain, web and apps

Seoul, KR--:-- GMT
Let’s Talk

Work/AI/KO

Vocabulary Flashcard App Built from Your Journal

Vocabulary Flashcard App Built from Your Journal

Turning my own journal sentences into vocabulary cards

Words from my own sentences

This started as a personal tool for one second-language learner. Instead of memorizing a word list prepared in advance, only the words pulled from sentences the learner wrote themselves become cards.

The pipeline takes a Korean sentence, translates it, pulls out the words, finds their dictionary forms, attaches pronunciation, and packs them into cards. Rather than have a person carry the work from tool to tool at each step, I wired the steps together so one sentence goes in and cards come out.

There were two constraints. Input arrives through a Telegram bot, and the final output has to land as the Anki deck (.apkg) file the learner already used. And it had to run without a resident server, spinning up only when a request comes in.

So this post covers two things only: how I ran this pipeline without a resident server, and how the Anki deck gets baked inside a Worker. FSRS review scheduling, offline sync, and the Expo web fallback each deserve their own post, so I'm folding them away here.

System boundary and external dependencies
System boundary and external dependencies

Fourteen bots, one app, no server

Cloudflare Workers + D1 + R2

A Telegram webhook that doesn't get a 200 back quickly enough triggers a flood of retries. A Worker that wakes only when a request arrives and runs at edges worldwide fit that traffic shape exactly. Audio goes to R2, and everything else fits in a single D1 (SQLite). It had to run without a resident server, effectively on the free tier, and in exchange I had to solve everything on workerd, where there is no Node API and no filesystem.

14 Workers in one pnpm workspace, a single D1

Telegram gives you one webhook per bot token. So I deploy a separate Worker per bot (bots/ru, bots/jp, 13 of them), while the learning logic lives once in packages/shared and the app-only Worker bots/app runs that same code. Secrets differ per Worker, but the data has to sit in one place. All 14 wrangler.toml files point at the same database_id, and migrations run as a single line in migrations/.

Translation stays synchronous; TTS, word extraction, and question generation move behind waitUntil

A three-sentence diary means three TTS calls. Waiting for all of them leaves the user staring at an empty screen for more than 20 seconds. I wait only for the translation so the card lands first, and run the rest after the response is sent. The body has to survive a failed background step, so Promise.allSettled lets each piece fail on its own, and once everything finishes I stamp processed_at so the app stops polling.

Deployment and infrastructure
Deployment and infrastructure

Translate now, the rest after

The core pipeline is one function, runLearningPipeline. It splits the Korean source into sentences, translates them into the target language, writes them into entries and sentences, and returns immediately. After that, in the background, it uploads per-sentence TTS to tts/{lang}/s/{id}.mp3 in R2, pulls out words and grammar chunks into words and grammar_chunks, and builds one "question this sentence answers" per sentence into sentence_questions. Meanwhile the app polls GET /api/entries/:id and shows an audio button and a question line attaching to the card one at a time.

Core data model
Core data model

Baking a whole Anki deck inside a Worker

An .apkg is really just one SQLite file and some mp3s zipped together. So the usual move is to stand up a Node server and hand the job to a genanki-style library: write a temporary collection.anki2 on the filesystem, fill it with better-sqlite3, zip the directory. The problem is that this project has no such server. The moment I stand up a resident instance for one Worker, the premise of going serverless collapses. And handing over a bare TSV with "add the audio yourself" would mean a system built to remove manual work handing the manual work right back.

I moved deck generation into the Worker. I import sql.js as a WASM module and instantiate that module directly inside initSqlJs's instantiateWasm hook. Two things snagged here. One: the emscripten glue sees WorkerGlobalScope, decides it is a web worker, and reads self.location.href, but workerd has no location. So I planted a harmless location with Object.defineProperty. Two: Workers forbid compiling a byte array at runtime, which the hook sidesteps by handing over a pre-compiled module.

On top of that I set up the schema 11 tables with plain db.run calls and pinned note GUIDs to saiwon:{lang}:{w|s|g}:{id}. Note and card ids come out of the GUID through an FNV hash so they reproduce. Audio is read from R2 and written into the zip with level: 0, uncompressed. Past 500 notes the .apkg splits into several pieces, and if there are two or more they get wrapped into a .zip bundle, then the result is uploaded to exports/ in R2 and served through a tokenized URL.

Pinning the GUID mattered most. Anki matches notes by GUID, so exporting again two months later with the same filter updates the existing cards instead of piling up duplicates. Whatever review history the user built on those cards survives untouched.

The uncompressed zip is deliberate too. mp3s are already compressed, so squeezing them again only burns CPU, and an uncompressed collection.anki2 is readable by older Anki versions as is. The 500-note chunking is a safety line to stay under the Worker memory limit. With a queue binding the work runs in the consumer, without one it trails the request through waitUntil, so the same code runs without a paid plan.

Reshaping two languages of data into an eight-language schema

The schema was bilingual by construction, with ru_translation and jp_translation sitting side by side in entries, and going to eight languages meant moving to one language per row.

Inside a single migration file, 0021, I deferred foreign key checks with PRAGMA defer_foreign_keys, parked the child rows in _mig_ temp tables, rebuilt the parent tables, and put the rows back.

The eight-language schema was adopted without dropping a single line of review history. How the deferred FKs and the child parking combine is a post of its own, so I'm only recording the outcome here.

What a user can do

Open the app for the first time and pick one language to learn
Open the app for the first time and pick one language to learn

Send a Korean sentence and get the translation, pronunciation, and words at once
Send a Korean sentence and get the translation, pronunciation, and words at once

Long press a sentence to star, tag, regenerate its question, or delete it
Long press a sentence to star, tag, regenerate its question, or delete it

Search the library in Korean and in the target language
Search the library in Korean and in the target language

Browse the vocabulary list and fix a level in the word detail
Browse the vocabulary list and fix a level in the word detail

Rate unclassified words and grammar one card at a time from 1 to 4
Rate unclassified words and grammar one card at a time from 1 to 4

Flip through the cards due today and grade them
Flip through the cards due today and grade them

Set a target date and a daily new-card count
Set a target date and a daily new-card count

Apply filters and export to an Anki deck
Apply filters and export to an Anki deck

Sign in with an emailed code and merge anonymous history into the account
Sign in with an emailed code and merge anonymous history into the account

Link the app account to the Telegram bot with a 6-digit code
Link the app account to the Telegram bot with a 6-digit code

Switch the study language and flip the whole app over to it
Switch the study language and flip the whole app over to it

Use the compare, dictionary form, etymology, kanji, and read-aloud tools
Use the compare, dictionary form, etymology, kanji, and read-aloud tools

1 / 1

Zero tests, a guessed chunk size

There is not one line of tests. pnpm typecheck and eslint are the whole story. For irreversible work like migration 0021 in particular, running it a few times against a local D1 copy and eyeballing the result was the entire verification. If I did it again I would at least keep one snapshot as a fixture and diff row counts and referential integrity before and after the migration automatically.

EXPORT_CHUNK_SIZE = 500 is honestly a guess, not a measurement. It's the line that felt like it wouldn't blow past Worker memory, and the real safe line shifts with how much audio hangs off each note. sql.js holds the whole collection in memory, so this is the first place in the system that will break. The Queues binding is ready in the code but commented out in wrangler.toml, so today even a large export trails the request through waitUntil.

Finally, of the eight languages, only Russian, Japanese, and Chinese have prompts and voices that are properly tuned. The other five ride the same pipeline, but their pronunciation notation rules and voices are still defaults. I didn't get everything polished, but at least I know where it's thin.

Read next

PDF to Quiz and Study Schedule App

LumiQuiz — 2024