Skip to content

Interested in AI, automation, blockchain, web and apps

Seoul, KR--:-- GMT
Let’s Talk

Work/Automation/KO

Personalized Intro Video Stitching and DM Delivery Tool

Personalized Intro Video Stitching and DM Delivery Tool

Stitching 500 personalized intros onto one shared reel daily

500 a day, never by hand

This was a request to automate video delivery for a B2B outbound campaign. An operator records a 39 second intro that greets one prospect by name, a 3 minute pitch is appended to it, and the result goes out as an Instagram DM.

The throughput is 500 a day. The intro differs per prospect, but the pitch behind it is the same file all 500 times. Each link also has to record how far the recipient watched and where they dropped off. On top of that comes a cost constraint: build 500 videos a day, keep them for 30 days, and stream them, with no human step in the loop and a low fixed monthly bill. These conditions decided the design.

System boundary and external dependencies
System boundary and external dependencies

Choosing the no re-encode path

FFmpeg concat demuxer + -c copy

When both inputs share the same encoding parameters, FFmpeg rewrites the container without decoding anything. Re-encoding the 3 minute pitch on every job costs about 5x the CPU, and that file is byte for byte identical across all 500 jobs, so there is no reason to touch it. If even one parameter differs, concat quietly emits a file whose audio drifts out of sync. So I compare with ffprobe first, and when something is off I normalize only the 39 second intro to the pitch's spec.

Cloudflare Queues + Containers

A single stitch takes tens of seconds, which does not fit inside a webhook response. Ingest validates, enqueues, returns 202, and stops there. Everything heavy runs behind the queue, and the queue owns retries entirely. Containers get no Cloudflare bindings. No env.VIDEOS, no env.DB, only environment variables and outbound internet. That constraint is what forced the presigned PUT design below.

R2 hosting plus my own /s/:id.mp4 Range streaming

Egress is $0. The same traffic on S3 plus CloudFront is $81 a month, and $270 at the 1080p assumption in the original brief. Sharing the link more often costs nothing, so there is no reason to ration shares. Mobile Safari refuses to play at all, never mind scrub, unless the 206 Content-Range is exact. I implemented Range parsing myself and put unit tests on it.

D1 + INSERT ... ON CONFLICT DO NOTHING RETURNING

The automation tool's retries and the recorder's re-publish both send the same intro_url twice. Read then write lets two concurrent requests through. One round trip settles it atomically, and only the caller that won the INSERT enqueues. D1's transaction support is thin, so idempotency had to fit inside a single SQL statement.

R2 instead of a YouTube upload, with the interface left in place

YouTube was the destination in the original brief. One upload costs 1,600 quota units against a default daily allowance of 10,000. That is a ceiling of six a day. Not a cost problem, a hard wall. A quota increase could still be approved, so YouTubeTarget is implemented down to resumable upload and OAuth refresh, then fenced off behind a quota guard. Turning it on is one line of DELIVERY_TARGET, not an integration project.

Deployment and infrastructure
Deployment and infrastructure

Video bytes that bypass the Worker

In one line: an automation scenario posts an intro URL, the ingest Worker validates it, filters duplicates, enqueues, and answers in milliseconds. The queue consumer picks up one message at a time, chooses a container slot with hash(job_id) % POOL_SIZE, and sends an HMAC signed POST /stitch. The container downloads the intro, compares it against the pitch with ffprobe, concatenates, pulls a thumbnail, and uploads to R2. The consumer writes the result to D1 and fires a signed callback.

One thing stands out here. The finished video bytes never pass through a Worker. The container has no R2 binding, so the consumer mints a presigned S3 PUT URL with a 30 minute TTL and the container streams straight to it. The signing key never leaves the Worker, and a leaked URL expires in 30 minutes and can only write the one key it names.

The reverse direction works the same way. When a prospect opens /w/:id, the delivery Worker reads the job row from D1 and server renders a landing page carrying OG tags, while the video itself is Range streamed by /s/:id.mp4 through the R2 binding. Playback events arrive at /e via sendBeacon, land in D1, and are forwarded to product analytics anonymously.

Core data model
Core data model

Staying on the copy path came down to how I read ffprobe

Feed two mp4 files to the concat demuxer and join them with -c copy. It usually works. That is exactly what makes it dangerous. Give it inputs whose parameters disagree and FFmpeg raises no error, it just produces a file whose audio slides further out of sync as it goes. The opposite extreme is just as common: play it safe and re-encode everything every time, which means re-encoding the same 3 minute pitch 500 times.

ffprobe compares codec, resolution, fps, pix_fmt, profile, and audio codec, sample rate and channels across both inputs, builds a list of mismatches, and takes -c copy only when that list is empty. When something is off, only the 39 second intro is normalized to the pitch's parameters, and the 3 minute pitch is never re-encoded under any circumstances. For fps the comparison reads r_frame_rate before avg_frame_rate. A genuine 30fps CFR intro reported avg=4365000/141839 (30.77), and left alone that would have pushed every job onto the re-encode path. I also write encode_path to the job row and to the logs, so a drifting upstream export preset shows up before the invoice does.

The copy path is file I/O and takes about 10 seconds. The re-encode path takes 60 to 90 seconds on 1 vCPU. Compare loosely and broken video reaches the customer. Skip the comparison and compute goes up 7x. Comparing strictly while limiting re-encoding to the 39 second intro caps the worst case at about a fifth of full re-encoding.

Container billing is not about CPU, it is about how long the instance stayed awake

Throughput is short, so raise max_concurrency. Still short, so add slots. And cold starts hurt, so give sleepAfter a generous 15 minutes. All three are natural calls, and all three are wrong in this system.

I measured a 500 job burst and changed three things. First, max_concurrency past POOL_SIZE × MAX_CONCURRENT_STITCHES means nothing. The overflow just waits on a semaphore inside the container, so going from 4 to 32 cut drain time by only 30% while slot wait went from 1.1 seconds to 44.5 seconds. I had moved the queue inside the container. Second, I added lanes rather than slots. Jobs are assigned by hash(job_id) % POOL_SIZE, which is uniform random rather than round robin, so more slots leaves the balls in bins skew intact. POOL_SIZE 8 / 4 lanes beat POOL_SIZE 16 / 2 lanes with half the instances (8.0 minutes against 8.9). Third, I dropped sleepAfter to 2 minutes and added a keepalive that touches storage every 20 seconds so a slot in the middle of an encode is not reclaimed.

Memory and disk are billed for as long as the instance is provisioned, and only vCPU is billed on actual use. Spread 500 jobs a day across 8 slots and each slot sees a job every 7.7 minutes, so a 15 minute timeout never expires and 8 slots stay awake around the clock. That works out to $250 a month against $37. Real compute is 1.84 vCPU hours a day, which is one always on vCPU idling 92% of the time. So this was never a compute budget problem, it was a provisioning problem, and the knob was sleep time rather than concurrency.

What a user can do

The automation scenario sends an intro URL and creates a job
The automation scenario sends an intro URL and creates a job

Sending the same intro URL again does not render it twice
Sending the same intro URL again does not render it twice

Pick a job off the queue, join intro and pitch, upload to R2
Pick a job off the queue, join intro and pitch, upload to R2

A failed job retries, or ends as failed at once when retrying is pointless
A failed job retries, or ends as failed at once when retrying is pointless

The finished result returns to the automation tool with a signature
The finished result returns to the automation tool with a signature

Open the link you were sent and watch the video made for you
Open the link you were sent and watch the video made for you

Tapping a link whose render is not finished still lets you wait
Tapping a link whose render is not finished still lets you wait

Rewind to watch a favorite stretch again
Rewind to watch a favorite stretch again

Leave the page mid watch and where you stopped is kept
Leave the page mid watch and where you stopped is kept

Open the video file directly, or save it as a file
Open the video file directly, or save it as a file

Paste the link into a DM and it expands into a card with the prospect's face
Paste the link into a DM and it expands into a card with the prospect's face

Watch what happens while a batch runs, in a table
Watch what happens while a batch runs, in a table

1 / 1

Pinned slots and stranded events

A few places still bother me, honestly.

Pinning jobs to slots by hash was a choice to send retries back to a warm instance, and the price was skew. I absorbed it by adding lanes, but that is closer to covering the symptom. If I built this again I would first measure the alternative: drop the pinning, keep a short work queue on the consumer side, and push each job into whichever slot is free. That said, the lane strategy only holds while the copy path dominates. If re-encoding grows, the work becomes CPU bound and 4 lanes on 1 vCPU push each other around, at which point lanes go back to 2 or the pool moves to 2 vCPU instances.

Using D1 as the job state store is worth revisiting too. The queue message and the D1 row live separate lives, so when a message exhausted its retries and disappeared, the row stayed queued and looked like it was processing forever. I added a DLQ consumer that settles the row as failed and fires the callback, but the two states should never have been able to diverge in the first place.

Analytics pile up in a US region. These events come from a page dedicated to a named individual, so even anonymous capture could count as personal data for an EU or UK campaign. Putting a consent banner in front of /e is as far as the implementation goes, and the rest is a legal call, which is what the docs say.

Finally, the tests cover idempotency, Range parsing, config validation, and one ffmpeg smoke test that runs inside the container image. They do not sweep the full range of production encoding combinations. It is not beautifully covered, but there is at least one test standing at the point where money leaks and one at the point where video breaks.

Read next

Search Tool for Scanned Textbooks and Uncaptioned Lectures

UniLens — 2026