Personalized Intro Video Stitching and DM Delivery Tool
Stitching 500 personalized intros onto one shared reel daily
500 a day, never by hand
This was a request to automate video delivery for a B2B outbound campaign. An operator records a 39 second intro that greets one prospect by name, a 3 minute pitch is appended to it, and the result goes out as an Instagram DM.
The throughput is 500 a day. The intro differs per prospect, but the pitch behind it is the same file all 500 times. Each link also has to record how far the recipient watched and where they dropped off. On top of that comes a cost constraint: build 500 videos a day, keep them for 30 days, and stream them, with no human step in the loop and a low fixed monthly bill. These conditions decided the design.

Choosing the no re-encode path
FFmpeg concat demuxer + -c copy
When both inputs share the same encoding parameters, FFmpeg rewrites the container without decoding anything. Re-encoding the 3 minute pitch on every job costs about 5x the CPU, and that file is byte for byte identical across all 500 jobs, so there is no reason to touch it. If even one parameter differs, concat quietly emits a file whose audio drifts out of sync. So I compare with ffprobe first, and when something is off I normalize only the 39 second intro to the pitch's spec.
Cloudflare Queues + Containers
A single stitch takes tens of seconds, which does not fit inside a webhook response. Ingest validates, enqueues, returns 202, and stops there. Everything heavy runs behind the queue, and the queue owns retries entirely. Containers get no Cloudflare bindings. No env.VIDEOS, no env.DB, only environment variables and outbound internet. That constraint is what forced the presigned PUT design below.
R2 hosting plus my own /s/:id.mp4 Range streaming
Egress is $0. The same traffic on S3 plus CloudFront is $81 a month, and $270 at the 1080p assumption in the original brief. Sharing the link more often costs nothing, so there is no reason to ration shares. Mobile Safari refuses to play at all, never mind scrub, unless the 206 Content-Range is exact. I implemented Range parsing myself and put unit tests on it.
D1 + INSERT ... ON CONFLICT DO NOTHING RETURNING
The automation tool's retries and the recorder's re-publish both send the same intro_url twice. Read then write lets two concurrent requests through. One round trip settles it atomically, and only the caller that won the INSERT enqueues. D1's transaction support is thin, so idempotency had to fit inside a single SQL statement.
R2 instead of a YouTube upload, with the interface left in place
YouTube was the destination in the original brief. One upload costs 1,600 quota units against a default daily allowance of 10,000. That is a ceiling of six a day. Not a cost problem, a hard wall. A quota increase could still be approved, so YouTubeTarget is implemented down to resumable upload and OAuth refresh, then fenced off behind a quota guard. Turning it on is one line of DELIVERY_TARGET, not an integration project.

Video bytes that bypass the Worker
In one line: an automation scenario posts an intro URL, the ingest Worker validates it, filters duplicates, enqueues, and answers in milliseconds. The queue consumer picks up one message at a time, chooses a container slot with hash(job_id) % POOL_SIZE, and sends an HMAC signed POST /stitch. The container downloads the intro, compares it against the pitch with ffprobe, concatenates, pulls a thumbnail, and uploads to R2. The consumer writes the result to D1 and fires a signed callback.
One thing stands out here. The finished video bytes never pass through a Worker. The container has no R2 binding, so the consumer mints a presigned S3 PUT URL with a 30 minute TTL and the container streams straight to it. The signing key never leaves the Worker, and a leaked URL expires in 30 minutes and can only write the one key it names.
The reverse direction works the same way. When a prospect opens /w/:id, the delivery Worker reads the job row from D1 and server renders a landing page carrying OG tags, while the video itself is Range streamed by /s/:id.mp4 through the R2 binding. Playback events arrive at /e via sendBeacon, land in D1, and are forwarded to product analytics anonymously.

Staying on the copy path came down to how I read ffprobe
Feed two mp4 files to the concat demuxer and join them with -c copy. It usually works. That is exactly what makes it dangerous. Give it inputs whose parameters disagree and FFmpeg raises no error, it just produces a file whose audio slides further out of sync as it goes. The opposite extreme is just as common: play it safe and re-encode everything every time, which means re-encoding the same 3 minute pitch 500 times.
ffprobe compares codec, resolution, fps, pix_fmt, profile, and audio codec, sample rate and channels across both inputs, builds a list of mismatches, and takes -c copy only when that list is empty. When something is off, only the 39 second intro is normalized to the pitch's parameters, and the 3 minute pitch is never re-encoded under any circumstances. For fps the comparison reads r_frame_rate before avg_frame_rate. A genuine 30fps CFR intro reported avg=4365000/141839 (30.77), and left alone that would have pushed every job onto the re-encode path. I also write encode_path to the job row and to the logs, so a drifting upstream export preset shows up before the invoice does.
The copy path is file I/O and takes about 10 seconds. The re-encode path takes 60 to 90 seconds on 1 vCPU. Compare loosely and broken video reaches the customer. Skip the comparison and compute goes up 7x. Comparing strictly while limiting re-encoding to the 39 second intro caps the worst case at about a fifth of full re-encoding.
Container billing is not about CPU, it is about how long the instance stayed awake
Throughput is short, so raise max_concurrency. Still short, so add slots. And cold starts hurt, so give sleepAfter a generous 15 minutes. All three are natural calls, and all three are wrong in this system.
I measured a 500 job burst and changed three things. First, max_concurrency past POOL_SIZE × MAX_CONCURRENT_STITCHES means nothing. The overflow just waits on a semaphore inside the container, so going from 4 to 32 cut drain time by only 30% while slot wait went from 1.1 seconds to 44.5 seconds. I had moved the queue inside the container. Second, I added lanes rather than slots. Jobs are assigned by hash(job_id) % POOL_SIZE, which is uniform random rather than round robin, so more slots leaves the balls in bins skew intact. POOL_SIZE 8 / 4 lanes beat POOL_SIZE 16 / 2 lanes with half the instances (8.0 minutes against 8.9). Third, I dropped sleepAfter to 2 minutes and added a keepalive that touches storage every 20 seconds so a slot in the middle of an encode is not reclaimed.
Memory and disk are billed for as long as the instance is provisioned, and only vCPU is billed on actual use. Spread 500 jobs a day across 8 slots and each slot sees a job every 7.7 minutes, so a 15 minute timeout never expires and 8 slots stay awake around the clock. That works out to $250 a month against $37. Real compute is 1.84 vCPU hours a day, which is one always on vCPU idling 92% of the time. So this was never a compute budget problem, it was a provisioning problem, and the knob was sleep time rather than concurrency.
What a user can do
Pinned slots and stranded events
A few places still bother me, honestly.
Pinning jobs to slots by hash was a choice to send retries back to a warm instance, and the price was skew. I absorbed it by adding lanes, but that is closer to covering the symptom. If I built this again I would first measure the alternative: drop the pinning, keep a short work queue on the consumer side, and push each job into whichever slot is free. That said, the lane strategy only holds while the copy path dominates. If re-encoding grows, the work becomes CPU bound and 4 lanes on 1 vCPU push each other around, at which point lanes go back to 2 or the pool moves to 2 vCPU instances.
Using D1 as the job state store is worth revisiting too. The queue message and the D1 row live separate lives, so when a message exhausted its retries and disappeared, the row stayed queued and looked like it was processing forever. I added a DLQ consumer that settles the row as failed and fires the callback, but the two states should never have been able to diverge in the first place.
Analytics pile up in a US region. These events come from a page dedicated to a named individual, so even anonymous capture could count as personal data for an EU or UK campaign. Putting a consent banner in front of /e is as far as the implementation goes, and the rest is a legal call, which is what the docs say.
Finally, the tests cover idempotency, Range parsing, config validation, and one ffmpeg smoke test that runs inside the container image. They do not sweep the full range of production encoding combinations. It is not beautifully covered, but there is at least one test standing at the point where money leaks and one at the point where video breaks.











