One daily run takes an idea from a queue, writes a 5–7 scene script, voices it with word-level timings, finds matching vertical footage, renders a captioned 9:16 video with music and motion, uploads it to YouTube, then comments, files it into a playlist and notifies me. A secure, installable dashboard shows every step live.
- pipeline stages, each with a fallback
- 6
- AI & media APIs orchestrated
- 4
- node:test suites, incl. a real FFmpeg render
- 13
- commits from first idea to production
- 24
The problem
Posting short-form video every day is the same routine every time: pick a topic, write a hook, record a voice-over, hunt for footage, cut it, caption it, upload it, fill in the metadata. That's the kind of work an agent should do.
The constraints made it interesting. It had to run on free tiers (a Render instance that sleeps when idle, the Gemini free tier, a free MongoDB cluster), never post twice or burn paid credits on a failure loop, and still leave room for a human to review a video before it goes public.
Architecture
A single Next.js 16 app is both the dashboard and the worker. External pingers wake it hourly; an idempotent cron endpoint decides whether today's run is due.
Trigger
GitHub Actions + cron-job.org
hourly POST /api/cron with a bearer secret
Due check
post time passed · nothing posted today · < 3 attempts
Pipeline
Script
Gemini Flash → Flash-Lite, or OpenAI · strict JSON
Voice
ElevenLabs Flash v2.5 → Gemini TTS
Footage
Pexels clip per scene query
Render
FFmpeg 720×1280 · word captions · music
Upload
YouTube Data API · thumbnail · playlist
State & signals
MongoDB
posts · runs · ideas · settings · analytics
Dashboard
SSE live steps · review queue · PWA
Alerts
Resend email · Web Push
The pipeline
- Script: the LLM returns a title, two alternative titles, thumbnail texts, a 3-second hook and 5–7 scenes (narration + a filmable stock-footage query) as strict JSON. It avoids recent topics and, once 5+ videos are two days old, learns from which hooks got the most views.
- Voice: ElevenLabs returns audio with word timings for synced captions; when quota runs low or a call fails, the free Gemini voice takes over, and a low-quota alert fires at 20% remaining.
- Render: FFmpeg builds a 9:16 video with word-by-word captions and highlighted key words, a title card, a slow pan on every clip, a progress bar and royalty-free background music.
- Publish: upload, custom thumbnail where YouTube allows it, add to a channel playlist, and post a first comment that asks viewers a question.
Review mode
In Review first mode the video uploads privately and waits. The MP4 is kept in MongoDB (up to 150 MB, swept after 7 days) so I can watch it on the dashboard, pick one of three titles or write my own, pick a thumbnail, then Approve & publish, or let YouTube auto-publish after N hours.
Security for a single-owner dashboard
- Password sign-in plus passkeys (Face ID, fingerprint, Windows Hello) via WebAuthn.
- HttpOnly signed session cookie; Sign out everywhere ends every other session at once, and changing the password signs out every device.
- Brute-force lockout (5 failed attempts per network per 15 minutes) with an email alert, plus an email on sign-in from a new network.
- CSRF origin checks on writes, CSP, X-Frame-Options, HSTS, strict referrer, no-store API responses, and a 90-day activity log.
Testing & CI
node:test suites cover the pipeline, scheduling, captions, error descriptions, auth and security, and a real FFmpeg render with no network. GitHub Actions runs lint, typecheck, tests and a production build on every push and pull request.
What broke, and how I fixed it
- GitHub's scheduled workflows are best-effort: on one day only 5 of ~20 hourly runs fired, none in the posting window. I added cron-job.org as a second pinger (safe, because the endpoint never posts twice) and a Post it now button when a post is an hour late.
- Next.js bundles lib/ separately into each route handler while the custom server loads it with require, which duplicated in-memory state (split logs, double runs). Moving pipeline state onto globalThis made it one shared instance.
- YouTube keeps uploads from unverified API projects private, and OAuth refresh tokens expire after 7 days while the consent screen is in Testing, so both are documented in the setup guide and surfaced in the run log.
Key decisions
Idempotent cron, external pingers
Free Render instances sleep, so the in-process cron can't be trusted. Any number of pings is safe: the app posts only when today's slot has passed and nothing has been posted yet.
A fallback at every paid step
Gemini Flash falls back to Flash-Lite on 429/503; ElevenLabs falls back to Gemini TTS. Retries use backoff, and a 3-attempt daily cap stops a broken key from burning credits.
A contract the prompt can't break
The output schema is appended to every prompt, default or custom, so the pipeline always receives the fields it needs no matter how the niche prompt is worded.
Human in the loop, optionally
Auto mode publishes directly; review mode uploads privately and waits for approval, which catches mistakes and adds the human input platforms increasingly expect.
What's next
- Edit the script on the dashboard before it renders
- Record my own voice-over in the browser, timed with forced alignment
- A fact-check pass that flags unsupported claims
- A trend scout that suggests ideas from what's working in the niche
Want something like this built, or someone who builds it?
Tell me about your project, or grab my résumé if you're hiring.