Subtitles are essential for accessibility, global reach, and viewer engagement. But generating and burning subtitles manually is tedious and time‑consuming. With AI‑powered speech‑to‑text, you can automate the entire process — from audio transcription to subtitle burn‑in — making your content accessible to a global audience.
This guide shows you how to build a fully automated subtitle generation pipeline using a declarative YAML template and a transpiler that generates the complete SQL migration. The pipeline:
- Extracts audio from uploaded videos
- Transcribes speech using OpenAI Whisper (state‑of‑the‑art multilingual model)
- Generates SRT subtitle files with timestamps
- Burns subtitles into the video using FFmpeg
- Delivers the final video with embedded subtitles
Key takeaways
- One YAML file — define your entire pipeline in a single, version‑controlled file.
- Automatic SQL generation — the transpiler produces the exact PostgreSQL migration.
- Whisper AI — state‑of‑the‑art speech‑to‑text supporting 99 languages.
- Zero manual intervention — upload a video, get a subtitled version automatically.
- Scalable and reliable — pgmq provides durable, transaction‑safe job queuing.
- Configurable — adjust language, model size, and subtitle styling.
- Visual graph — understand your pipeline at a glance with an SVG diagram.
- Per‑run grouping – all outputs for a single upload are stored under a unique
runIdfolder.
The Gap: Accessibility at Scale
Subtitles are no longer optional. They're essential for:
- Accessibility — deaf and hard‑of‑hearing viewers
- Global reach — viewers who speak different languages
- Engagement — videos with subtitles have higher watch time
- Compliance — accessibility regulations (ADA, WCAG, etc.)
But generating subtitles manually is expensive and doesn't scale. What if the pipeline could be fully automated — triggered by the upload itself, transcribing and burning subtitles in the background? And what if you could define that pipeline in a declarative YAML file that you can version, share, and reuse?
This guide shows you exactly how to build that pipeline.
Architecture Overview
Extract Audio → Whisper AI → Generate SRT → Burn Subtitles
Subtitled Video → Public Folder → pg_notify → User Notified
The pipeline consists of:
- Supabase Storage — two buckets:
private-uploads(per‑user) andpublic-processed(per‑user). - RLS policies — restrict access to each user's own folders.
- PostgreSQL trigger — fires on
INSERTintostorage.objects. - pgmq — message queue for job processing (uses the existing
renderqueue). - ffmpeglab-runner — extracts audio, runs Whisper, burns subtitles.
- pg_notify — real‑time status updates.
- YAML transpiler — reads the pipeline definition and generates the SQL migration + SVG graph.
Important: This pipeline uses the existing render and logpiece tables from the FFmpegLab server. It does not create new tables — it only adds the pipeline components.
What the Pipeline Delivers
| Output | Format | Location |
|---|---|---|
| Subtitled Video | MP4 (H.264) with embedded subtitles | public-processed/{userId}/{pipelineId}/{runId}/subtitled/ |
| SRT Subtitle File | .srt (standalone subtitles) | public-processed/{userId}/{pipelineId}/{runId}/subtitles/ |
| Real‑time notifications | pg_notify channels | N/A |
| Job tracking | render table | Existing FFmpegLab table |
| Logs | logpiece table | Existing FFmpegLab table |
Prerequisites
- A Supabase project (cloud or self‑hosted).
- ffmpeglab-server and ffmpeglab-runner deployed.
- Whisper.cpp installed on the runner (see Setting Up Whisper).
- The
renderandlogpiecetables must already exist (created by the FFmpegLab server migrations). - Access to your Supabase database (psql or the Supabase SQL Editor).
- Deno installed to run the transpiler.
The YAML‑Driven Approach
While you can write the SQL directly, the recommended way is to use the YAML transpiler. This gives you:
- Declarative pipeline definition — define steps, triggers, and buckets in clean YAML.
- Automatic SQL generation — the transpiler produces the exact PostgreSQL migration.
- Visual pipeline graph — generate an SVG diagram of your pipeline with
--svg. - Reusable templates — share and version your pipeline definitions.
The transpiler is a single TypeScript file that reads your YAML and generates the SQL migration. It runs with Deno and has zero external dependencies (except yaml for parsing).
The YAML Template
Create a file called subtitle-pipeline.yaml with the following content. It defines the buckets, RLS policies, and each processing step. The runId section configures how the per‑run ID is generated — in this case, deterministically from the input file name.
name: "Automated Subtitle Generation Pipeline" pipelineId: "subtitle-pipeline" runId: mode: "deterministic" template: "{baseFilename}" description: "Extract audio, transcribe with Whisper, burn subtitles into video" version: "1.0.0" editor: compressionLevel: 23 preset: "medium" aspectRatio: "16:9" framerate: 30 opacity: 1.0 output: "mp4" storage: output_bucket: "public-processed" buckets: - name: "private-uploads" public: false allowed_mime_types: - "video/mp4" - "video/quicktime" - "video/x-msvideo" - "video/webm" - "video/mpeg" - name: "public-processed" public: true allowed_mime_types: - "video/mp4" - "text/plain" rls_policies: - name: "Users can upload to their own folder" operation: "INSERT" role: "authenticated" condition: | bucket_id = 'private-uploads' AND (storage.foldername(name))[1] = auth.uid()::text - name: "Users can download from their own folder" operation: "SELECT" role: "authenticated" condition: | bucket_id = 'private-uploads' AND (storage.foldername(name))[1] = auth.uid()::text - name: "Public read access to processed media" operation: "SELECT" role: "anon" condition: | bucket_id = 'public-processed' - name: "Service role can manage processed media" operation: "ALL" role: "service_role" condition: | bucket_id = 'public-processed' - name: "Users can read their own processed media" operation: "SELECT" role: "authenticated" condition: | bucket_id = 'public-processed' AND (storage.foldername(name))[1] = auth.uid()::text steps: # Step 1: Extract audio from video - id: "extract_audio" trigger: name: "handle_extract_audio" event: "INSERT" table: "storage.objects" condition: | NEW.bucket_id = 'private-uploads' AND NEW.metadata->>'mimetype' LIKE 'video/%' command: -i $MEDIA_1 -ac 1 -ar 16000 -vn -y $OUTPUT_PATH inputs: ["INPUT_FILE"] outputs: ["OUTPUT_FILE"] output_path: "{{userId}}/{{pipelineId}}/{{runId}}/temp/{{baseFilename}}.wav" editor: output: "wav" preset: "fast" selectedCode: "custom" width: 0 height: 0 compressionLevel: 0 next_bucket: "private-uploads" keep: false # Step 2: Transcribe with Whisper - id: "transcribe" trigger: name: "handle_transcribe" event: "INSERT" table: "storage.objects" condition: | NEW.bucket_id = 'private-uploads' AND NEW.name LIKE '%.wav' AND NEW.name NOT LIKE '%.emptyFolderPlaceholder' command: whisper --model $WHISPER_MODEL --language $WHISPER_LANG --output-srt $OUTPUT_PATH $MEDIA_1 inputs: ["INPUT_FILE"] outputs: ["OUTPUT_FILE"] output_path: "{{userId}}/{{pipelineId}}/{{runId}}/subtitles/{{baseFilename}}.srt" editor: output: "srt" preset: "medium" selectedCode: "custom" next_bucket: "private-uploads" keep: false # Step 3: Burn subtitles into video - id: "burn_subtitles" trigger: name: "handle_burn_subtitles" event: "INSERT" table: "storage.objects" condition: | NEW.bucket_id = 'private-uploads' AND NEW.name LIKE '%.srt' AND NEW.name NOT LIKE '%.emptyFolderPlaceholder' command: -i $MEDIA_1 -vf "subtitles=$MEDIA_2:force_style='FontName=Arial,FontSize=24,PrimaryColour=&HFFFFFF,OutlineColour=&H000000,Outline=2'" -c:a copy -y $OUTPUT_PATH inputs: ["INPUT_FILE", "SUBTITLE_FILE"] outputs: ["OUTPUT_FILE"] output_path: "{{userId}}/{{pipelineId}}/{{runId}}/subtitled/{{baseFilename}}_subtitled.mp4" editor: output: "mp4" preset: "medium" selectedCode: "custom" next_bucket: "public-processed" keep: true render: project_name: "subtitle-generation" status: "queued" public: false
The keep: true flag on the last step tells the transpiler to send the output directly to the final bucket (public-processed). Intermediate steps use next_bucket to pass the result to the next step's trigger. The runId is computed deterministically from the input file name (using mode: "deterministic" and template: "{baseFilename}"). This ensures all steps in the sequential pipeline compute the same run ID, grouping all outputs for a single upload under one folder.
Running the Transpiler
Download the transpiler and the SVG generator:
Run the transpiler to generate the migration files:
Add the --svg flag to also generate a visual graph of your pipeline:
The output will be:
Apply the migration to your Supabase database:
Visualising the Pipeline
The generated SVG gives you a clear overview of your pipeline. Steps marked with KEEP are green – their outputs are permanently stored in the final bucket. Edges are labelled with the bucket they use for data flow.
In the graph above, the steps run sequentially. Each step triggers on the private-uploads bucket but with different file type conditions (video, WAV, SRT). The final step outputs to public-processed with the subtitled video. All steps share the same runId, so all outputs are grouped under {userId}/subtitle-pipeline/{runId}/.
FFmpeg Commands
The YAML steps define the following FFmpeg commands using placeholders:
$MEDIA_1— The path to the downloaded input file (resolved by the runner).$OUTPUT_PATH— The temporary path for the output file (resolved by the runner).
1. Extract Audio
-ac 1— Convert to mono (Whisper works best with mono audio)-ar 16000— Resample to 16kHz (Whisper's optimal sample rate)-vn— Disable video output (audio only)-y— Overwrite output file without prompting
2. Whisper Transcription
--model— Path to the Whisper model (base, small, medium, large)--language— Language code (auto, en, es, fr, etc.)--output-srt— Output SRT subtitle file
3. Burn Subtitles into Video
subtitles=$MEDIA_2— Input SRT file (second input)force_style— Styling options:FontName=Arial— Font familyFontSize=24— Font size in pointsPrimaryColour=&HFFFFFF— White text (HEX: &HBBGGRR)OutlineColour=&H000000— Black outlineOutline=2— Outline width-c:a copy— Copy audio stream without re-encoding-y— Overwrite output file without prompting
Setting Up Whisper
The runner needs Whisper installed. You can use the Python version or whisper.cpp for better performance.
Option 1: whisper.cpp (Recommended for CPU)
Option 2: Python Whisper (More Features)
Docker Integration
Configure ffmpeglab-runner
The runner needs to be configured to poll the render queue and execute the provided commands. The transpiler uses the existing render queue.
.env file or Docker Compose configuration.# 2. For each job:
# a. Download the input file from private-uploads
# b. Parse the 'commands' array from the job payload
# c. Execute each command:
# - Extract audio (ffmpeg)
# - Run Whisper transcription
# - Burn subtitles (ffmpeg)
# d. Upload outputs to public-processed
# e. Update render table with progress and logs
# f. Mark job complete and delete from queue
Monitor the Pipeline
You can monitor the pipeline using SQL queries and notifications.
render table for job status.Customising the Pipeline
Change Whisper Model Size
Replace $WHISPER_MODEL with a different model path:
ggml-base.bin— Fast, ~50x real-time, ~85% accuracyggml-small.bin— Balanced, ~30x real-time, ~90% accuracyggml-medium.bin— Accurate, ~15x real-time, ~95% accuracyggml-large.bin— Best accuracy, ~5x real-time
Change Subtitle Language
Set WHISPER_LANGUAGE to a specific language code:
en— Englishes— Spanishfr— Frenchde— Germanzh— Chineseja— Japanesehi— Hindiauto— Auto-detect
Customise Subtitle Styling
Modify the force_style in the burn step:
Add Multiple Language Subtitles
To generate subtitles in multiple languages, duplicate the transcribe step with different language settings and burn steps for each language.
Frequently Asked Questions (FAQ)
How does the automated subtitle pipeline work?
When a video is uploaded, a PostgreSQL trigger fires and pushes a job to pgmq. The ffmpeglab-runner extracts the audio, transcribes it using Whisper AI, generates an SRT subtitle file, and burns it into the video using FFmpeg's subtitles filter. The result is a video with embedded subtitles, stored in a public bucket.
What AI model is used for transcription?
The pipeline uses OpenAI's Whisper model (whisper.cpp implementation) for speech-to-text transcription. Whisper is a state-of-the-art multilingual model that supports 99 languages and runs efficiently on CPU, GPU, or Apple Silicon.
What languages are supported?
Whisper supports 99 languages including English, Spanish, French, German, Chinese, Japanese, Hindi, and many more. You can specify the language in the pipeline configuration or let Whisper auto-detect it.
Can I customize the subtitle styling?
Yes. The FFmpeg subtitles filter supports styling options including font, size, color, shadow, outline, and positioning. You can customize these in the runner's FFmpeg command.
How accurate is the transcription?
Whisper achieves state-of-the-art accuracy, with word error rates (WER) as low as 2-5% for English in clean audio. Accuracy depends on audio quality, background noise, and accents. Using a larger model (medium or large) improves accuracy.
Does this pipeline create new tables?
No. The pipeline uses the existing render and logpiece tables from the FFmpegLab server. It only adds storage buckets, RLS policies, and the trigger function — no table conflicts.
Final Word
You now have a fully automated subtitle generation pipeline defined in YAML and generated via a transpiler. With PostgreSQL triggers, pgmq, and Supabase Storage, you get:
- AI-powered transcription — state-of-the-art Whisper model
- 99 language support — global reach
- Real‑time notifications — know when processing is complete
- Full observability — monitoring views and logs
- Scalable architecture — handles any volume of videos
- Declarative YAML — version-controlled, reusable pipeline definitions
- Per‑run grouping via deterministic
runId– all outputs for one upload stay together.
The pipeline is production‑ready, scalable, and configurable. It uses the existing render and logpiece tables from the FFmpegLab server, so there are no table conflicts — just pure, automated subtitle generation.