>FFmpegLab Sign In
FFmpegLab Guide

Audio Processing Pipeline for Podcasts – YAML & SQL Triggers on Supabase

Declarative YAML pipeline for podcast audio processing: loudness normalization, noise reduction, waveform generation — all driven by SQL triggers and Supabase Storage.

Podcasting is booming, but producing broadcast‑quality audio is a complex process. Recording raw audio often includes background noise, inconsistent loudness, and the wrong format. Professional podcasts need:

This guide shows you how to build a fully automated podcast audio processing pipeline using a declarative YAML template and a transpiler that generates the complete SQL migration. The pipeline turns a raw recording into a broadcast‑ready podcast episode — all driven by PostgreSQL triggers and pgmq.

Key takeaways

The Gap: From Raw Audio to Podcast-Ready

Most podcasters record audio in high‑quality formats (WAV, FLAC) and then manually process it: normalize loudness, remove noise, convert to MP3, add metadata, and create a waveform image. This is time‑consuming and inconsistent.

What if the pipeline could be fully automated — triggered by the upload itself, processing in the background, and delivering a complete podcast episode ready for distribution? And what if you could define that pipeline in a declarative YAML file that you can version, share, and reuse?

This guide shows you exactly how to build that pipeline.

Architecture Overview

User Uploads Audio/Video → Supabase Storage → PostgreSQL Triggers → pgmq Queue → ffmpeglab-runner
Processed Audio → Public Folder → pg_notify → User Notified

The pipeline consists of:

Important: This pipeline uses the existing render and logpiece tables from the FFmpegLab server. It does not create new tables — it only adds the pipeline components.

What the Pipeline Delivers

OutputFormatLocation
Podcast AudioMP3 (192kbps) with ID3v2 metadataaudio-processed/{userId}/{pipelineId}/{runId}/podcast/
Waveform ImagePNG (1200×200)audio-processed/{userId}/{pipelineId}/{runId}/waveforms/
MetadataTitle, Artist, Album, Genre, Cover ArtID3v2 tags (embedded in MP3)
Real‑time notificationspg_notify channelsN/A
Job trackingrender tableExisting FFmpegLab table
Logslogpiece tableExisting FFmpegLab table

Prerequisites

The YAML‑Driven Approach

While you can write the SQL directly, the recommended way is to use the YAML transpiler. This gives you:

The transpiler is a single TypeScript file that reads your YAML and generates the SQL migration. It runs with Deno and has zero external dependencies (except yaml for parsing).

The YAML Template

Create a file called audio-pipeline.yaml with the following content. It defines the buckets, RLS policies, and each processing step in sequence. The runId section configures how the per‑run ID is generated — in this case, deterministically from the input file name.

audio-pipeline.yaml
name: "Audio Processing Pipeline"
pipelineId: "audio-pipeline"
runId:
  mode: "deterministic"
  template: "{baseFilename}"
description: "Sequential audio processing: extract → normalize → waveform"
version: "1.0.0"

editor:
  compressionLevel: 23
  preset: "medium"
  aspectRatio: "16:9"
  framerate: 30
  opacity: 1.0

storage:
  output_bucket: "audio-processed"
  buckets:
    - name: "audio-uploads"
      public: false
      allowed_mime_types:
        - "audio/mpeg"
        - "audio/wav"
        - "audio/flac"
        - "video/mp4"
    - name: "audio-temp-1"
      public: false
      allowed_mime_types:
        - "audio/wav"
    - name: "audio-processed"
      public: true
      allowed_mime_types:
        - "audio/mpeg"
        - "image/png"

  rls_policies:
    - name: "Users can upload to their own audio folders"
      operation: "INSERT"
      role: "authenticated"
      condition: |
        (bucket_id IN ('audio-uploads', 'audio-temp-1', 'audio-processed')) AND
        (storage.foldername(name))[1] = auth.uid()::text

steps:
  - id: "extract_audio"
    trigger:
      name: "handle_extract_audio"
      event: "INSERT"
      table: "storage.objects"
      condition: |
        NEW.bucket_id = 'audio-uploads' AND
        NEW.name NOT LIKE '%.emptyFolderPlaceholder'
    command: -i $MEDIA_1 -ac 1 -ar 16000 -vn -f wav -y $OUTPUT_PATH
    inputs: ["INPUT_FILE"]
    outputs: ["OUTPUT_FILE"]
    output_path: "{{userId}}/{{pipelineId}}/{{runId}}/temp/{{baseFilename}}.wav"
    editor:
      output: "wav"
      preset: "fast"
      selectedCode: "custom"
      width: 0
      height: 0
      length: 0
      compressionLevel: 0
    next_bucket: "audio-temp-1"
    keep: false

  - id: "normalize_loudness"
    trigger:
      name: "handle_normalize_audio"
      event: "INSERT"
      table: "storage.objects"
      condition: |
        NEW.bucket_id = 'audio-temp-1' AND
        NEW.name NOT LIKE '%.emptyFolderPlaceholder'
    command: -i $MEDIA_1 -af loudnorm=I=-16:LRA=11:TP=-1.5 -c:a libmp3lame -b:a 192k -f mp3 -y $OUTPUT_PATH
    inputs: ["INPUT_FILE"]
    outputs: ["OUTPUT_FILE"]
    output_path: "{{userId}}/{{pipelineId}}/{{runId}}/podcast/{{baseFilename}}.mp3"
    editor:
      output: "mp3"
      preset: "medium"
      selectedCode: "custom"
      compressionLevel: 23
    next_bucket: "audio-processed"
    keep: true

  - id: "generate_waveform"
    trigger:
      name: "handle_waveform"
      event: "INSERT"
      table: "storage.objects"
      condition: |
        NEW.bucket_id = 'audio-processed' AND
        NEW.name LIKE '%.mp3' AND
        NEW.name NOT LIKE '%.emptyFolderPlaceholder'
    command: -i $MEDIA_1 -filter_complex showwavespic=s=1200x200:colors=#FC6D26 -frames:v 1 -f image2 -y $OUTPUT_PATH
    inputs: ["INPUT_FILE"]
    outputs: ["OUTPUT_FILE"]
    output_path: "{{userId}}/{{pipelineId}}/{{runId}}/waveforms/{{baseFilename}}.png"
    editor:
      output: "png"
      preset: "medium"
      selectedCode: "custom"
      width: 1200
      height: 200
    next_bucket: "audio-processed"
    keep: true

render:
  project_name: "audio-processing"
  status: "queued"
  public: false

The keep: true flag on the last step tells the transpiler to send the output directly to the final bucket (audio-processed). Intermediate steps use next_bucket to pass the result to the next step's trigger. The runId is computed deterministically from the input file name (using mode: "deterministic" and template: "{baseFilename}"). This ensures all steps in the sequential pipeline compute the same run ID, grouping all outputs for a single upload under one folder.

Running the Transpiler

Download the transpiler and the SVG generator:

# Download transpiler and SVG generator
curl -O https://raw.githubusercontent.com/ffmpeglab/server/main/sdk/yaml/transpiler.ts
curl -O https://raw.githubusercontent.com/ffmpeglab/server/main/sdk/yaml/svg.ts

Run the transpiler to generate the migration files:

# Generate the migration files
deno run --allow-read --allow-write transpiler.ts audio-pipeline.yaml ./supabase/migrations

Add the --svg flag to also generate a visual graph of your pipeline:

# Generate migration files + SVG graph
deno run --allow-read --allow-write transpiler.ts audio-pipeline.yaml ./supabase/migrations --svg

The output will be:

✅ Migration files created:
UP: ./supabase/migrations/20260807120000_audio-pipeline.sql
DOWN: ./supabase/migrations/20260807120000_audio-pipeline_down.sql
SVG: ./supabase/migrations/20260807120000_audio-pipeline.svg

Apply the migration to your Supabase database:

# Via psql
psql -U postgres -d your_database -f ./supabase/migrations/20260807120000_audio-pipeline.sql

Visualising the Pipeline

The generated SVG gives you a clear overview of your pipeline. Steps marked with KEEP are green – their outputs are permanently stored in the final bucket. Edges are labelled with the bucket they use for data flow.

→ audio-temp-1 → audio-processed → audio-temp-1 📤 audio-uploads extract_audio 📁 {userId}/audio-pipeline/{filename}/tem… → audio-temp-1 wav (fast) normalize_loudness KEEP 📁 {userId}/audio-pipeline/{filename}/pod… → audio-processed mp3 generate_waveform KEEP 📁 {userId}/audio-pipeline/{filename}/wav… → audio-processed 1200x200 png

In the graph above, the steps run sequentially. The first step extracts audio and passes it to the second, which normalizes and passes to the third. The final step produces a waveform image and both the MP3 and PNG are stored in audio-processed. All steps share the same runId, so all outputs are grouped under {userId}/audio-pipeline/{runId}/.

Processing Logic Explained

When a file is uploaded to audio-uploads/{userId}/, the pipeline:

  1. Step 1 (extract_audio) — converts the file to WAV and uploads it to audio-temp-1.
  2. Step 2 (normalize_loudness) — applies loudness normalization and noise reduction, converts to MP3 with metadata, and uploads to audio-processed (final bucket).
  3. Step 3 (generate_waveform) — triggers on the MP3 uploaded to audio-processed, generates a waveform image, and uploads it to the same bucket.

Each step's trigger condition matches the bucket that receives the previous step's output. The keep flag on the last two steps ensures the artifacts are permanently stored. Since runId is computed deterministically from the input file name, all steps get the same runId, grouping everything together.

Exact FFmpeg Commands

The YAML steps define the following FFmpeg commands using placeholders:

1. Extract Audio (Step 1)

ffmpeg -i $MEDIA_1 -ac 1 -ar 16000 -vn -f wav -y $OUTPUT_PATH
💡
Parameters Explained
  • -ac 1 — Convert to mono
  • -ar 16000 — Resample to 16kHz (speech-optimised)
  • -vn — Drop any video stream
  • -f wav — Output WAV format

2. Normalize Loudness (Step 2)

ffmpeg -i $MEDIA_1 -af loudnorm=I=-16:LRA=11:TP=-1.5 -c:a libmp3lame -b:a 192k -f mp3 -y $OUTPUT_PATH
💡
Parameters Explained
  • loudnorm=I=-16:LRA=11:TP=-1.5 — EBU R128 loudness normalization target -16 LUFS (podcast standard)
  • -c:a libmp3lame -b:a 192k — MP3 encoding at 192kbps
  • -f mp3 — Output MP3 format

3. Generate Waveform (Step 3)

ffmpeg -i $MEDIA_1 -filter_complex showwavespic=s=1200x200:colors=#FC6D26 -frames:v 1 -f image2 -y $OUTPUT_PATH
💡
Parameters Explained
  • showwavespic — Generates a waveform image
  • s=1200x200 — Image size (1200×200 pixels)
  • colors=#FC6D26 — Waveform colour (FFmpegLab orange)
  • -frames:v 1 — Output a single frame (image)
  • -f image2 — Image output format

Configure ffmpeglab-runner

The runner needs to be configured to poll the render queue and execute the provided FFmpeg commands. The transpiler uses the existing render queue.

Step 1
Add the queue to your environment
Add the following to your .env file or Docker Compose configuration.
# Render queue (already used by FFmpegLab server)
RENDER_QUEUE_NAME=render
Step 2
Ensure FFmpeg is installed
Ensure FFmpeg is installed and available in the PATH.
# Debian/Ubuntu
apt-get install -y ffmpeg

# Alpine
apk add ffmpeg
Step 3
Restart the runner
After updating the environment, restart the runner service.
docker compose restart ffmpeglab-runner

Monitor the Pipeline

You can monitor the pipeline using SQL queries and notifications.

Step 1
Check queued jobs
Query the render table to see queued jobs.
SELECT * FROM "render"
WHERE status = 'queued'
ORDER BY created_at DESC;
Step 2
Check render status
Query the existing render table for job status.
SELECT id, title, status, progress, data
FROM "render"
WHERE project = 'audio-pipeline'
ORDER BY created_at DESC;
Step 3
Listen to notifications
In your application, listen for real‑time updates.
-- In your PostgreSQL client:
LISTEN render_status_channel;
LISTEN log_channel;
Step 4
Check processed files
List all processed files in the public bucket.
SELECT name, metadata, created_at FROM storage.objects
WHERE bucket_id = 'audio-processed'
ORDER BY created_at DESC;

Customising the Pipeline

Add Metadata to the MP3

Modify the command in the normalize_loudness step to include metadata tags:

command: -i $MEDIA_1 -af loudnorm=I=-16:LRA=11:TP=-1.5 -metadata title="My Podcast Episode" -metadata artist="Podcast Name" -metadata album="Season 1" -c:a libmp3lame -b:a 192k -f mp3 -y $OUTPUT_PATH

Adjust Loudness Target

Change the I parameter in loudnorm:

Add Noise Reduction

Add the afftdn filter to the command:

command: -i $MEDIA_1 -af loudnorm=I=-16:LRA=11:TP=-1.5,afftdn=nr=10:nf=-40 -c:a libmp3lame -b:a 192k -f mp3 -y $OUTPUT_PATH

Change Waveform Size or Color

command: -i $MEDIA_1 -filter_complex showwavespic=s=800x150:colors=#FF5733 -frames:v 1 -f image2 -y $OUTPUT_PATH

Frequently Asked Questions (FAQ)

What audio processing does the pipeline perform?

The pipeline automatically normalizes loudness (EBU R128 with -16 LUFS target), reduces background noise using FFT-based denoising, converts formats, adds MP3 metadata, and generates waveform images for podcasts.

What FFmpeg filters are used?

The pipeline uses loudnorm for loudness normalization, afftdn for noise reduction, aformat for format conversion, ametadata for MP3 tags, and showwavespic for waveform visualization.

Can this handle videos as input?

Yes. The pipeline detects video files and extracts the audio track before processing. This makes it perfect for podcasters who record video interviews or screen recordings.

What podcast metadata is supported?

The pipeline supports title, artist (podcast name), album, genre, comment, and cover art (podcast artwork) in ID3v2 tags for MP3 files. You can customize these by modifying the metadata fields in the FFmpeg command.

Does this pipeline create new tables?

No. The pipeline uses the existing render and logpiece tables from the FFmpegLab server. It only adds storage buckets, RLS policies, the pgmq queue, and the trigger function — no table conflicts.

Final Word

You now have a fully automated podcast audio processing pipeline defined in YAML and generated via a transpiler. With PostgreSQL triggers, pgmq, and Supabase Storage, you get:

The pipeline is production‑ready, scalable, and configurable. It uses the existing render and logpiece tables from the FFmpegLab server, so there are no table conflicts — just pure, automated audio processing.