>FFmpegLab Sign In
FFmpegLab Guide

Video Labeling with DNN Filters – YAML & SQL + Object Detection Pipeline

Declarative YAML pipeline for automated video labeling using FFmpeg's DNN filters — detect objects, faces, and scenes with SQL triggers and Supabase Storage.

Video understanding is one of the most powerful applications of AI. Being able to automatically detect objects, faces, and scenes in videos opens up a world of possibilities — from content moderation and search to analytics and accessibility.

FFmpeg's DNN filters — dnn_detect and dnn_classify — make this possible. With pre‑trained models like YOLO, ResNet, and face detection models, you can label every frame of your video with rich metadata.

This guide shows you how to build a fully automated video labeling pipeline using a declarative YAML template and a transpiler that generates the complete SQL migration. The pipeline:

Key takeaways

The Gap: Video Understanding at Scale

Video is the most data‑rich medium we have. But without understanding what's in the video, it's just pixels. Manually labeling videos is impossible at scale. Automated video labeling using DNN filters makes it practical:

What if the pipeline could be fully automated — triggered by the upload itself, running DNN detection and classification in the background? And what if you could define that pipeline in a declarative YAML file that you can version, share, and reuse?

This guide shows you exactly how to build that pipeline.

Architecture Overview

User Uploads Video → Supabase Storage → PostgreSQL Trigger → pgmq Queue → ffmpeglab-runner
dnn_detect → Bounding Boxes → dnn_classify → Labels
JSON Metadata → Public Folder → pg_notify → User Notified

The pipeline consists of:

Important: This pipeline uses the existing render and logpiece tables from the FFmpegLab server. It does not create new tables — it only adds the pipeline components.

What the Pipeline Delivers

OutputFormatLocation
Detection MetadataJSON (bounding boxes, labels, confidence)public-processed/{userId}/{pipelineId}/{runId}/labels/
Labeled VideoMP4 (with bounding boxes overlaid)public-processed/{userId}/{pipelineId}/{runId}/labeled/
Real‑time notificationspg_notify channelsN/A
Job trackingrender tableExisting FFmpegLab table
Logslogpiece tableExisting FFmpegLab table

Prerequisites

The YAML‑Driven Approach

While you can write the SQL directly, the recommended way is to use the YAML transpiler. This gives you:

The transpiler is a single TypeScript file that reads your YAML and generates the SQL migration. It runs with Deno and has zero external dependencies (except yaml for parsing).

The YAML Template

Create a file called labeling-pipeline.yaml with the following content. It defines the buckets, RLS policies, and each processing step. The runId section configures how the per‑run ID is generated — in this case, deterministically from the input file name.

labeling-pipeline.yaml
name: "Video Labeling with DNN Filters"
pipelineId: "video-labeling"
runId:
  mode: "deterministic"
  template: "{baseFilename}"
description: "Detect objects, faces, and scenes using FFmpeg DNN filters"
version: "1.0.0"

editor:
  compressionLevel: 23
  preset: "medium"
  aspectRatio: "16:9"
  framerate: 30
  opacity: 1.0
  output: "mp4"

storage:
  output_bucket: "public-processed"
  buckets:
    - name: "dnn-uploads"
      public: false
      allowed_mime_types:
        - "video/mp4"
        - "video/quicktime"
        - "video/x-msvideo"
        - "video/webm"
        - "video/mpeg"
    - name: "public-processed"
      public: true
      allowed_mime_types:
        - "video/mp4"
        - "application/json"

  rls_policies:
    - name: "Users can upload to their own folder"
      operation: "INSERT"
      role: "authenticated"
      condition: |
        bucket_id = 'dnn-uploads' AND
        (storage.foldername(name))[1] = auth.uid()::text

    - name: "Users can download from their own folder"
      operation: "SELECT"
      role: "authenticated"
      condition: |
        bucket_id = 'dnn-uploads' AND
        (storage.foldername(name))[1] = auth.uid()::text

    - name: "Public read access to processed media"
      operation: "SELECT"
      role: "anon"
      condition: |
        bucket_id = 'public-processed'

    - name: "Service role can manage processed media"
      operation: "ALL"
      role: "service_role"
      condition: |
        bucket_id = 'public-processed'

    - name: "Users can read their own processed media"
      operation: "SELECT"
      role: "authenticated"
      condition: |
        bucket_id = 'public-processed' AND
        (storage.foldername(name))[1] = auth.uid()::text

steps:
  # Step 1: Object/Face Detection (dnn_detect)
  - id: "detect_objects"
    trigger:
      name: "handle_detect"
      event: "INSERT"
      table: "storage.objects"
      condition: |
        NEW.bucket_id = 'dnn-uploads' AND
        NEW.metadata->>'mimetype' LIKE 'video/%'
    command: -i $MEDIA_1 -vf "dnn_detect=dnn_backend=$DNN_BACKEND:model=$DETECT_MODEL:input=data:output=detection_out:confidence=$DETECT_CONFIDENCE:labels=$DETECT_LABELS,showinfo" -f null -
    inputs: ["INPUT_FILE"]
    outputs: ["OUTPUT_FILE"]
    output_path: "{{userId}}/{{pipelineId}}/{{runId}}/labels/{{baseFilename}}_detections.json"
    editor:
      output: "json"
      preset: "medium"
      selectedCode: "custom"
    next_bucket: "dnn-uploads"
    keep: false

  # Step 2: Scene/Emotion Classification (dnn_classify)
  - id: "classify_scenes"
    trigger:
      name: "handle_classify"
      event: "INSERT"
      table: "storage.objects"
      condition: |
        NEW.bucket_id = 'dnn-uploads' AND
        NEW.name LIKE '%.json' AND
        NEW.name NOT LIKE '%.emptyFolderPlaceholder'
    command: -i $MEDIA_1 -vf "dnn_classify=dnn_backend=$DNN_BACKEND:model=$CLASSIFY_MODEL:input=data:output=$CLASSIFY_OUTPUT:confidence=$CLASSIFY_CONFIDENCE:labels=$CLASSIFY_LABELS,showinfo" -f null -
    inputs: ["INPUT_FILE"]
    outputs: ["OUTPUT_FILE"]
    output_path: "{{userId}}/{{pipelineId}}/{{runId}}/labels/{{baseFilename}}_classifications.json"
    editor:
      output: "json"
      preset: "medium"
      selectedCode: "custom"
    next_bucket: "dnn-uploads"
    keep: false

  # Step 3: Overlay bounding boxes on video
  - id: "overlay_labels"
    trigger:
      name: "handle_overlay"
      event: "INSERT"
      table: "storage.objects"
      condition: |
        NEW.bucket_id = 'dnn-uploads' AND
        NEW.name LIKE '%.json' AND
        NEW.name NOT LIKE '%.emptyFolderPlaceholder'
    command: -i $MEDIA_1 -vf "dnn_detect=dnn_backend=$DNN_BACKEND:model=$DETECT_MODEL:input=data:output=detection_out:confidence=$DETECT_CONFIDENCE:labels=$DETECT_LABELS,drawbox=x=1005:y=813:w=81:h=92:color=red" -c:v libx264 -crf 18 -y $OUTPUT_PATH
    inputs: ["INPUT_FILE"]
    outputs: ["OUTPUT_FILE"]
    output_path: "{{userId}}/{{pipelineId}}/{{runId}}/labeled/{{baseFilename}}_labeled.mp4"
    editor:
      output: "mp4"
      preset: "medium"
      selectedCode: "custom"
    next_bucket: "public-processed"
    keep: true

render:
  project_name: "video-labeling"
  status: "queued"
  public: false

The keep: true flag on the last step tells the transpiler to send the output directly to the final bucket (public-processed). Intermediate steps use next_bucket to pass the result to the next step's trigger. The runId is computed deterministically from the input file name (using mode: "deterministic" and template: "{baseFilename}"). This ensures all steps in the sequential pipeline compute the same run ID, grouping all outputs for a single upload under one folder.

Running the Transpiler

Download the transpiler and the SVG generator:

# Download transpiler and SVG generator curl -O https://raw.githubusercontent.com/ffmpeglab/server/main/sdk/yaml/transpiler.ts curl -O https://raw.githubusercontent.com/ffmpeglab/server/main/sdk/yaml/svg.ts

Run the transpiler to generate the migration files:

# Generate the migration files deno run --allow-read --allow-write transpiler.ts labeling-pipeline.yaml ./supabase/migrations

Add the --svg flag to also generate a visual graph of your pipeline:

# Generate migration files + SVG graph deno run --allow-read --allow-write transpiler.ts labeling-pipeline.yaml ./supabase/migrations --svg

The output will be:

✅ Migration files created: UP: ./supabase/migrations/20260807120000_video-labeling.sql DOWN: ./supabase/migrations/20260807120000_video-labeling_down.sql SVG: ./supabase/migrations/20260807120000_video-labeling.svg

Apply the migration to your Supabase database:

# Via psql psql -U postgres -d your_database -f ./supabase/migrations/20260807120000_video-labeling.sql

Visualising the Pipeline

The generated SVG gives you a clear overview of your pipeline. Steps marked with KEEP are green – their outputs are permanently stored in the final bucket. Edges are labelled with the bucket they use for data flow.

→ dnn-uploads → dnn-uploads → public-processed 📤 dnn-uploads detect_objects 📁 {userId}/video-labeling/{filename}/lab… → dnn-uploads json classify_scenes 📁 {userId}/video-labeling/{filename}/lab… → dnn-uploads json overlay_labels KEEP 📁 {userId}/video-labeling/{filename}/lab… → public-processed

In the graph above, the steps run sequentially. The first step detects objects and faces, the second classifies scenes and emotions, and the final step overlays bounding boxes on the video. All steps share the same runId, so all outputs are grouped under {userId}/video-labeling/{runId}/.

FFmpeg Commands

The YAML steps define the following FFmpeg commands using placeholders:

1. Object / Face Detection (dnn_detect)

ffmpeg -i $MEDIA_1 -vf "dnn_detect=dnn_backend=openvino:model=/app/models/detection/face-detection-adas-0001.xml:input=data:output=detection_out:confidence=0.6:labels=/app/models/detection/face-detection-adas-0001.label,showinfo" -f null -
💡
dnn_detect Parameters Explained
  • dnn_detect — FFmpeg's object detection filter
  • dnn_backend=openvino — DNN backend (openvino, tensorflow, native)
  • model — Path to the model file
  • input=data — Input tensor name
  • output=detection_out — Output tensor name
  • confidence=0.6 — Confidence threshold
  • labels — Path to labels file
  • showinfo — Prints detection results to console

2. Classification (dnn_classify)

ffmpeg -i $MEDIA_1 -vf "dnn_classify=dnn_backend=openvino:model=/app/models/classification/emotions-recognition-retail-0003.xml:input=data:output=prob_emotion:confidence=0.3:labels=/app/models/classification/emotions-recognition-retail-0003.label,showinfo" -f null -
💡
dnn_classify Parameters Explained
  • dnn_classify — FFmpeg's classification filter
  • output=prob_emotion — Output tensor for probabilities
  • confidence=0.3 — Confidence threshold

3. Overlay Detection Results (drawbox)

ffmpeg -i $MEDIA_1 -vf "dnn_detect=...,drawbox=x=1005:y=813:w=81:h=92:color=red" -c:v libx264 -crf 18 -y $OUTPUT_PATH
💡
Overlay Parameters Explained
  • drawbox — FFmpeg filter for drawing rectangles
  • x, y, w, h — Position and size of the bounding box
  • color=red — Color of the bounding box
  • -c:v libx264 -crf 18 — High-quality encoding
  • -y — Overwrite output file

Model Management

Directory Structure

/app/models/ ├── detection/ │ ├── face-detection-adas-0001.xml │ ├── face-detection-adas-0001.bin │ └── face-detection-adas-0001.label ├── classification/ │ ├── emotions-recognition-retail-0003.xml │ ├── emotions-recognition-retail-0003.bin │ └── emotions-recognition-retail-0003.label └── object-detection/ ├── yolo-v3-tiny.xml ├── yolo-v3-tiny.bin └── coco.names

Downloading Models

# Create directories mkdir -p models/detection models/classification models/object-detection # Face Detection Model (OpenVINO) wget -O models/detection/face-detection-adas-0001.xml \ https://github.com/guoyejun/ffmpeg_dnn/raw/main/models/openvino/2021.1/face-detection-adas-0001.xml wget -O models/detection/face-detection-adas-0001.bin \ https://github.com/guoyejun/ffmpeg_dnn/raw/main/models/openvino/2021.1/face-detection-adas-0001.bin wget -O models/detection/face-detection-adas-0001.label \ https://github.com/guoyejun/ffmpeg_dnn/raw/main/models/openvino/2021.1/face-detection-adas-0001.label # Emotion Recognition Model (OpenVINO) wget -O models/classification/emotions-recognition-retail-0003.xml \ https://github.com/guoyejun/ffmpeg_dnn/raw/main/models/openvino/2021.1/emotions-recognition-retail-0003.xml wget -O models/classification/emotions-recognition-retail-0003.bin \ https://github.com/guoyejun/ffmpeg_dnn/raw/main/models/openvino/2021.1/emotions-recognition-retail-0003.bin wget -O models/classification/emotions-recognition-retail-0003.label \ https://github.com/guoyejun/ffmpeg_dnn/raw/main/models/openvino/2021.1/emotions-recognition-retail-0003.label

Configure ffmpeglab-runner

The runner needs to be configured to poll the render queue and execute the provided commands. The transpiler uses the existing render queue.

Step 1
Add the queue to your environment
Add the following to your .env file or Docker Compose configuration.
# Render queue (already used by FFmpegLab server) RENDER_QUEUE_NAME=render # Detection model DNN_DETECT_MODEL=/app/models/detection/face-detection-adas-0001.xml DNN_DETECT_LABELS=/app/models/detection/face-detection-adas-0001.label # Classification model DNN_CLASSIFY_MODEL=/app/models/classification/emotions-recognition-retail-0003.xml DNN_CLASSIFY_LABELS=/app/models/classification/emotions-recognition-retail-0003.label DNN_CLASSIFY_OUTPUT=prob_emotion # DNN backend DNN_BACKEND=openvino # Confidence thresholds DETECT_CONFIDENCE=0.6 CLASSIFY_CONFIDENCE=0.3
Step 2
Implement the processing loop
The runner should execute the following steps for each job:
# 1. Connect to Supabase and listen to the queue

# 2. For each job:
# a. Download the input file
# b. Parse the 'commands' array
# c. Execute each command:
# - Run dnn_detect (extract JSON metadata)
# - Run dnn_classify (extract JSON metadata)
# - Run overlay (generate labeled video)
# d. Upload outputs to public-processed
# e. Update render table
# f. Mark job complete and delete from queue
Step 3
Restart the runner
After updating the environment, restart the runner service.
docker compose restart ffmpeglab-runner

Monitor the Pipeline

You can monitor the pipeline using SQL queries and notifications.

Step 1
Check queued jobs
Query the render table to see queued jobs.
SELECT * FROM "render" WHERE status = 'queued' ORDER BY created_at DESC;
Step 2
Check render status
Query the existing render table for job status.
SELECT id, title, status, progress, data FROM "render" WHERE project = 'video-labeling' ORDER BY created_at DESC;
Step 3
Listen to notifications
In your application, listen for real‑time updates.
-- In your PostgreSQL client: LISTEN render_status_channel; LISTEN log_channel;
Step 4
Check processed files
List all processed files in the public bucket.
SELECT name, metadata, created_at FROM storage.objects WHERE bucket_id = 'public-processed' ORDER BY created_at DESC;

Customising the Pipeline

Use a Different Detection Model (YOLO)

Replace DETECT_MODEL and DETECT_LABELS with YOLO paths:

# YOLO model configuration DNN_DETECT_MODEL=/app/models/object-detection/yolo-v3-tiny.xml DNN_DETECT_LABELS=/app/models/object-detection/coco.names DNN_DETECT_INPUT=image DNN_DETECT_OUTPUT=detector/yolo-v3-tiny/Identity_1

Update the YAML command:

command: -i $MEDIA_1 -vf "dnn_detect=dnn_backend=$DNN_BACKEND:model=$DETECT_MODEL:input=$DETECT_INPUT:output=$DETECT_OUTPUT:confidence=$DETECT_CONFIDENCE:labels=$DETECT_LABELS,showinfo" -f null -

Customise Overlay Styling

Modify the drawbox parameters in the overlay step:

command: -i $MEDIA_1 -vf "dnn_detect=...,drawbox=x=1005:y=813:w=81:h=92:color=green:thickness=3" -c:v libx264 -crf 18 -y $OUTPUT_PATH

Add Multi‑Frame Detection

To detect objects across multiple frames, use the dnn_detect filter with frame caching:

command: -i $MEDIA_1 -vf "dnn_detect=...,select='gte(n,10)'" -frames:v 100 -f null -

Frequently Asked Questions (FAQ)

What does the video labeling pipeline do?

The pipeline automatically detects and labels objects, faces, and scenes in uploaded videos using FFmpeg's dnn_detect and dnn_classify filters. It generates bounding boxes, labels, and confidence scores, and stores them as JSON metadata alongside the video.

What models are used?

The pipeline supports multiple models including YOLO for object detection, ResNet for classification, and face detection models. The default configuration uses OpenVINO models for face detection (face-detection-adas-0001), classification (emotions-recognition-retail-0003), and can be extended for general object detection.

What is the output format?

The pipeline outputs a JSON file containing all detection results: frame number, bounding box coordinates, label/class name, confidence score, and timestamp. This can be used for search, analytics, or further processing.

Can I use custom models?

Yes. You can use any OpenVINO or TensorFlow model that works with FFmpeg's dnn_detect or dnn_classify filters. You'll need to provide the model files and configure the input/output tensor names.

Does this pipeline create new tables?

No. The pipeline uses the existing render and logpiece tables from the FFmpegLab server. It only adds storage buckets, RLS policies, and the trigger function — no table conflicts.

Final Word

You now have a fully automated video labeling pipeline defined in YAML and generated via a transpiler. With PostgreSQL triggers, pgmq, and Supabase Storage, you get:

The pipeline is production‑ready, scalable, and extensible — you can swap in any DNN model for your specific use case.