I indexed 37h of my YouTube library using an RTX 4090 and local ML models in 24h
by Ilias Haddad

Following my recent blog post and Hacker News post. where I ran the desktop app on my M1 Max. Today, I’m using the self-hosted version running in Docker with an NVIDIA RTX 4090 (24 GB VRAM).
The content is also fundamentally more demanding: long podcast episodes with at least two faces in every frame, coding tutorials packed with on-screen text, and screen recordings. GoPro footage is mostly wide outdoor shots.
But NVIDIA was faster than my M1 Max.
Here’s the full spec of the rented server:

Here’s my YouTube channel: https://www.youtube.com/@iliashaddad_dev
The key difference from the GoPro run: this time I used the source-available Docker version (github.com/IliasHad/edit-mind) rather than the desktop app, running on a machine with an RTX 4090 and CUDA 12.8.0 (I rented from a cloud provider)
Summary
| Metric | Value |
|---|---|
| Edit Mind version | v0.30.0 |
| Hardware | NVIDIA RTX 4090 (24 GB VRAM) |
| Runtime | Docker + CUDA 12.8.0 |
| Videos indexed | 87 |
| Total footage duration | 37h 37m 56s |
| Total compute time | 24h 13m 1s |
| Total frames analyzed | 54,054 |
The indexing process
The self-hosted version runs the same pipeline as the desktop app. Point it at a folder, and it works through each video in stages:
- Transcription: the full audio track is transcribed using OpenAI Whisper
- Frame analysis: the video is divided into scenes (1-2.5 seconds), and each frame runs through a set of ML plugins: face recognition, on-screen text detection, object detection, scene description, dominant color, and shot type
- Embedding: scene data is embedded into a local vector database across three collections: text, visual, and audio
Stage breakdown
| Stage | Total | Avg/video | % of compute |
|---|---|---|---|
| Frame analysis | 20h 41m 30s | 14m 26s | 85.4% |
| Transcription | 1h 26m 36s | 1m 0s | 6.0% |
| Audio embedding | 54m 13s | 37s | 3.7% |
| Visual embedding | 42m 8s | 29s | 2.9% |
| Text embedding | 27m 14s | 19s | 1.9% |
| Scene creation | 1m 20s | < 1s | 0.1% |
- Coding tutorials and screen recordings:
TextDetectionPluginIt is the most expensive plugin - Podcast and interview episodes:
FaceRecognitionPluginandDescriptorPlugindominate, the face model runs significantly harder than it ever does on my GoPro videos.
Fastest video
A Day in the Life of a Remote Software Engineer in Malaysia | Work, Life, and Culture.mp4
| Metric | Value |
|---|---|
| Footage duration | 4m 25s |
| Total compute time | 1m 7s |
| Speed ratio | 0.25× |
| Frames analyzed | 107 |
| Frame analysis time | 49s |
| Transcription time | 6s |
Slowest video
Note: (excl. short clips)
| Metric | Value |
|---|---|
| Footage duration | 58m 45s |
| Total compute time | 53m 52s |
| Speed ratio | 0.917× |
| Frames analyzed | 1,411 |
| Frame analysis time | 47m 32s |
| Transcription time | 3m 33s |
Frame analysis plugin averages
| Plugin | Avg duration/video | Avg time/frame |
|---|---|---|
| DescriptorPlugin | 127.6s | 0.293s |
| TextDetectionPlugin | 120.6s | 0.277s |
| FaceRecognitionPlugin | 103.7s | 0.238s |
| DominantColorPlugin | 13.5s | 0.031s |
| ObjectDetectionPlugin | 5.8s | 0.013s |
| ShotTypePlugin | 0.03s | < 0.001s |
Longest video
EP 10: Building a Shopify App In Public (Part 1).mp4
| Metric | Value |
|---|---|
| Footage duration | 3h 12m 9s |
| Total compute time | 1h 52m 43s |
| Speed ratio | 0.587× |
| Frames analyzed | 4,612 |
| Frame analysis time | 1h 40m 59s |
| Transcription time | 4m 17s |
Plugin breakdown:
| Plugin | Time | Time/frame | % of frame analysis |
|---|---|---|---|
| TextDetectionPlugin | 1h 8m 47s | 0.895s | 68.1% |
| DescriptorPlugin | 23m 34s | 0.307s | 23.4% |
| FaceRecognitionPlugin | 3m 42s | 0.048s | 3.7% |
| DominantColorPlugin | 1m 52s | 0.024s | 1.9% |
| ObjectDetectionPlugin | 46s | 0.010s | 0.8% |
Here are a couple of example video clips using these prompts:
Generate me a compilation of videos where transcription has the word "welcome” - https://youtu.be/OButxl1-A0k
Generate me a compilation of videos where transcription has the word "edit mine" - https://youtu.be/OzauaIc5kOk
Here are a couple of example screenshots using the search:




Wrapping up
The big takeaway: an RTX 4090 running Edit Mind in Docker via CUDA is faster than my M1 Max desktop app for this kind of workload, even on more demanding content.
And the best part is you can use the desktop app as a client for the self-hosted Docker server to utilize the video editing software integration and native MacOS desktop experience
If you wanna check the full processing jobs data, here’s the JSON file. And you can search the video file name in my YouTube video to watch the full video.
What’s next?
If you want to try it yourself, the self-hosted version is at github.com/IliasHad/edit-mind. Add a folder, and it handles the rest. The desktop app (with integrations for Final Cut Pro, DaVinci Resolve, and Adobe Premiere Pro) is available at edit-mind.com.
If you have ideas for performance improvements, new plugins, or just want to share what you're working on in the same space, open an issue or create a pull request.
- Video