Notes on Transcoding from Building a Content Management System
Transcoding is where a content management system (CMS) for video, photos and audio spends most of its time and resources. Here are the principles we settled on while building one — keeping originals untouched, purpose-built copies, quality-based encoding, watermarking, camera RAW and operating with failure in mind.

We're building a content management system (CMS) that collects and manages video, photos and audio. The visible features are things like AI search and watermarking, but the place that actually consumes the most time and resources is transcoding. From the moment a video arrives through playback, scene analysis and external delivery, every step means encoding the video one more time.
This post walks through the principles we settled on, step by step.
1. Leave originals untouched; make a copy for each purpose
The first principle is to store the original as-is and create a separate copy for each use.
| Purpose | Copy |
|---|---|
| Web playback | Streaming (HLS) copy |
| Scene analysis | Low-resolution frames for analysis |
| Audio preview | Lightweight compressed audio |
| External delivery | Per-recipient watermarked copies (by quality tier) |
We never play or export the original directly. That brings two benefits:
- The original never has to leave the system.
- Copies can always be regenerated, so changing settings is easy.
2. First, check whether the video can actually be decoded
Every upload is checked twice before processing: first we read the file's information (duration, resolution, track layout), then we actually decode the first frame.
A file can end in .mp4 and still use an unknown codec or be corrupted — more often than you'd think. Rather than failing halfway through a long conversion, it's far easier to operate if the system stops within the first few seconds and leaves a human-readable reason such as "unsupported codec" or "no video track".
3. Playback copies: encode by quality, not fixed bitrate
Web playback uses HLS streaming. Resolution is capped, and smaller originals are never upscaled. The video is split into segments of a set length with a keyframe at each segment, so clicking a scene jumps to the exact moment.
At first we used a hardware encoder at a fixed bitrate. It was fast, but a video full of still shots used the same space as one full of motion. So we switched to quality-based encoding with a bitrate cap. On our samples, file size dropped by OO–OO%, while the difference in the VMAF quality metric was under OO.
For when speed matters most, the hardware encoder is still available behind a single setting — whether size or speed matters more depends on the situation.
4. Scene detection is decoding work too
Before AI can describe each scene, the video has to be split into scenes. That also means decoding the whole video from start to finish, so it costs about as much as transcoding.
- Analyze at a reduced size — finding scene changes doesn't need full resolution, and shrinking the frames speeds things up dramatically.
- Thresholds differ by video — motion graphics with few hard cuts don't split well under the same threshold, so we adjust it step by step based on the result.
- Normalize scene length — very short scenes are merged into the previous one and very long ones are split, so AI gets segments of a workable length.
- Several representative frames — one frame per scene misses content, so frames from several points in each scene go to the AI together.
5. A different copy for each recipient: watermarking runs on top of transcoding
When we send video to an external partner, each recipient gets a separate copy carrying an invisible identifier. The video is decoded frame by frame, the identifier is embedded, and the result is encoded again. We chain these steps into a single stream so even long videos are processed without loading the whole file into memory. Recipients can get high, standard or low quality, and audio copies are made per tier in the same way.
The key point is that the watermark must survive whatever transcoding happens afterwards. Leaked videos usually turn up trimmed, downgraded and re-encoded. So we deliberately tested with a short clip re-encoded at lower quality — OO of OO bits of the identifier matched, enough to identify the recipient. We also have a fallback for cases where the picture is heavily damaged.
6. Formats FFmpeg can't decode: camera RAW
Original footage from broadcast and film sets often comes in proprietary camera RAW formats that general-purpose tools such as FFmpeg can't decode (RED R3D, Blackmagic RAW, ARRIRAW, Sony X-OCN and others). We split the architecture so these files are handled as follows:
- Connect the manufacturer's SDK or an editing tool as a converter.
- Convert to a high-quality intermediate file.
- Pass the intermediate file to the existing steps (streaming, scene detection, tagging, watermarking).
MXF uses the same file extension for regular broadcast codecs and RAW, so you only know which it is after opening the file. Photo RAW is excluded from video conversion and routed through the image pipeline instead.
7. Operations: failure will happen
Transcoding is slow and resource-hungry, so we designed it assuming failure from the start.
- Job queue — conversion requests go into a queue and workers take them one at a time, so a burst of requests doesn't bring the server down.
- Automatic retry for temporary errors — retried a limited number of times with increasing intervals.
- Report permanent errors immediately — a corrupted file or one with no video track won't succeed on retry, so the reason is shown right away without retrying.
- Operations dashboard — queued, in-progress and failed counts for conversion, watermarking and email jobs, plus storage usage, at a glance.
For reference, measured on a single laptop, processing one hour of video took about OO minutes, including streaming conversion, scene detection and AI tagging.
Wrapping up
In a CMS, transcoding doesn't stop at producing a playable file. Analysis for search, copies for security and leak tracing all run on top of it. That's why these three things ended up deciding the cost and stability of the whole system:
- Leave originals untouched.
- Make a copy for each purpose.
- Encode by quality, and operate assuming failure.
If you need a system that handles media such as video and audio — or want to add AI search and analysis on top — SJ System can work with you from design through operation.
- #Transcoding
- #CMS
- #Video processing
- #HLS
- #Watermarking
- #FFmpeg
Need an AI that answers precisely from your own material?
We start by diagnosing where RAG fits in your work — and the team that builds it runs it to the end.