Skip to content

comma.ai Video Compression Challenge

Ongoing video-compression research, with three submissions that have appeared on comma.ai’s leaderboard.

Stack
Python, PyTorch, MLX, CUDA, entropy coding
Year
2026
Status
Active research
Repo
adpena/comma-lab
Challenge
Leaderboard

I’m working on comma.ai’s lossy video compression challenge: making a driving video smaller while preserving what the challenge’s segmentation and motion-estimation models see. That means testing both the file size and the effect of compression on those models.

I use comma-lab for experiments, decoders, evaluation tools, and research records. tac contains reusable training, compression, and validation code. I develop across Apple Silicon and Linux GPU environments and check submission archives with the challenge’s evaluator.

Leaderboard submissions

My submissions #107, #110, and #140 have all appeared on the official leaderboard:

These submissions build on credited public work. #110 uses #101’s HNeRV representation; #140 builds on the learned renderer and pose components in #130 and #135, including #133’s contributions. The pull requests describe my optimization, selection, coding, and validation work.

For a more visual introduction to related research, The Witness Machine is an interactive notebook about allocating compression error around what a perception model needs. Its demonstrations are separate from official challenge scores.

The research is ongoing. These leaderboard results apply to the challenge’s fixed video and evaluators.

Public submission reports · 600 samples each

What goes into the score?

The challenge charges for file size and for changes in two perception models’ outputs. Lower is better.

Archive size
180,002 B
Original video
37,545,489 B
Score, approximately
0.1480
Segmentation distortion
0.02014
Pose distortion
0.00798
File-size cost
0.11986

100 × SegNet distortion + √(10 × PoseNet distortion) + 25 × archive/original size. Recomputed from the rounded components in PR #140. Evaluation: CUDA · Tesla T4. Hardware differs between reports, and these values are not a live ranking.

A frame and the model’s segmentation

Reference driving-video frame showing a road and surrounding landscape
Reference frame 393
The reference frame divided into five colored segmentation classes
Model classes, reduced for display

This reference pair comes from Witness Machine. Its cached segmentation was produced on macOS CPU and is a visual diagnostic, separate from the submission reports above. It does not show a reconstruction from any of these three archives.