Skip to content
daily-hour-news·

🔬Kandinsky 6.0 Video: MIT-Licensed Joint Audio-Video Models

TL;DR

Kandinsky 6.0 Video is a family of open diffusion models that generate video and synchronized audio together: Lite at 3B parameters and Pro at 29B. RL post-training cut Pro's speech word error rate by 47%.

Kandinsky 6.0 Video is a family of open diffusion models that generate video and synchronized audio together: Lite at 3B parameters and Pro at 29B. RL post-training cut Pro's speech word error rate by 47%. Code, checkpoints and Diffusers support are public.

Kandinsky 6.0 Video: MIT-Licensed Joint Audio-Video Models — daily-hour-news

Key Points

1

Outputs 5-second Full-HD (1920x1080) clips through built-in super-resolution, with 44 kHz audio and lip-sync

2

Dual-stream CrossDiT joins a pretrained video stream and a new audio stream by bidirectional cross-attention

3

Pipeline runs SFT, RL post-tuning, then distillation

4

Human evals show a significant speech-quality edge over LTX 2.5

5

Released under the MIT license

Why It Matters

Joint audio-video generation was a closed-model feature a year ago. An MIT-licensed 29B option puts it in reach of anyone with a few GPUs.

Quick Facts

Kandinskyvideo generationtext-to-audio-videodiffusionopen sourceMIT licenseHugging Face

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,556 builders reading daily.

Also get