What actually shipped

The release is a family, not a single model. Kandinsky 6.0 Video Lite (3 billion parameters) and Pro (29 billion) both handle text-to-audio-video and image-to-audio-video: five-second clips with 44kHz sound generated in the same pass as the picture, lip-sync included, upscaled to 1080p by a built-in super-resolution model. The paper, posted to arXiv this week, describes a dual-stream CrossDiT architecture: a pretrained video stream joined to a newly trained audio stream through bidirectional cross-attention, so the two modalities stay aligned instead of being stitched together after the fact.

The training recipe is the current frontier playbook. Continuous pretraining: first the audio stream alone on large audio corpora, then both streams jointly on paired audio-video data. Then supervised fine-tuning, reinforcement-learning post-training, and distillation. The GitHub repo went public on October 5 and holds code, checkpoints, and diffusers integration.

The grades are their own

The paper's evaluation is a side-by-side human comparison in which Pro “clearly outperforms” its predecessor Kandinsky 5.0 and is “competitive” with leading audio-video models, “particularly in speech quality.” Every one of those judgments comes from the authors. The repository is about a day and a half old, with four commits and a handful of stars. Nobody independent has run the weights yet.

That does not make the release uninteresting; it makes the quality claims provisional. Five seconds is short, 1080p arrives through a super-resolution pass rather than native generation, and hardware requirements are not yet documented in the release. Treat the benchmarks as a direction to verify, not a verdict to quote.

The license is the actual story

This is the detail that separates Kandinsky 6.0 from the summer's other “open” releases. When MiniMax published H3's open weights in August, the Community License excluded the United States, the European Union, the United Kingdom, and South Korea from local deployment, and added a $20 million revenue ceiling before companies needed separate written authorization. The hosted API stayed global; the weights were open in name and geo-fenced in practice.

Kandinsky 6.0 ships under plain MIT. No territories, no revenue tiers, no attribution requirements beyond the license notice. Code, checkpoints, and diffusers integration, all of it MIT. For anyone in the markets MiniMax fenced out, this is the first open video-with-audio model of the season that can be downloaded and run locally without a lawyer in the room.

One compliance footnote

Kandinsky Lab is Sber AI, the AI arm of Sberbank, and Sberbank has sat on US and EU sanctions lists since 2022. The MIT license may be clean, but pulling code from a sanctioned entity is a separate compliance question for Western teams. Check that before checking GPU requirements.

Sources

  1. [1] Kandinsky 6.0 Video technical report, arXiv:2610.05608 (Oct 2026)Read source
  2. [2] kandinskylab/kandinsky-6 on GitHub (MIT, public Oct 5, 2026)Read source
  3. [3] explainx.ai — “MiniMax H3 Open Weights: License Excludes US/EU”Read source