The actual new stuff

Vidu's Q4 Preview comes in two modes. Image-to-Video animates a single frame with a text prompt, 3 to 16 seconds. Reference-to-Video is the interesting one: up to 15 image references plus up to three audio references (MP3, 3 to 12 seconds each), with the model keeping character, wardrobe, props, environment, and voice consistent across the shot. Output tops out at 4K with 10-bit color, up to 16 seconds, and the original aspect ratio preserved.

The flagship promise is expressive audio and video as a single generation: dialogue, sound effects, and camera moves rendered together, with automatic camera switching. Audio that stays in sync and in character is the part of the stack everyone is racing to fix right now, and it is the hardest part to judge from launch demos.

The five-times claim, calculator required

"Up to five times the output for the same budget" is the line ShengShu wants quoted. Read it as what it is: a cost-per-quality argument with the quality held fixed by assumption. The comparison is against "comparable" output specs and billing conditions that the announcement doesn't name. Five times the output only matters if each of those five outputs is worth watching; price per attempt means nothing without the success rate.

The $0.014 per second is launch pricing on a preview, not a price list. The announcement says final pricing, supported resolutions, features, and usage terms may vary by plan and region — and that the point of the preview is to gather production feedback before the final Q4 model ships. Fine print worth noting if you're building a pipeline: generated creation URLs stay valid for 24 hours.

The practical angle: pricing the attempt, not the asset

Still, the framing is honest about how AI video actually gets made. Nobody buys one generation; they buy dozens of them, hunting the version where the expression lands or the camera doesn't drift. A production-ready shot routinely takes more than one attempt — ShengShu's own announcement says so. Competing on iteration cost instead of single-shot specs is the right lever, and it explains who this is aimed at: ad teams testing concepts, e-commerce teams building product variants, indie studios doing coverage they can't afford to shoot.

What to watch

Vidu was among the first video models out of China (U-ViT architecture, Tsinghua lineage) and has been climbing steadily upmarket toward advertising and e-commerce workflows. The Q4 preview is ShengShu doing something sensible: putting an unfinished model in front of real productions and asking what breaks. The question worth carrying into the final release: does the synced, expressive audio hold up on anything other than the launch examples?

Sources

  1. [1] Unite.AI — Vidu Releases Q4 Preview of Next-Generation Flagship AI Video ModelRead source
  2. [2] Vidu official product pageRead source
  3. [3] Vidu API documentationRead source