Vidu's Q4 Preview is built around references: up to fifteen images and three audio clips. The useful way to read that capacity is not as a bigger upload allowance, but as a production kit. Give each slot a job, direct the model in shots, and spend resolution only after the scene works.

Vidu Q4 Preview title treatment on Vidu's official site
Vidu Q4 Preview, via Vidu's official site.

The kit: 15 slots, zero excuses

Vidu's Reference to Video takes three or more images and lets you feed different angles of a character or object to hold consistency, per the official documentation. Q4 Preview raises the ceiling to fifteen images plus three audio clips. That is not "more inputs." It is the difference between a mood board and a storyboard.

Spend your slots like a production designer. Three on the character: front, three-quarter, profile. Two on wardrobe, worn and shown flat if the outfit changes mid-scene. Two or three on key props. The prop shots should sit on plain backgrounds, not embedded in scenes, so the model learns the object, not the room. Three or four on environment plates: the wide, the mid, the detail that sells the location. Keep one or two in reserve for the frame you wish you had. When the fifteenth slot goes to a blurry screenshot, that is the frame your video will copy.

One discipline matters more than count: consistency across the kit. Same lighting mood, same style register, same character design. Fifteen references that disagree with each other do not average out. They drift. Save the kit in Vidu's My References feature so you can reuse it; rebuilding it each time is how you get a different protagonist every Monday.

Three audio clips are a direction, not three files

The three audio references guide voice consistency and emotional delivery, according to the launch release. The mistake is uploading three takes of the same line in the same tone. Give the model a range: the neutral read, the emotional peak, the quiet moment. Clean, dry voice, no music bed. If the line in the video is a whisper over machinery, one of the three references should demonstrate exactly that. The model cannot extrapolate a delivery style it has not heard.

Write it like a call sheet

Vidu's own example prompt for Q4 is instructive, and it sits on the public homepage. It opens with a subject lock: which reference image owns the scene, and what must stay consistent (face, helmet, jersey, football, smoke). Then duration and aspect ratio. Then shot beats with hard cuts: 0–3s close-up, 3–7s tracking shot, 7–10s frontal medium. Then explicit instructions for text on wardrobe and what not to add (no extra players, no subtitles). Sound design gets one line, and it bans background music.

Steal that structure, not the content. Subject lock, then blocking, then cuts, then constraints, then sound. The constraints section is doing more work than most creators realize: every "do not add" is one fewer thing to reroll.

Vidu Q4 example prompt showing subject lock and timed shot directions
Vidu's own Q4 example prompt reads like a call sheet: subject lock, shot beats, hard cuts. Image: Vidu official site.

Draft cheap, pay for the final

The launch price is $0.014 per generated second. A ten-second clip costs fourteen cents. That is the price of one take, and Vidu's own pitch is that a production-ready shot takes more than one attempt.

So do not test at 4K. The resolution ladder runs 540p, 720p, 1080p, 2K, 4K. Block the scene at 540p: does the camera move make sense, do the cuts land. Check performance at 720p: faces, lip-sync, dialogue timing. Pay for 2K or 4K only for the take that survives. Shengshu says Q4 Preview delivers "up to five times as much output for the same budget under comparable output specifications and billing conditions." Read that as vendor arithmetic against their own prior pricing, not a universal law. And launch pricing is promotional, which means it has an expiration date even if nobody prints one.

Where it breaks

Two things the release will not say plainly. First, Q4 Preview is a preview, released explicitly so creators can "help shape the final release." Behaviors you depend on can change underneath you. Second, reference limits do not fix reference quality. Contradictory kits produce characters with wandering faces and voices that change rooms mid-sentence, and no amount of prompting fully repairs that.

*How this was assembled:* this workflow is built from Vidu's official Q4 launch release, the Reference to Video documentation on vidu.com, and third-party beta notes. We have not run paid generations on Q4 Preview ourselves; prices are public launch rates plus arithmetic, and the verdicts are ours.

Sources

  1. [1] Shengshu Technology, “Vidu Launches Q4 Preview, Making Flagship AI Video Creation Accessible to Everyone” (PR Newswire, Oct 7, 2026)Read source
  2. [2] Vidu official site, Reference to Video FAQRead source
  3. [3] SelfyzAI first-batch beta testing of Vidu Q4 (EIN Presswire, Oct 9, 2026)Read source