A different kind of model
MiniMax 3 makes up to 15 seconds of 2K video, with stereo sound generated in the same pass as the picture. The shift from earlier models is not just resolution. It accepts character references, movement references, sound, and direction all at once, and blends them into the same clip.

One input, one job
The fastest way to a usable result is not to use every input. It is to assign each part a single, clear role. The reference image is the exact opening frame and the character's identity. Direction covers the tone and emotional feel of the shot, nothing more. Subject action is one clear action, stated in the order it happens. Camera is one dominant movement a real camera could physically make. Environment sets the location, lighting, and whatever moves in the background. Audio is dialogue, ambience, or music only where the shot actually needs it. Consistency locks the face, wardrobe, product, or lighting that must not change. And the final frame states where the subject and camera end.

Start with one idea
Before touching any reference, decide what the shot shows. Not the whole story. Just the shot. "A creator walks into a bright studio, sits at a desk, and opens a laptop." One subject, one action, one payoff. That is the kind of shot most people actually need: product demos, lifestyle content, social clips. Packing three camera moves, a wardrobe change, and an explosion into fifteen seconds is how good ideas become expensive mistakes. Generations consume credits whether they succeed or not, so the simple version is also the cheap version.

Add complexity last
Once the simple shot works, add one thing. A second movement, an extra sound cue, a harder lighting setup. Removing elements after the model has turned the scene into something unusable is harder than building it up.
Source note
Inspired by a walkthrough on Promethean AI, rewritten in our own editorial voice.
Sources
- [1] promethean-ai.comRead source