We shipped video this week. Seven model families, one credits pool, no per-provider contracts to manage. Before we get into what’s in the lineup, a quick note on how we picked — because the “which video models are best” answer is much less obvious than the chat-side equivalent.
The four axes that matter
Chat models compete mostly on intelligence. Video models compete on four axes at once, and almost no model wins all four:
- Visual quality — sharpness, motion coherence, how often a hand has five fingers.
- Speed and cost — how long until you have something to look at, and what it costs to iterate.
- Control — references, seeds, durations, camera moves, lip-sync.
- Audio — whether the model emits sound that matches the visuals, or whether you have to bolt foley on after.
A 4K cinematic generator is wrong for an X reply. A fast social-tier model is wrong for a hero shot. So the lineup is plural by design — we want one right tool per job, not one model trying to be everything.
What’s in
HappyHorse 1.0 is the multi-reference option. It accepts text, a single image, or a set of up to nine image references, with seed and resolution controls for repeatable visual exploration.
Veo 3.1 is the cinematic option. 4K-native, the cleanest text-to-shot we’ve seen, and ridiculously expensive — so we route to it when the prompt looks like a hero shot (“wide-angle, dolly in, golden hour”) and not when it looks like a social clip.
Seedance 2.0 supports up to nine image references and generated audio. It’s a different kind of tool: less for “type a prompt, get a video” and more for “give me a director’s brief.” We’ll cover it separately on May 3.
Kling 3.0 Omni is the motion-fluidity specialist. Best multi-shot continuity in our tests, especially when you need a character to walk through several beats without their face morphing between them.
Grok Imagine is the social-native option — fast, drafts-quality, real- feeling clips. We added it last week and have a fuller write-up coming.
Wan 2.7 is the reference-and-keyframe option, with start/end framing, multi-image references, seeds, and up to 1080p output.
Hailuo 02 is the straightforward text-to-video option for 6- or 10-second generations.
What didn’t make it
Three serious models almost made the launch and didn’t, for different reasons.
The first didn’t ship a non-watermarked tier yet — we won’t surface a model that brands your output. The second has great visuals but no API for seed control, which makes iteration painful in a multi-take workflow. The third is just expensive in a way that doesn’t pencil out — even for Studio-tier users, the per-clip cost would consume too much campaign capacity.
We’ll revisit all three when their tiering changes.
One pool, eight models
The reason any of this works is that you’re not buying eight separate subscriptions. You spend the same shared credits whichever model you pick, and you can see the cost per generation before you commit. Creator is 3,000 credits a month, Power is 6,000, and Studio is 14,000. The exact estimate changes with the model, duration, resolution, and current provider pricing, so the studio calculates it from your settings before every generation.
That ratio is the part we’re proudest of. The point of the studio is that you don’t have to pre-commit to a provider before you know what your prompt needs — and video, more than any other modality, punishes that kind of pre-commitment.
Video is available anywhere credits are available. The one-time Free Trial is for trying the workflow; Creator, Power, and Studio provide the recurring capacity for regular campaign production.