

Foundation Video Generation Models
Foundation models that generate video from text, image or reference inputs — assessed on raw fidelity, motion physics, prompt adherence, duration, native audio and API availability.
Top 15 solutions
Click any solution to instantly generate a full AI-powered research report.
Top-tier fidelity with native synchronised audio/dialogue, strong physics and camera control; available via Gemini, Flow and Vertex AI.
Highly realistic physics and world-simulation with native audio, cameo/likeness features and a consumer social app.
Leading motion quality and human movement, long-duration clips, motion brush and strong image-to-video; aggressive release cadence.
Multi-shot narrative generation with strong prompt adherence and speed; distributed via Dreamina/CapCut and BytePlus API.
Consistent characters and scenes from references, plus in-video editing (Aleph); tuned for professional control.
Reasoning-driven video model with native HDR output and draft mode for fast iteration.
Excellent physics and instruction following at low cost; popular via API aggregators.
Open-weights video model family enabling self-hosting, fine-tuning and ComfyUI pipelines.
Strong reference-to-video consistency across multiple subjects; anime and stylised strength.
Fast, effects-driven generation optimised for social content creators.
Open, fast model with native audio and 4K output designed for real-time iteration.
Playful effects (Pikaffects, Pikaswaps) aimed at consumer creators.
Open-source high-parameter video model with strong community adoption.
Commercially-safe model trained on licensed footage for studio/filmmaker use.
Open-weights research model with good motion; limited commercial tooling.
