AI Video Model Hub
22 text-to-video models from 11 vendors, grouped by provider so you can find the best AI video model for your use case. Every model runs on the same credit balance — pick one, or run several on the same prompt and compare the clips side by side.
The two specs that decide whether a model can do your job at all are maximum clip length and native audio. Clip lengths in this catalog run from 4 seconds to 30 seconds, and 14 of 22 models generate synchronized sound. Use the filters below to narrow by vendor, audio and clip length.
Not sure where to start? Open any model profile for its credits, turnaround and strengths, or jump to the head-to-head pages to see two models judged on the same prompt.
22 of 22 models
Alibaba
4 modelsWan 3.0 Prime
Alibaba's Wan 3.0 Prime text-to-video model. Generates cinematic videos from text prompts with adaptive aspect ratio, 480P, 720P, or 1080P resolution, and configurable duration.
- Latency
- ~90s
- Max clip
- 15s
- Audio
- No
Wan 3.0
Alibaba's Wan 3.0 text-to-video model. Generates cinematic videos from text prompts with adaptive aspect ratio, 480P, 720P, or 1080P resolution, and configurable duration.
- Latency
- ~70s
- Max clip
- 15s
- Audio
- No
HappyHorse 1.1
Alibaba's HappyHorse 1.1 text-to-video model. Generates videos from a text prompt with stronger dynamic expressiveness, better visual quality, and improved instruction following over 1.0. Configurable resolution, aspect ratio, and duration (3-15s).
- Latency
- ~55s
- Max clip
- 15s
- Audio
- No
HappyHorse 1.0
Alibaba's HappyHorse 1.0 text-to-video model. Generates videos from a text prompt with configurable resolution, aspect ratio, and duration (3-15s).
- Latency
- ~50s
- Max clip
- 15s
- Audio
- No
Lightricks
1 modelLTX-2.5 Fast
Lightricks LTX-2.5 Fast is a fast video generation model for text-to-video and image-to-video workflows, with synchronized audio, configurable duration, resolution, and frame rate.
- Latency
- ~25s
- Max clip
- 20s
- Audio
- Yes
ByteDance
4 modelsSeedance 2.5
ByteDance's next-generation video model with a unified multimodal reference-to-video architecture. Generates video from text, up to 30 reference images, 10 reference videos, and 10 reference audio clips — including audio-only input with no image or video required. Supports first/last-frame image-to-video, video editing, video extension, intelligent duration (including automatic selection), and adaptive aspect ratio.
- Latency
- ~120s
- Max clip
- 20s
- Audio
- Yes
Seedance 2.0 Mini
ByteDance's compact, cost-efficient video generation model from the Seedance 2.0 family. Supports text-to-video, image-to-video, reference video, and reference audio for background music. Ideal for high-volume workloads where speed and cost matter.
- Latency
- ~20s
- Max clip
- 12s
- Audio
- Yes
Seedance 2.0 Fast
Faster variant of ByteDance's Seedance 2.0 video model. Trades some quality for speed while sharing the same multimodal architecture. Supports text-to-video, image-to-video, native audio generation, multimodal references (images, videos, audio), video editing, and video extension.
- Latency
- ~28s
- Max clip
- 12s
- Audio
- Yes
Seedance 2.0
ByteDance's next-generation video model with a unified multimodal architecture. Generates high-quality video with synchronized audio from text, images, video clips, and audio inputs. Supports multimodal references (up to 9 images, 3 videos, 3 audio files), native audio generation, video editing, video extension, intelligent duration, and adaptive aspect ratio.
- Latency
- ~80s
- Max clip
- 12s
- Audio
- Yes
Black Forest Labs
1 modelFLUX 3 Video
FLUX 3 Video is Black Forest Labs' video generation model. It generates video from a text prompt (t2v), animates one or more reference images (i2v), or continues an existing clip (v2v), with synchronized audio, up to FHD resolution, and 5-20 second durations.
- Latency
- ~100s
- Max clip
- 20s
- Audio
- Yes
Pruna AI
1 modelP-Video
Pruna's P-Video is a premium video generation model supporting text-to-video, image-to-video, and audio-conditioned generation up to 1080p at 24 or 48 fps, with configurable duration up to 20 seconds.
- Latency
- ~95s
- Max clip
- 20s
- Audio
- Yes
RunwayML
2 modelsAleph 2
RunwayML's video editing model. Edit one frame to update your whole video, make changes across multiple shots, and work with up to 30 seconds of video. Supports keyframe-guided editing for precise control over specific moments in the clip.
- Latency
- ~150s / 30s clip
- Max clip
- 30s
- Audio
- No
Gen-4.5
RunwayML's video generation model supporting both text-to-video and image-to-video with customizable duration, aspect ratio, and content moderation controls.
- Latency
- ~90s
- Max clip
- 20s
- Audio
- No
xAI
1 modelGrok Imagine Video
xAI's video generation model. Generates, edits, and extends videos from text and image inputs with native synchronized audio including dialogue, sound effects, and music. Supports multiple creative modes (normal, fun, custom).
- Latency
- ~60s
- Max clip
- 15s
- Audio
- Yes
Vidu
2 modelsVidu Q3 Pro
Vidu Q3 Pro is a high-quality video generation model supporting text-to-video, image-to-video, and start/end-frame-to-video workflows with audio and up to 16-second clips.
- Latency
- ~85s
- Max clip
- 16s
- Audio
- Yes
Vidu Q3 Turbo
Vidu Q3 Turbo is a faster version of Vidu Q3 optimized for lower latency video generation while maintaining audio support and up to 16-second clips.
- Latency
- ~35s
- Max clip
- 16s
- Audio
- Yes
PixVerse
2 modelsPixVerse v6
Pixverse v6 is the latest Pixverse video model with support for up to 15-second videos, customizable duration from 1 to 15 seconds, and audio generation.
- Latency
- ~50s
- Max clip
- 15s
- Audio
- Yes
PixVerse v5.6
Pixverse v5.6 is a video generation model supporting text-to-video and image-to-video with audio generation, customizable aspect ratios, and up to 1080p output.
- Latency
- ~45s
- Max clip
- 10s
- Audio
- Yes
MiniMax
2 modelsHailuo 2.3 Fast
A lower-latency version of Hailuo 2.3 that preserves core motion quality, visual consistency, and stylization while enabling faster iteration.
- Latency
- ~30s
- Max clip
- 10s
- Audio
- No
Hailuo 2.3
A high-fidelity video generation model optimized for realistic human motion, cinematic VFX, expressive characters, and strong prompt and style adherence across text-to-video and image-to-video workflows.
- Latency
- ~65s
- Max clip
- 10s
- Audio
- No
Veo 3.1 Fast
A faster version of Veo 3.1 optimized for lower latency while maintaining high-quality video and audio output.
- Latency
- ~28s
- Max clip
- 8s
- Audio
- Yes
Veo 3.1
Google's latest video generation model with improved quality, motion, and audio generation.
- Latency
- ~75s
- Max clip
- 8s
- Audio
- Yes