AI Video Model Hub

22 text-to-video models from 11 vendors, grouped by provider so you can find the best AI video model for your use case. Every model runs on the same credit balance — pick one, or run several on the same prompt and compare the clips side by side.

The two specs that decide whether a model can do your job at all are maximum clip length and native audio. Clip lengths in this catalog run from 4 seconds to 30 seconds, and 14 of 22 models generate synchronized sound. Use the filters below to narrow by vendor, audio and clip length.

Not sure where to start? Open any model profile for its credits, turnaround and strengths, or jump to the head-to-head pages to see two models judged on the same prompt.

Vendor
Audio
Clip length
Sort

22 of 22 models

Alibaba

4 models
Wan 3.0 Prime60 credits
Wan 3.0 Prime

Alibaba's Wan 3.0 Prime text-to-video model. Generates cinematic videos from text prompts with adaptive aspect ratio, 480P, 720P, or 1080P resolution, and configurable duration.

CinematicAdaptive aspect ratioUp to 1080P
Latency
~90s
Max clip
15s
Audio
No
View model
Wan 3.040 credits
Wan 3.0

Alibaba's Wan 3.0 text-to-video model. Generates cinematic videos from text prompts with adaptive aspect ratio, 480P, 720P, or 1080P resolution, and configurable duration.

CinematicAdaptive aspect ratioUp to 1080P
Latency
~70s
Max clip
15s
Audio
No
View model
HappyHorse 1.132 credits
HappyHorse 1.1

Alibaba's HappyHorse 1.1 text-to-video model. Generates videos from a text prompt with stronger dynamic expressiveness, better visual quality, and improved instruction following over 1.0. Configurable resolution, aspect ratio, and duration (3-15s).

Dynamic motionInstruction followingConfigurable
Latency
~55s
Max clip
15s
Audio
No
View model
HappyHorse 1.020 credits
HappyHorse 1.0

Alibaba's HappyHorse 1.0 text-to-video model. Generates videos from a text prompt with configurable resolution, aspect ratio, and duration (3-15s).

ConfigurableBudget friendly3-15s clips
Latency
~50s
Max clip
15s
Audio
No
View model

Lightricks

1 model
LTX-2.5 Fast18 credits
LTX-2.5 Fast

Lightricks LTX-2.5 Fast is a fast video generation model for text-to-video and image-to-video workflows, with synchronized audio, configurable duration, resolution, and frame rate.

FastSynchronized audioConfigurable fps
Latency
~25s
Max clip
20s
Audio
Yes
View model

ByteDance

4 models
Seedance 2.595 credits
Seedance 2.5

ByteDance's next-generation video model with a unified multimodal reference-to-video architecture. Generates video from text, up to 30 reference images, 10 reference videos, and 10 reference audio clips — including audio-only input with no image or video required. Supports first/last-frame image-to-video, video editing, video extension, intelligent duration (including automatic selection), and adaptive aspect ratio.

Multimodal referencesVideo editingAudio-only input
Latency
~120s
Max clip
20s
Audio
Yes
View model
Seedance 2.0 Mini14 credits
Seedance 2.0 Mini

ByteDance's compact, cost-efficient video generation model from the Seedance 2.0 family. Supports text-to-video, image-to-video, reference video, and reference audio for background music. Ideal for high-volume workloads where speed and cost matter.

Cost-efficientHigh volumeBackground music
Latency
~20s
Max clip
12s
Audio
Yes
View model
Seedance 2.0 Fast42 credits
Seedance 2.0 Fast

Faster variant of ByteDance's Seedance 2.0 video model. Trades some quality for speed while sharing the same multimodal architecture. Supports text-to-video, image-to-video, native audio generation, multimodal references (images, videos, audio), video editing, and video extension.

FastNative audioMultimodal references
Latency
~28s
Max clip
12s
Audio
Yes
View model
Seedance 2.072 credits
Seedance 2.0

ByteDance's next-generation video model with a unified multimodal architecture. Generates high-quality video with synchronized audio from text, images, video clips, and audio inputs. Supports multimodal references (up to 9 images, 3 videos, 3 audio files), native audio generation, video editing, video extension, intelligent duration, and adaptive aspect ratio.

Native audioMultimodal referencesVideo extension
Latency
~80s
Max clip
12s
Audio
Yes
View model

Black Forest Labs

1 model
FLUX 3 Video75 credits
FLUX 3 Video

FLUX 3 Video is Black Forest Labs' video generation model. It generates video from a text prompt (t2v), animates one or more reference images (i2v), or continues an existing clip (v2v), with synchronized audio, up to FHD resolution, and 5-20 second durations.

Synchronized audioUp to FHDReference images
Latency
~100s
Max clip
20s
Audio
Yes
View model

Pruna AI

1 model
P-Video70 credits
P-Video

Pruna's P-Video is a premium video generation model supporting text-to-video, image-to-video, and audio-conditioned generation up to 1080p at 24 or 48 fps, with configurable duration up to 20 seconds.

Premium24/48 fpsAudio-conditioned
Latency
~95s
Max clip
20s
Audio
Yes
View model

RunwayML

2 models
Aleph 2110 credits
Aleph 2

RunwayML's video editing model. Edit one frame to update your whole video, make changes across multiple shots, and work with up to 30 seconds of video. Supports keyframe-guided editing for precise control over specific moments in the clip.

Video editingKeyframe controlUp to 30s
Latency
~150s / 30s clip
Max clip
30s
Audio
No
View model
Gen-4.578 credits
Gen-4.5

RunwayML's video generation model supporting both text-to-video and image-to-video with customizable duration, aspect ratio, and content moderation controls.

Text & image to videoCustomizableModeration controls
Latency
~90s
Max clip
20s
Audio
No
View model

xAI

1 model
Grok Imagine Video65 credits
Grok Imagine Video

xAI's video generation model. Generates, edits, and extends videos from text and image inputs with native synchronized audio including dialogue, sound effects, and music. Supports multiple creative modes (normal, fun, custom).

Native audioDialogue & SFXCreative modes
Latency
~60s
Max clip
15s
Audio
Yes
View model

Vidu

2 models
Vidu Q3 Pro68 credits
Vidu Q3 Pro

Vidu Q3 Pro is a high-quality video generation model supporting text-to-video, image-to-video, and start/end-frame-to-video workflows with audio and up to 16-second clips.

High qualityStart/end frameUp to 16s
Latency
~85s
Max clip
16s
Audio
Yes
View model
Vidu Q3 Turbo34 credits
Vidu Q3 Turbo

Vidu Q3 Turbo is a faster version of Vidu Q3 optimized for lower latency video generation while maintaining audio support and up to 16-second clips.

Low latencyAudio supportUp to 16s
Latency
~35s
Max clip
16s
Audio
Yes
View model

PixVerse

2 models
PixVerse v645 credits
PixVerse v6

Pixverse v6 is the latest Pixverse video model with support for up to 15-second videos, customizable duration from 1 to 15 seconds, and audio generation.

Up to 15sAudio generationCustomizable duration
Latency
~50s
Max clip
15s
Audio
Yes
View model
PixVerse v5.630 credits
PixVerse v5.6

Pixverse v5.6 is a video generation model supporting text-to-video and image-to-video with audio generation, customizable aspect ratios, and up to 1080p output.

Up to 1080pAudio generationCustom aspect ratios
Latency
~45s
Max clip
10s
Audio
Yes
View model

MiniMax

2 models
Hailuo 2.3 Fast28 credits
Hailuo 2.3 Fast

A lower-latency version of Hailuo 2.3 that preserves core motion quality, visual consistency, and stylization while enabling faster iteration.

Low latencyMotion qualityFast iteration
Latency
~30s
Max clip
10s
Audio
No
View model
Hailuo 2.352 credits
Hailuo 2.3

A high-fidelity video generation model optimized for realistic human motion, cinematic VFX, expressive characters, and strong prompt and style adherence across text-to-video and image-to-video workflows.

Realistic human motionCinematic VFXStyle adherence
Latency
~65s
Max clip
10s
Audio
No
View model

Google

2 models
Veo 3.1 Fast48 credits
Veo 3.1 Fast

A faster version of Veo 3.1 optimized for lower latency while maintaining high-quality video and audio output.

Low latencyAudio outputHigh quality
Latency
~28s
Max clip
8s
Audio
Yes
View model
Veo 3.185 credits
Veo 3.1

Google's latest video generation model with improved quality, motion, and audio generation.

Improved motionAudio generationFlagship
Latency
~75s
Max clip
8s
Audio
Yes
View model