Zhipu AI · Video generation
CogVideoX-5B
Zhipu AI's open-source 5B parameter video generation model, integrating 3D Variational Autoencoder (3D VAE) compression. Supports local deployment (24GB VRAM recommended), with per-clip cloud APIs available on fal.ai.
Developer
Zhipu AI
Type
Video generation
Status
Active
Last confirmed
Sep 25, 2026
Next review Oct 25, 2026
Official specs
| Clip length | 6 seconds |
|---|---|
| What it can do | Text to video, Image to video |
| Highest resolution | 720p |
| Audio | Video only, no audio |
5B parameter model with expert Transformer architecture and 3D VAE, generating 6-second clips.
From Zhipu AI's own page, confirmed on Sep 25, 2026. This is what the model can do; whether a given channel sells it, and at which settings, is in the prices below. Official spec page
Where it is available
None of the subscription tools this site tracks list it; for now it is sold only through the APIs below.
Pay-as-you-go API prices
| API provider | API model name | Resolution | Audio | Official price | Per second | Minimum per generation |
|---|---|---|---|---|---|---|
| fal | fal-ai/cogvideox-5b | 720p | No audio | $0.20 per clip (6 s) | $0.033 | — |
Prices as each provider's official pricing page states them, confirmed Sep 18, 2026. No subscription: you pay for the seconds generated. Prices exclude tax.
How the status is checked
Whether a model still runs is checked separately from prices: prices change monthly, but a model can disappear between two price syncs, quietly making every page that mentions it wrong. So each model record has its own confirmation date and review deadline; overdue records are flagged, never assumed to still be running.
Last confirmed on Sep 25, 2026; due to be rechecked by Oct 25, 2026.
Recommended Alternatives & Budget-Friendly Picks
Based on the same model category (Video generation), public specs, and availability across channels, here are recommended alternatives:
A third-party model listed on both Luma's and Runway's official pricing pages. Luma publishes a per-second rate; Runway lists it in its plans, but part of Runway's rate table sits behind "View more", and the page we read has no rate for this model. Kling's own site lists it as the “Kling 3.0 Model Series”, made up of two models, VIDEO 3.0 and VIDEO 3.0 Omni. Credits per second on Kling's international web membership (kling.ai) were read from the generation page's official quote after signing in on Sep 19, 2026, and are linear in length.
A video model from MiniMax (the company behind Hailuo), listed in Luma's official rate table. Audio is included by default, up to 1440p (2K). It is the default model in the generator on Hailuo's own homepage, and PixVerse's site has a page for it too.
Tencent's open-source 13B parameter video generation model, natively supporting Chinese and English prompts with 720p generation. Supports ComfyUI local deployment (full version requires 60GB+ VRAM, quantized requires 24GB), with pay-per-second cloud APIs available on fal.ai.