LIVE
Formalizing Fermat's Last Theorem04/09/26 · Anthropic|OpenAI and Anthropic suffer simultaneous outages with no official explanation04/09/26|Show HN: Open-Source eInk Bike Computer04/09/26|Corporate America increasingly turns to open-source AI04/09/26|Study finds Google's AI Mode surfaces pricier products than classic search04/09/26 · Google|Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Formalizing Fermat's Last Theorem04/09/26 · Anthropic|OpenAI and Anthropic suffer simultaneous outages with no official explanation04/09/26|Show HN: Open-Source eInk Bike Computer04/09/26|Corporate America increasingly turns to open-source AI04/09/26|Study finds Google's AI Mode surfaces pricier products than classic search04/09/26 · Google|Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|
🇨🇳 China🎬 Video generationReleased August 2024

CogVideoX 5B

Zhipu AI
Open-sourceSelf-hostRGPD
Verdict

A solid pick for self-hosting and full data control.

Overview

CogVideoX-5B is an open-weight video generation model from Zhipu AI, released under an Apache 2.0 licence that permits commercial use. It produces coherent short clips from detailed text prompts and, in its quantised form, runs on as little as 4.4GB of VRAM, making it practical on a single consumer GPU. The catch: prompts must be written in English and capped at 226 tokens, and generation stays slow, up to 90 seconds per clip on an H100. Because the weights are open, it can be self-hosted entirely within Europe, keeping data on-premise.

Skill profile

Not disclosed

Strengths

  • Open-source and self-hostable

Limitations

  • API pricing not disclosed

Who is it for

A good fit if…
  • you want to control cost or self-host
  • you have GDPR constraints
Skip it if…

    Ideal use cases

    • Short clip prototyping
    • On-prem video generation
    • Internal creative demos
    • Generative video R&D
    • Quick marketing content

    Access & availability

    Self-hostable (open weights)

    Key specifications

    Context
    0
    Input price
    Output price
    Speed
    Price not auditedBenchmarks not audited

    Privacy

    GDPR-compliantHosting : self

    Run it locally

    Runs on a laptop18k686
    QuantizationDiskRAM / VRAMTypical hardware
    Q4 · recommended2.8 GB5 GBAny recent PC/Mac
    Q8 · balanced5.4 GB8 GB16GB PC / M1+ Mac / 8GB GPU
    FP16 · max quality10 GB13 GB32GB PC / 24GB Mac / RTX 4070 Ti+

    Estimates for a moderate context. Long contexts need more RAM (KV cache).

    Deploy

    Copy-ready commands generated from this card. Adjust context length and GPU count to your hardware.

    Estimated memory · FP16 13 Go · Q4 5 Go

    OpenAI-compatible server for production on NVIDIA GPUs.

    pip install vllm
    vllm serve THUDM/CogVideoX-5b \
      --max-model-len 32768 \
      --tensor-parallel-size 1 \
      --dtype auto
    THUDM/CogVideoX-5b
    Advanced data · for experts
    Architecture, modalities, detailed cost, full benchmarks
    Architecture
    Size
    5B
    Cutoff date
    Inputs
    text
    Outputs
    text
    License
    apache-2.0
    Hosting
    self
    Hallucination score
    Estimated cost (API)
    Typical exchange (~3k in / 1k out)
    Not disclosed
    1M in + 1M out
    Not disclosed
    All benchmarks
    Arena Elo
    MMLU
    GPQA
    HumanEval
    SWE-Bench
    MATH

    Similar models