LIVE
Formalizing Fermat's Last Theorem04/09/26 · Anthropic|OpenAI and Anthropic suffer simultaneous outages with no official explanation04/09/26|Show HN: Open-Source eInk Bike Computer04/09/26|Corporate America increasingly turns to open-source AI04/09/26|Study finds Google's AI Mode surfaces pricier products than classic search04/09/26 · Google|Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Formalizing Fermat's Last Theorem04/09/26 · Anthropic|OpenAI and Anthropic suffer simultaneous outages with no official explanation04/09/26|Show HN: Open-Source eInk Bike Computer04/09/26|Corporate America increasingly turns to open-source AI04/09/26|Study finds Google's AI Mode surfaces pricier products than classic search04/09/26 · Google|Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|
🇫🇷 France🎧 Speech recognitionReleased June 2025

Voxtral Small 24B

Mistral AI
Open-sourceSelf-hostRGPD
Verdict

A solid pick for self-hosting and full data control.

Overview

Voxtral Small builds on the Mistral Small 3 text backbone and adds a full audio stack: transcription, translation and spoken-language understanding, plus function calling straight from voice. It handles up to 30 minutes of audio for transcription and 40 minutes for understanding, across eight languages including French and German. One real limit: system prompts aren't supported yet, and running it in bf16 needs roughly 55GB of GPU RAM. Released under Apache 2.0, it runs entirely on your own infrastructure, so audio data never leaves your servers.

Skill profile

Not disclosed

Strengths

  • Open-source and self-hostable

Limitations

  • API pricing not disclosed

Who is it for

A good fit if…
  • you want to control cost or self-host
  • you have GDPR constraints
Skip it if…

    Ideal use cases

    • Multilingual meeting transcription
    • Call summarization pipelines
    • Voice-driven business assistants
    • Voice-triggered workflow automation

    Access & availability

    Self-hostable (open weights)

    Key specifications

    Context
    32K
    Input price
    Output price
    Speed
    Price not auditedBenchmarks not audited

    Privacy

    GDPR-compliantHosting : self

    Run it locally

    Desktop with solid GPU / 32GB+ Mac166k523
    QuantizationDiskRAM / VRAMTypical hardware
    Q4 · recommended13.7 GB17 GB32GB PC / 24GB Mac / RTX 4070 Ti+
    Q8 · balanced25.7 GB30 GB64GB Mac / dual 24GB GPUs
    FP16 · max quality48 GB55 GB96-128GB Mac Studio / 4× GPUs

    Estimates for a moderate context. Long contexts need more RAM (KV cache).

    Deploy

    Copy-ready commands generated from this card. Adjust context length and GPU count to your hardware.

    Estimated memory · FP16 55 Go · Q4 17 Go

    OpenAI-compatible server for production on NVIDIA GPUs.

    pip install vllm
    vllm serve mistralai/Voxtral-Small-24B-2507 \
      --max-model-len 32000 \
      --tensor-parallel-size 1 \
      --dtype auto
    mistralai/Voxtral-Small-24B-2507
    Advanced data · for experts
    Architecture, modalities, detailed cost, full benchmarks
    Architecture
    dense
    Size
    24B
    Cutoff date
    Inputs
    text
    Outputs
    text
    License
    apache-2.0
    Hosting
    self
    Hallucination score
    Estimated cost (API)
    Typical exchange (~3k in / 1k out)
    Not disclosed
    1M in + 1M out
    Not disclosed
    All benchmarks
    Arena Elo
    MMLU
    GPQA
    HumanEval
    SWE-Bench
    MATH

    Similar models