LIVE
Formalizing Fermat's Last Theorem04/09/26 · Anthropic|OpenAI and Anthropic suffer simultaneous outages with no official explanation04/09/26|Show HN: Open-Source eInk Bike Computer04/09/26|Corporate America increasingly turns to open-source AI04/09/26|Study finds Google's AI Mode surfaces pricier products than classic search04/09/26 · Google|Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|Formalizing Fermat's Last Theorem04/09/26 · Anthropic|OpenAI and Anthropic suffer simultaneous outages with no official explanation04/09/26|Show HN: Open-Source eInk Bike Computer04/09/26|Corporate America increasingly turns to open-source AI04/09/26|Study finds Google's AI Mode surfaces pricier products than classic search04/09/26 · Google|Study measures how coding agents pick third-party tools03/09/26|OpenAI's GPT-6 Astra sets new state-of-the-art on ARC-AGI-303/09/26 · OpenAI|ESPO: A Prompt Optimization Method That Beats GEPA With Shorter, More Stable Prompts03/09/26|"Last Translation Benchmark": a large-scale collaborative effort for a definitive MT benchmark03/09/26|NVIDIA pushes local AI at IFA 2026 with RTX Spark PCs and a personal inference router03/09/26 · NVIDIA|Google DeepMind unveils WeatherNext 3, its most accurate weather AI model yet03/09/26 · Google DeepMind|OpenAI Launches $1B Daybreak Program for Frontline Cybersecurity Defenders03/09/26 · OpenAI|
🇫🇷 France🎧 Speech recognitionReleased June 2025

Voxtral Mini 3B

Mistral AI
Open-sourceSelf-hostRGPD
Verdict

A solid pick for self-hosting and full data control.

Overview

Voxtral Mini 3B is Mistral AI's audio-enabled extension of Ministral 3B, built to transcribe, translate and summarise speech across eight languages within a 32k-token window that covers up to 40 minutes of audio per pass. Its real limitation: system prompts aren't supported yet, which complicates integration into agentic pipelines. Released under Apache 2.0 and light enough to run on a laptop, it keeps all data on-premise — a genuine plus for European teams prioritising data sovereignty.

Skill profile

Not disclosed

Strengths

  • Open-source and self-hostable

Limitations

  • API pricing not disclosed

Who is it for

A good fit if…
  • you want to control cost or self-host
  • you have GDPR constraints
Skip it if…

    Ideal use cases

    • Confidential meeting transcription
    • Multilingual voice translation
    • Automated call summarization
    • Voice-triggered API calls
    • On-prem subtitle generation

    Access & availability

    Self-hostable (open weights)

    Key specifications

    Context
    32K
    Input price
    Output price
    Speed
    Price not auditedBenchmarks not audited

    Privacy

    GDPR-compliantHosting : self

    Run it locally

    Runs on a laptop518k675
    QuantizationDiskRAM / VRAMTypical hardware
    Q4 · recommended1.7 GB4 GBAny recent PC/Mac
    Q8 · balanced3.2 GB6 GBAny recent PC/Mac
    FP16 · max quality6 GB9 GB16GB PC / M1+ Mac / 8GB GPU

    Estimates for a moderate context. Long contexts need more RAM (KV cache).

    Deploy

    Copy-ready commands generated from this card. Adjust context length and GPU count to your hardware.

    Estimated memory · FP16 9 Go · Q4 4 Go

    OpenAI-compatible server for production on NVIDIA GPUs.

    pip install vllm
    vllm serve mistralai/Voxtral-Mini-3B-2507 \
      --max-model-len 32000 \
      --tensor-parallel-size 1 \
      --dtype auto
    mistralai/Voxtral-Mini-3B-2507
    Advanced data · for experts
    Architecture, modalities, detailed cost, full benchmarks
    Architecture
    dense
    Size
    3B
    Cutoff date
    Inputs
    text
    Outputs
    text
    License
    apache-2.0
    Hosting
    self
    Hallucination score
    Estimated cost (API)
    Typical exchange (~3k in / 1k out)
    Not disclosed
    1M in + 1M out
    Not disclosed
    All benchmarks
    Arena Elo
    MMLU
    GPQA
    HumanEval
    SWE-Bench
    MATH

    Similar models