EN DIRECT
The Download: the Kids issue arrives, and Bill Gates reveals his AI fears26/08/26|Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights26/08/26|How loveholidays is making everyone a builder with Codex26/08/26 · OpenAI|Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers26/08/26 · Hugging Face|Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses25/08/26 · OpenAI|Granite 4.2 LLMs: How They're Built25/08/26 · Hugging Face|OpenAI Jalapeño: Better than Nvidia Blackwell25/08/26 · OpenAI|Anthropic tells staff to work from home due to possible security team strike25/08/26 · Anthropic|OpenAI restores 5-hour Codex and Work limits for ChatGPT Plus users25/08/26 · OpenAI|Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original25/08/26 · Hugging Face|The full stack behind abundant intelligence25/08/26 · OpenAI|Jalapeño’s first results show industry-leading speed and efficiency in AI inference25/08/26 · OpenAI|The Download: the Kids issue arrives, and Bill Gates reveals his AI fears26/08/26|Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights26/08/26|How loveholidays is making everyone a builder with Codex26/08/26 · OpenAI|Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers26/08/26 · Hugging Face|Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses25/08/26 · OpenAI|Granite 4.2 LLMs: How They're Built25/08/26 · Hugging Face|OpenAI Jalapeño: Better than Nvidia Blackwell25/08/26 · OpenAI|Anthropic tells staff to work from home due to possible security team strike25/08/26 · Anthropic|OpenAI restores 5-hour Codex and Work limits for ChatGPT Plus users25/08/26 · OpenAI|Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original25/08/26 · Hugging Face|The full stack behind abundant intelligence25/08/26 · OpenAI|Jalapeño’s first results show industry-leading speed and efficiency in AI inference25/08/26 · OpenAI|
ReleaseOpenAI

Jalapeño’s first results show industry-leading speed and efficiency in AI inference

Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.

25 août 20261 min de lecturePublié parOpenAI

Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.

Tags
inference

À lire aussi