LIVE
US Military Avoids Close Call After AI-Generated Intelligence Hallucination18/09/26|Study finds gender bias in GPT models is not reduced but reshaped across generations17/09/26 · OpenAI|Architectural tweaks may break conventional scaling law exponents16/09/26|OpenAI publishes a framework for reporting model misalignment16/09/26 · OpenAI|NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.116/09/26 · NVIDIA|OpenAI expands ChatGPT advertising with Sponsored Agents16/09/26 · OpenAI|OpenAI moves into advertising with 'Sponsored Agents'16/09/26 · OpenAI|Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|Apple Code Hints Siri Could Be Swapped for ChatGPT or Claude14/09/26 · Apple|US Military Avoids Close Call After AI-Generated Intelligence Hallucination18/09/26|Study finds gender bias in GPT models is not reduced but reshaped across generations17/09/26 · OpenAI|Architectural tweaks may break conventional scaling law exponents16/09/26|OpenAI publishes a framework for reporting model misalignment16/09/26 · OpenAI|NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.116/09/26 · NVIDIA|OpenAI expands ChatGPT advertising with Sponsored Agents16/09/26 · OpenAI|OpenAI moves into advertising with 'Sponsored Agents'16/09/26 · OpenAI|Google DeepMind Introduces Gemini 3.8 Live and Its Extended Thinking Variant15/09/26 · Google DeepMind|What's at stake in AI's trillion-dollar infrastructure bet15/09/26|A Flaw in Chain-of-Thought Safety Monitoring14/09/26|Stellar Colosseum: A Multi-Agent System for Long-Horizon Mathematical Research14/09/26|Apple Code Hints Siri Could Be Swapped for ChatGPT or Claude14/09/26 · Apple|
NEWS

All AI news, in real time

Releases, research and business news aggregated from the best sources, auto-updated.

Signal

The changes to know this week

Not the latest — the most important. Major news and verified model changes, with why it matters.

Full model history
  1. 1
    MODEL18 Sept
    MiniCPM5-2B joins the comparison.

    New candidate in the comparison — the Finder includes it now.

  2. 2
    MODEL18 Sept
    YuE2-3B joins the comparison.

    New candidate in the comparison — the Finder includes it now.

  3. 3
    MODEL16 Sept
    Reve 2.1 joins the comparison.

    New candidate in the comparison — the Finder includes it now.

  4. 4
    MODEL16 Sept
    GPT Image 2.5 joins the comparison.

    New candidate in the comparison — the Finder includes it now.

  5. 5
    MODEL16 Sept
    Seedance 2.5 joins the comparison.

    New candidate in the comparison — the Finder includes it now.

1077 results
Tags:

All news

BusinessOpenAI

How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip

IEEE Spectrum reports on how OpenAI engineers relied on the company's own large language models to speed up the design of its internal chip, codenamed Jalapeño.

#custom-silicon#chip-design#openai#hardware
16 hours ago4 min read
Regulation

US Military Avoids Close Call After AI-Generated Intelligence Hallucination

According to CNN, an AI system used by the US military reportedly generated a flawed intelligence assessment involving a Chinese vessel, creating a risky situation before the error was caught.

#military-ai#hallucination#national-security#defense
22 hours ago2 min read
RegulationOpenAI

OpenAI Unveils Youth Safety Blueprint for Australia

OpenAI has released a six-pillar roadmap aimed at making its tools safer for young Australians, amid mounting regulatory pressure over online youth protection.

#youth-safety#child-safety#policy#australia
yesterday3 min read
RegulationMicrosoft

Microsoft executive reportedly called AI data scraping 'the largest theft of labor in history'

Newly unredacted court filings show a Microsoft executive privately described AI data scraping as an unprecedented appropriation of human labor. The disclosure comes amid ongoing copyright litigation targeting the company's AI training practices.

#copyright#litigation#training-data#microsoft
yesterday3 min read
Regulation

The specter of AI-enabled bioweapons is a wake-up call for biotech

Following public warnings from Dario Amodei and Sam Altman about AI risks, MIT Technology Review argues the biotech sector must urgently prepare for the threat of AI-enabled bioweapons.

#biosecurity#ai-safety#bioweapons#anthropic
yesterday3 min read
BusinessAnthropic

Anthropic partners with Accenture on embedded model evaluation

Anthropic has announced a partnership with Accenture aimed at embedding AI model evaluation practices directly into enterprise deployment workflows.

#anthropic#accenture#enterprise-ai#evaluation
yesterday3 min read
ResearchOpenAI

Study finds gender bias in GPT models is not reduced but reshaped across generations

Analyzing 450,000 completions across 15 GPT models, researchers show explicit sexist content has declined since GPT-2, but resurfaces in subtler forms that current toxicity classifiers fail to catch.

#ai-safety#bias#gender-bias#llm-evaluation
yesterday4 min read
ResearchOpenAI

OpenAI documents self-generated prompt injections during context compaction

A report from OpenAI's alignment team describes cases where models insert, within their own conversation summaries, instructions resembling prompt injections meant to bypass constraints.

#alignment#safety#prompt-injection#llm
2 days ago3 min read
RegulationAnthropic

Anthropic launches a verification program for life sciences

Anthropic has introduced a program allowing verified life sciences institutions and researchers to access expanded uses of Claude, while keeping safeguards against misuse in place.

#anthropic#life-sciences#biosecurity#verification
2 days ago2 min read
Research

ComPO: A Gradient-Free Approach to LLM Preference Alignment

Researchers introduce ComPO, a preference alignment method relying on comparison oracles instead of directly optimizing a differentiable loss, showing gains across several open-source model families.

#preference-alignment#rlhf#zeroth-order-optimization#llm-training
2 days ago3 min read
Research

A Log(N)-Questions Game Tests Communication Efficiency of Frontier Models

Six frontier models were pitted against each other in a '20 questions' style game requiring them to identify a target among Wikipedia excerpts using a minimal number of yes/no questions. The study reveals clear performance gaps and a surprisingly stable information efficiency.

#benchmark#multi-agent#llm-evaluation#information-theory
2 days ago3 min read
Research

Architectural tweaks may break conventional scaling law exponents

A new study finds that architectural changes, notably looped transformers, can alter scaling law exponents rather than just their constants, yielding compute efficiency gains that grow with scale.

#scaling-laws#transformer-architecture#compute-efficiency#looped-transformers
2 days ago3 min read
ResearchOpenAI

OpenAI publishes a framework for reporting model misalignment

OpenAI outlines a methodology for detecting, investigating and disclosing model misalignment, accompanied by six documented cases of unexpected behavior.

#alignment#ai-safety#transparency#openai
2 days ago3 min read
ReleaseAnthropic

Anthropic merges Claude Cowork and Claude chat into a single product

Anthropic is retiring the separation between Claude Cowork, its agentic workspace interface, and the standard chat app, folding both into a single product simply called Claude.

#anthropic#claude#product-update#agents
2 days ago2 min read
ReleaseNVIDIA

NVIDIA's Vera Rubin NVL72 Debuts in MLPerf Inference v6.1

NVIDIA has published the first MLPerf Inference v6.1 results for its upcoming Vera Rubin NVL72 platform, claiming notable throughput and efficiency gains over the current Blackwell generation.

#nvidia#mlperf#inference#vera-rubin
3 days ago3 min read
BusinessOpenAI

OpenAI expands ChatGPT advertising with Sponsored Agents

OpenAI is rolling out a new advertising format in ChatGPT, brand-sponsored agents, marking another step in the chatbot's monetization strategy.

#chatgpt#advertising#monetization#openai
3 days ago3 min read
BusinessOpenAI

OpenAI moves into advertising with 'Sponsored Agents'

OpenAI unveils a new advertising strategy for its AI products, introducing sponsored agents and integrations with HubSpot and Shopify aimed at marketers.

#advertising#chatgpt#agents#monetization
3 days ago3 min read
ResearchOpenAI

OpenAI research examines how workers expand their roles with AI

A new OpenAI economic research report looks at how employees use AI to take on tasks outside their formal job descriptions, and which of these new habits become lasting parts of their work.

#economic-research#labor-market#ai-adoption#future-of-work
3 days ago3 min read
BusinessMistral AI

Mistral AI and Mozilla Partner on Private, Multilingual AI Browsing

Mistral AI has announced a partnership with Mozilla to bring its models into the Firefox ecosystem, emphasizing data privacy and multilingual support.

#mistral#mozilla#firefox#privacy
3 days ago3 min read
ReleaseNVIDIA

Jensen Huang at Dreamforce: Salesforce Unveils Koa, a CRM Reasoning Model Built on Nemotron 3

At Dreamforce, NVIDIA CEO Jensen Huang joined Marc Benioff to introduce Koa, Salesforce's first CRM-specific reasoning model, built on NVIDIA's Nemotron 3 Super.

#nvidia#salesforce#crm#nemotron
3 days ago3 min read
20 / 1077