LIVE
OpenAI called the Hugging Face attack unprecedented. But we’ve been here before.27/07/26 · OpenAI|ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding27/07/26 · OpenAI|Efficient LLM-Generated Shuttling Compilers for Complex Trapped-Ion Architectures27/07/26 · Anthropic|Show HN: FeyNoBg – Automatic background removal model and training library27/07/26 · Google|Elevated errors on Claude Opus 527/07/26 · Anthropic|NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics27/07/26 · Hugging Face|Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security27/07/26 · NVIDIA|How AI is expanding what people do at work27/07/26 · OpenAI|NVIDIA Harnesses Vera CPU to Speed Up Design of Next-Generation CPUs and GPUs27/07/26 · NVIDIA|Our position on open-weights models27/07/26 · Anthropic|Reimagining Independence: How Meta’s AI Models Are Helping the University of Pittsburgh Transform Assistive Robotics27/07/26 · Meta|Cognizant and Anthropic expand their partnership to bring Claude to enterprise clients27/07/26 · Anthropic|OpenAI called the Hugging Face attack unprecedented. But we’ve been here before.27/07/26 · OpenAI|ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding27/07/26 · OpenAI|Efficient LLM-Generated Shuttling Compilers for Complex Trapped-Ion Architectures27/07/26 · Anthropic|Show HN: FeyNoBg – Automatic background removal model and training library27/07/26 · Google|Elevated errors on Claude Opus 527/07/26 · Anthropic|NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics27/07/26 · Hugging Face|Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security27/07/26 · NVIDIA|How AI is expanding what people do at work27/07/26 · OpenAI|NVIDIA Harnesses Vera CPU to Speed Up Design of Next-Generation CPUs and GPUs27/07/26 · NVIDIA|Our position on open-weights models27/07/26 · Anthropic|Reimagining Independence: How Meta’s AI Models Are Helping the University of Pittsburgh Transform Assistive Robotics27/07/26 · Meta|Cognizant and Anthropic expand their partnership to bring Claude to enterprise clients27/07/26 · Anthropic|
ResearchOpenAI

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D…

July 27, 20261 min readPublished byarXiv

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (\textit{e.g.}, Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.

Tags
llmmultimodalragopen-sourcebenchmark

Read also

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding · nAIvigate