VERITAS: A Multi-agent Co-scientist for Verifiable Image-Derived Hypothesis Testing
Abstract
Scientix001Cc research based on multimodal clinical data (includ-ing medical imaging) requires coordinating clinical, radiological, pro-gramming, and biostatistical expertise, a fragmented process that bottle-necks discovery. We present Veritas (Verix001Cable Epistemic Reasoningfor Image-Derived Hypothesis Testing via Agentic Systems), a clinicalco-scientist: a multi-agent system that autonomously tests natural-language hypotheses and produces a fully auditable evidence trail, trac-ing every conclusion through executable outputs from analysis plan tosegmentation masks to statistical code to x001Cnal verdict. Unlike prior AI-scientist systems, which mainly operate on tabular or text data, Veritasgrounds autonomous discovery directly in medical images. It decomposesthe workx001Dow into four phases handled by role-specialized agents, and in-troduces an epistemic evidence label framework that mechanically classi-x001Ces outcomes as Supported, Refuted, Underpowered, or Invalid byjointly evaluating signix001Ccance, ex001Bect direction, and study power. This dis-tinction is critical in medical imaging, where non-signix001Ccant results oftenrex001Dect insux001Ecient sample size rather than absent ex001Bects. We construct atiered benchmark of 64 hypotheses spanning six complexity levels acrosscardiac and brain glioma MRI datasets. Veritas reaches 81.4% verdictaccuracy with frontier models and 71.2% with locally-hosted open-weightmodels (8x001530B), outperforming all single-model baselines in both classes.It also produces the highest rate of independently verix001Cable statisticaloutputs (86.6%), so even its failures remain diagnosable through arti-fact inspection. Structured multi-agent decomposition thus substitutesfor model scale while preserving the verix001Cability that scientix001Cc discoverydemands. We release code, hypothesis bank, and evaluation pipeline athttps://github.com/LucZot/veritas.