Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
Abstract
Automatically detecting abnormal events in videos is cru-cial for modern autonomous systems, yet existing Video Anomaly De-tection (VAD) benchmarks lack the scene diversity, balanced anomalycoverage, and temporal complexity needed to reliably assess real-worldperformance. Meanwhile, the community is increasingly moving towardVideo Anomaly Understanding (VAU), which requires deeper seman-tic and causal reasoning but remains difficult to benchmark due to theheavy manual annotation effort it demands. In this paper, we introducePistachio, a new VAD/VAU benchmark constructed entirely through acontrolled, generation-based pipeline. By leveraging recent advances invideo generation models, Pistachio provides precise control over scenes,anomaly types, and temporal narratives, effectively eliminating the bi-ases and limitations of Internet-collected datasets. Our pipeline inte-grates scene-conditioned anomaly assignment, multi-step storyline gen-eration, and a temporally consistent long-form synthesis strategy thatproduces coherent 41-second videos with minimal human intervention.Extensive experiments demonstrate the scale, diversity, and complexityof Pistachio, revealing new challenges for existing methods and motivat-ing future research on dynamic and multi-event anomaly understanding.