From Script to Shot: A Benchmark for Grounding Screenplays in Movies
Abstract
Aligning screenplay scenes to video shots is a foundationaltask for narrative video understanding. Unlike subtitles or captions,screenplays combine dialogue, visual direction, and narrative cues in asingle document, making them ill-suited for standard video-text retrievaland leaving prior methods to exploit only subtitle overlaps. Methodsranging from alignment based on textual subtitle to contrastive encodersand temporal grounding networks can potentially bridge this gap, yetno benchmark systematically compares these paradigms or isolates thecontributions of dialogue and visual information to alignment quality.We introduce “From Script to Shot”, a benchmark of 50 feature filmsspanning nine decades with over 55K shot-level scene alignment an-notations verified by human annotators. To disentangle text matchingfrom visual grounding, we evaluate fourteen approaches—from sparsetext retrieval to recent vision-language models such as FG-CLIP 2 andQwen3-VL-Embedding, on an identical pipeline. Dialogue-based retrievalperforms poorly on non-dialogue scenes, and contrastive encoders aloneunderperform direct text matching. Adaptive multimodal fusion achievesthe strongest overall results, reaching 0.599 mSIoU when BM25 is fusedwith a Qwen3-VL-Embedding, yet even the strongest encoder leaves aconsistent gap between dialogue and non-dialogue scenes, exposing thelimits of current visual encoding for narrative grounding. By quantify-ing this gap, the benchmark provides a controlled testbed for studyingscreenplay-conditioned alignment. We demonstrate downstream utilitythrough zero-shot movie scene segmentation and screenplay-guided videogeneration, confirming that the alignment signal generalizes beyond thebenchmark itself. We release the full dataset, parser and parser outputs,toolkit, and baselines at https://github.com/jungucho92/script2shotto support research on long-form multimodal video understanding.2 J. Cho et al. 사용 폰트: 노토산스한국:https://fonts.google.com/noto/specimen/Noto+Sans+KRScreenplay -Video Alignment Prior works(dialogue -based)Screenplay SCENE 69 dialogue D D D D D Dthe script of a movie,H INT HOTEL ROOM-HOURS LATERincluding actinginstructions and D HAZEL Good morning.subtitle#69 scene directions. D FRANNIE Actually, it'sfive o'clock.Good morning. … it's five o’ … How was the park? Good morning. … it's five o’ … How was the park?D HAZEL How was the park? D D D D D D. D FRANNIE Never made it.D HAZEL Mom, what do you mean? Never made it. Mom, what … … do you mean? Never made it. Mom, what … … do you mean?#70 H headingSCENE 70 mixedfectly tailored BLACK H INT HOTEL ROOM– LATERN D D D D DN Frannie opens the door to find GusGus is here. (TO GUS) in a perfectly tailored BLACK SUIT.partial Looking sharp. Thank you ma’am . Looking sharp. Thank you ma’am . Hazel! Gus is here.D FRANNIE Hazel! Gus is here. (TOD dialogue GUS) Looking sharp.D N D N Dbathroom. She wears a D GUS Thank you ma’am.Not predictedhe looks... N A few beats later, Hazel appears in aHazel! Gus is here. Wow. Wow.pale blue sundress.D GUS Wow.N narrativeSCENE 86 non-dialogueey're readyE CLAIM to go.– DAY no subtitle ✘ N N NShot 1465 H INT AIRPORT– DAYs Michael standing Not predicted Not predicted Not predictedN Michael with sign: “My Beautiful Family.”that says - instead of #86 He kisses his wife, and hugs Hazel.y (and Gus)." Uponcourse. He kisses his Notation SCREENPLAY VIDEO SHOTS (KEYFRAMES)ake his hand butFig. 1: Screenplay-video alignment maps each shot to its corresponding screenplayscene. Each scene combines a scene heading (H), narrative direction (N), and dialogue(D). Shots carry subtitle text only when spoken dialogue is present. Prior dialogue-based methods align subtitle-matched shots correctly but fail entirely on non-dialoguescenes, whereas our benchmark evaluates all scene types under a unified protocol.