Skip to yearly menu bar Skip to main content


Show Detail
 
Filter Rooms:  

MON 7 SEP
3 p.m.
(ends 8:00 PM)

TUE 8 SEP
7:30 a.m.
(ends 3:30 PM)
8:45 a.m.
noon
Lunch:
(ends 1:45 PM)
1:15 p.m.
Workshop:
(ends 5:30 PM)
1:45 p.m.
Workshop:
(ends 6:00 PM)

WED 9 SEP
7:30 a.m.
(ends 3:30 PM)
noon
Lunch:
(ends 1:45 PM)

THU 10 SEP
7:30 a.m.
(ends 3:30 PM)
8 a.m.
Facilities:
(ends 5:00 PM)
Facilities:
(ends 5:00 PM)
9 a.m.
Orals 9:00-10:30
[9:00] Provable and Robust Wavefront Sensing via Self-Reference Interferometry
[9:15] Broadband Wide Field of View Imaging with Computational Mirrors
[9:30] Poppy: Polarization-Based Plug-and-Play Guidance for Enhancing Surface Normal Estimation
[9:45] A second-order theory of texture for depth from focus
[10:00] PRISM-VO: Scale-Aware Visual Odometry Using Photometric Plenoptic Bundle Adjustment
[10:15] Cross-View Yaw Estimation in Location Uncertainty with Line-Aligning Yaw Scoring
(ends 10:30 AM)
Facilities:
(ends 4:00 PM)
Facilities:
(ends 4:00 PM)
Spotlights 9:00-9:15
[9:00] GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction
[9:05] Wat3R: Underwater 3D Geometry Learning without Underwater Annotations
[9:10] GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
Q&As 9:15-9:18
[9:15] Q&A
Spotlights 9:18-9:33
[9:18] DreamWorld: Geometry-Grounded Video Diffusion for 3D-Consistent World Modeling
[9:23] Edit3r: Instant 3D Scene Editing from Sparse Unposed Images
[9:28] PriSplat: Propagating Reliable Multi-view Information for Distractor-Free 3DGS
Q&As 9:33-9:36
[9:33] Q&A
Spotlights 9:36-9:51
[9:36] ReSplat: Learning Recurrent Gaussian Splatting
[9:41] SA-ResGS: Self-Augmented Residual 3D Gaussian Splatting for Next Best View Selection
[9:46] GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification
Q&As 9:51-9:54
[9:51] Q&A
Spotlights 9:54-10:09
[9:54] SubSplat: High-Resolution Pixel-aligned 3DGS via Sub-pixel Gaussian Reparameterization
[9:59] Deformable Triangle Splatting: Flexible Primitives for Real-Time Radiance Field Rendering
[10:04] Neural Harmonic Textures for High-Quality Primitive Based Neural Reconstruction
Q&As 10:09-10:12
[10:09] Q&A
Spotlights 10:12-10:27
[10:12] CubicSplat: Differentiable Vector Graphics via Error-Bounded Forward Relaxation
[10:17] RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
[10:22] TetraSDF: Analytic Isosurface Extraction with Multi-resolution Tetrahedral Grid
Q&As 10:27-10:30
[10:27] Q&A
(ends 10:30 AM)
Spotlights 9:00-9:15
[9:00] CFM: Language-aligned Concept Foundation Model for Vision
[9:05] AdaBoosting Text Prompts for Vision-Language Models
[9:10] Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
Q&As 9:15-9:18
[9:15] Q&A
Spotlights 9:18-9:33
[9:18] Molmo-Point: Better Pointing for VLMs with Grounding Tokens
[9:23] Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
[9:28] Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models
Q&As 9:33-9:36
[9:33] Q&A
Spotlights 9:36-9:51
[9:36] Learning to Deny: Action Denial in Multimodal Large Language Models
[9:41] PuzLM: Solving Jigsaw Puzzles with Sequence-to-Sequence Language Models
[9:46] GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing
Q&As 9:51-9:54
[9:51] Q&A
Spotlights 9:54-10:09
[9:54] On Test-Time Scaling for Vision-Language Models
[9:59] URoPE: Universal Relative Position Embedding across Geometric Spaces
[10:04] Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
Q&As 10:09-10:12
[10:09] Q&A
Spotlights 10:12-10:22
[10:12] EmbedCopilot: Evaluating Vision-Language Models for Hardware-Aware Embedded System Development
[10:17] SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
(ends 10:22 AM)
10:30 a.m.
Posters 10:30-12:30
(ends 12:30 PM)
Break:
(ends 11:00 AM)
11 a.m.
noon
Lunch:
(ends 1:30 PM)
12:30 p.m.
Mentorship:
(ends 1:30 PM)
1:30 p.m.
Art:
(ends 2:30 PM)
Orals 1:30-2:45
[1:30] Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
[1:45] OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
[2:00] OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
[2:15] Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
[2:30] FEEL (Force-Enhanced Egocentric Learning): A Dataset for Physical Action Understanding
(ends 2:45 PM)
Spotlights 1:30-1:45
[1:30] Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation
[1:35] Mimicking Radiologists: A Coarse-to-Fine Framework with Structural Sparse Tokens for Dual-LLM Computed Tomography Report Generation
[1:40] VesselTok: Tokenizing Vessel-like 3D Biomedical Graph Representations for Reconstruction and Generation
Q&As 1:45-1:48
[1:45] Q&A
Spotlights 1:48-2:03
[1:48] RPM-Distill: Physiology-guided Adaptive Cross-modal Distillation for Robust Remote Physiological Measurement
[1:53] Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
[1:58] Ice Cloud Geometry Retrieval with Calibrated Uncertainty from Passive Satellite Imagery
Q&As 2:03-2:06
[2:03] Q&A
Spotlights 2:06-2:21
[2:06] Event-based Sparse-view Background-Oriented Schlieren Tomography
[2:11] Zero-shot Depth from Defocus
[2:16] Bayesian Self-Attention with Local Pixel Correlations for Lightweight Denoising Transformers
Q&As 2:21-2:24
[2:21] Q&A
Spotlights 2:24-2:39
[2:24] PolarAPP: Beyond Polarization Demosaicking for Polarimetric Applications
[2:29] Stokes-Informed Diffusion for Robust Linear Polarization Estimation
[2:34] Off the Planckian Locus: Using 2D Chromaticity to Improve In-Camera Color
Q&As 2:39-2:42
[2:39] Q&A
(ends 2:42 PM)
Spotlights 1:30-1:45
[1:30] On the Reliability of Cue Conflict and Beyond
[1:35] One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models
[1:40] TR-MoE: Temporal Reliability-Aware Mixture-of-Experts for Robust Tracking
Q&As 1:45-1:48
[1:45] Q&A
Spotlights 1:48-2:03
[1:48] Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer
[1:53] LoRC: Detecting AI-Generated Images via Low-Rank Collapse in the Semantic-Residual Space
[1:58] Proteus: Model Leakage-Induced Adversarial Attack in Federated Learning
Q&As 2:03-2:06
[2:03] Q&A
Spotlights 2:06-2:21
[2:06] Beyond Filter Pruning: Top-K Spatial Selection for Efficient Neural Networks
[2:11] Spectral Gating via Damped Oscillations for Adaptive Implicit Neural Representations
[2:16] When Higher Order Hurts: Pre-Asymptotic Order Collapse in Generative ODE Sampling — A Theory of Discretization–Learning Interaction
Q&As 2:21-2:24
[2:21] Q&A
Spotlights 2:24-2:39
[2:24] Structured-Noise Masked Modeling for Video, Audio and Beyond
[2:29] Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
[2:34] Multi-Anchor Distillation with Text-Guided Analytic Classifier for Continual Learning
Q&As 2:39-2:42
[2:39] Q&A
(ends 2:42 PM)
2:45 p.m.
Break:
(ends 3:00 PM)
3 p.m.
Keynote:
Kristen Grauman
(ends 4:00 PM)
4:30 p.m.
Break:
(ends 5:00 PM)
Posters 4:30-6:30
(ends 6:00 PM)
5 p.m.

FRI 11 SEP
8 a.m.
Facilities:
(ends 5:00 PM)
Facilities:
(ends 5:00 PM)
8:30 a.m.
(ends 3:30 PM)
9 a.m.
Facilities:
(ends 4:00 PM)
Facilities:
(ends 4:00 PM)
Orals 9:00-10:30
[9:00] Steerable Vision Transformers
[9:15] Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
[9:30] World Knowledge in the Weights: Reading Concept Circuits of Vision Transformers
[9:45] What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
[10:00] Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
[10:15] Make Geometry Matter for Spatial Reasoning
(ends 10:30 AM)
Spotlights 9:00-9:15
[9:00] RoMa v2: Harder Better Faster Denser Feature Matching
[9:05] LoMa: Local Feature Matching Revisited
[9:10] Gravity-aware partially calibrated absolute pose estimation from affine- or rotation-covariant features
Q&As 9:15-9:18
[9:15] Q&A
Spotlights 9:18-9:33
[9:18] Rolling Shutter Camera Self-Calibration
[9:23] Lost in the Tail: Addressing Geographic Imbalance in Urban Visual Place Recognition
[9:28] OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization
Q&As 9:33-9:36
[9:33] Q&A
Spotlights 9:36-9:51
[9:36] E-MOTION: A Dataset for Event-Based Scene Flow Estimation with Independent Moving Objects
[9:41] SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers
[9:46] Training-free Controllable Motion Generation under Heterogeneous Constraints
Q&As 9:51-9:54
[9:51] Q&A
Spotlights 9:54-10:09
[9:54] OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer
[9:59] VKSR: Scalable Kernel Surface Reconstruction Using Vecchia's Approximation
[10:04] RADmesh: Remesh-Aware Mesh Deformation
Q&As 10:09-10:12
[10:09] Q&A
Spotlights 10:12-10:27
[10:12] Face Anything: 4D Face Reconstruction from Any Image Sequence
[10:17] HoloTetSphere: Unified TetSphere Mesh Reconstruction for Physical Simulations
[10:22] InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360° Image
Q&As 10:27-10:30
[10:27] Q&A
(ends 10:30 AM)
Spotlights 9:00-9:15
[9:00] Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
[9:05] Mind-to-Face: Neural-Driven Photorealistic Avatar Synthesis via EEG Decoding
[9:10] GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild Videos
Q&As 9:15-9:18
[9:15] Q&A
Spotlights 9:18-9:33
[9:18] OmniFace: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer
[9:23] OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data
[9:28] MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
Q&As 9:33-9:36
[9:33] Q&A
Spotlights 9:36-9:51
[9:36] Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video
[9:41] Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
[9:46] EgoExoMoCap: Distributed Human Motion Capture via Ego- and Exocentric Body Tracking from Head-Mounted Devices
Q&As 9:51-9:54
[9:51] Q&A
Spotlights 9:54-10:09
[9:54] Interaction-Aware 4D Gaussian Splatting for Dynamic Hand-Object Interaction Reconstruction
[9:59] OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
[10:04] CameraAnything: Refilming Videos with Arbitrary Camera Control
Q&As 10:09-10:12
[10:09] Q&A
Spotlights 10:12-10:27
[10:12] RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
[10:17] A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models
[10:22] Grounding World Simulation Models in a Real-World Metropolis
Q&As 10:27-10:30
[10:27] Q&A
(ends 10:30 AM)
10:30 a.m.
Break:
(ends 11:00 AM)
11 a.m.
Posters 10:30-12:30
(ends 12:30 PM)
noon
Lunch:
(ends 1:30 PM)
12:30 p.m.
Mentorship:
(ends 1:30 PM)
Mentorship:
(ends 1:30 PM)
1:30 p.m.
Orals 1:30-2:45
[1:30] Heat Kernel Textures -- the Geodesic Gaussians That Do Not Splat
[1:45] NURBS Splatting: A Unified Differentiable Rendering Framework for Vector Graphics
[2:00] TriFlow: Generating Artist-Like 3D Mesh Topology via Nearest-Vertex Vector Fields
[2:15] Repurposing Geometric Foundation Models for Multi-view Diffusion
[2:30] MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation
(ends 2:45 PM)
Spotlights 1:30-1:45
[1:30] UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer
[1:35] RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
[1:40] Scaling Multi-Reference Image Generation with Dynamic Reward Optimization
Q&As 1:45-1:48
[1:45] Q&A
Spotlights 1:48-2:03
[1:48] AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
[1:53] Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
[1:58] EditHF-1M: A Million-Scale Rich Human Preference Feedback for Image Editing
Q&As 2:03-2:06
[2:03] Q&A
Spotlights 2:06-2:21
[2:06] WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
[2:11] Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
[2:16] GIDE: Unlocking Diffusion LLMs for Precise Training-Free Image Editing
Q&As 2:21-2:24
[2:21] Q&A
Spotlights 2:24-2:39
[2:24] Push–Pull Attentional Anchoring for Diffusion Concept Erasure
[2:29] When Rubrics Fail: Error Enumeration as Reward for Reference-Free RL Post-Training
[2:34] Drift-AR: Single-Step Visual Autoregressive Generation via Anti-Symmetric Drifting
Q&As 2:39-2:42
[2:39] Q&A
(ends 2:42 PM)
Spotlights 1:30-1:45
[1:30] Silhouette-based Gait Foundation Model
[1:35] CMCC-ReID: Cross-Modality Clothing-Change Person Re-Identification
[1:40] QVAM: Query-guided View-aware Adaptive Modulation for Aerial-Ground Person Re-Identification
Q&As 1:45-1:48
[1:45] Q&A
Spotlights 1:48-2:03
[1:48] Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search
[1:53] PADFormer: Pose-agnostic Anomaly Detection from Sparse View Images
[1:58] RiO-DETR: DETR for Real-time Oriented Object Detection
Q&As 2:03-2:06
[2:03] Q&A
Spotlights 2:06-2:21
[2:06] Align and Segment: Unsupervised Learning for Building Segmentation From Misaligned Labels
[2:11] Exclusivity-Guided Mask Learning for Semi-Supervised Crowd Instance Segmentation and Counting
[2:16] Instance Segmentation as Tracking: A New Paradigm for Multi-Small-Object Tracking with Event Cameras
Q&As 2:21-2:24
[2:21] Q&A
Spotlights 2:24-2:39
[2:24] SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors
[2:29] PhenoLeaf-TS: A Time-Series Benchmark for Leaf Instance Segmentation, Tracking, and Growth Stage Classification
[2:34] Event-based Gaze Control Systems for Real-time Spin Estimation in Professional Ball Games
Q&As 2:39-2:42
[2:39] Q&A
(ends 2:42 PM)
2:45 p.m.
Break:
(ends 3:00 PM)
3 p.m.
4 p.m.
Posters 4:00-6:00
(ends 6:00 PM)
5 p.m.
6 p.m.
Reception:
(ends 9:00 PM)

SAT 12 SEP
8 a.m.
Facilities:
(ends 5:00 PM)
Facilities:
(ends 5:00 PM)
8:30 a.m.
(ends 1:00 PM)
9 a.m.
Facilities:
(ends 4:00 PM)
Facilities:
(ends 1:00 PM)
Keynote:
Jamie Shotton
(ends 10:00 AM)
10 a.m.
Panel:
(ends 10:30 AM)
10:30 a.m.
Break:
(ends 11:00 AM)
11 a.m.
Posters 10:30-12:30
(ends 12:30 PM)
noon
Lunch:
(ends 1:30 PM)
12:30 p.m.
Mentorship:
(ends 1:30 PM)
1:30 p.m.
Orals 1:30-3:00
[1:30] Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
[1:45] Geometric Context Transformer for Streaming 3D Reconstruction
[2:00] LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
[2:15] GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
[2:30] Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
[2:45] Seeing Through the Weights: Privacy Leakage in Scene Coordinate Regression
(ends 3:00 PM)
Spotlights 1:30-1:45
[1:30] DoCoG: Mask-based Multi-Type Grounded Chain-of-Thought for Document QA
[1:35] SyntheticDoc: A Large Synthetic Dataset for Document Unwarping and Illumination Correction
[1:40] PanoRec: Spatially-Structured Sequence Modeling for Multi-Granularity Panoramic Retrieval
Q&As 1:45-1:48
[1:45] Q&A
Spotlights 1:48-2:03
[1:48] NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
[1:53] Incentivizing Vision Language Models to Search for Long Video Question Answering
[1:58] Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
Q&As 2:03-2:06
[2:03] Q&A
Spotlights 2:06-2:21
[2:06] HumanOmni-Speaker: Identifying Who said What and When
[2:11] UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation
[2:16] See & Sniff: Learning Visuo-Olfactory Representations
Q&As 2:21-2:24
[2:21] Q&A
Spotlights 2:24-2:39
[2:24] Bridging Vision and Language Concepts through Optimal Transport Semantic Flow
[2:29] CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
[2:34] When Sinks Help or Hurt: Unified Framework for Attention Sink in MLLMs
Q&As 2:39-2:42
[2:39] Q&A
Spotlights 2:42-2:52
[2:42] How to Teach Large Multimodal Models New Skills
[2:47] Event-Driven Video Generation
(ends 2:52 PM)
Spotlights 1:30-1:45
[1:30] SPEAR: A Simulator for Photorealistic Embodied AI Research
[1:35] Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
[1:40] Unordered Landmark Visual Navigation
Q&As 1:45-1:48
[1:45] Q&A
Spotlights 1:48-2:03
[1:48] R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation
[1:53] Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum
[1:58] Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch
Q&As 2:03-2:06
[2:03] Q&A
Spotlights 2:06-2:21
[2:06] Cooking beyond Frames: A Stereo Event Camera Dataset in the Kitchen
[2:11] TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception
[2:16] VOCA: Visual Odometry with Codec Awareness
Q&As 2:21-2:24
[2:21] Q&A
Spotlights 2:24-2:39
[2:24] PriorEye: Geospatial Visual Priors for End-to-End Autonomous Driving
[2:29] Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior
[2:34] MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model
Q&As 2:39-2:42
[2:39] Q&A
Spotlights 2:42-2:52
[2:42] CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization
[2:47] VIPS: Vehicle-Infrastructure Cooperative Planning Benchmark via Pseudo-Simulation
(ends 2:52 PM)
3 p.m.
Posters 3:00-5:00
(ends 5:00 PM)