Skip to yearly menu bar Skip to main content


Show Detail
Timezone: Europe/Stockholm
 
Filter Rooms:  

TUE 8 SEP
9 a.m.
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 1:00 PM)
Workshop:
(ends 1:00 PM)
2 p.m.
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)

WED 9 SEP
9 a.m.
Workshop:
(ends 6:00 PM)
Workshop:
(ends 1:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 1:00 PM)
Workshop:
(ends 1:00 PM)
Workshop:
(ends 6:00 PM)
Workshop:
(ends 1:00 PM)
2 p.m.
Workshop:
(ends 6:00 PM)
Tutorial:
(ends 6:00 PM)

THU 10 SEP
8 a.m.
9 a.m.
Spotlights 9:00-10:15
[9:00] GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction
[9:05] Wat3R: Underwater 3D Geometry Learning without Underwater Annotations
[9:10] GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
[9:15] DreamWorld: Geometry-Grounded Video Diffusion for 3D-Consistent World Modeling
[9:20] Edit3r: Instant 3D Scene Editing from Sparse Unposed Images
[9:25] PriSplat: Propagating Reliable Multi-view Information for Distractor-Free 3DGS
[9:30] ReSplat: Learning Recurrent Gaussian Splatting
[9:35] SA-ResGS: Self-Augmented Residual 3D Gaussian Splatting for Next Best View Selection
[9:40] GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification
[9:45] SubSplat: High-Resolution Pixel-aligned 3DGS via Sub-pixel Gaussian Reparameterization
[9:50] Deformable Triangle Splatting: Flexible Primitives for Real-Time Radiance Field Rendering
[9:55] Neural Harmonic Textures for High-Quality Primitive Based Neural Reconstruction
[10:00] CubicSplat: Differentiable Vector Graphics via Error-Bounded Forward Relaxation
[10:05] RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
[10:10] TetraSDF: Analytic Isosurface Extraction with Multi-resolution Tetrahedral Grid
(ends 10:30 AM)
Orals 9:00-10:30
[9:00] Provable and Robust Wavefront Sensing via Self-Reference Interferometry
[9:15] Broadband Wide Field of View Imaging with Computational Mirrors
[9:30] Poppy: Polarization-Based Plug-and-Play Guidance for Enhancing Surface Normal Estimation
[9:45] A second-order theory of texture for depth from focus
[10:00] PRISM-VO: Scale-Aware Visual Odometry Using Photometric Plenoptic Bundle Adjustment
[10:15] Cross-View Yaw Estimation in Location Uncertainty with Line-Aligning Yaw Scoring
(ends 10:30 AM)
Spotlights 9:00-10:10
[9:00] CFM: Language-aligned Concept Foundation Model for Vision
[9:05] AdaBoosting Text Prompts for Vision-Language Models
[9:10] Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
[9:15] Molmo-Point: Better Pointing for VLMs with Grounding Tokens
[9:20] Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
[9:25] Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models
[9:30] Learning to Deny: Action Denial in Multimodal Large Language Models
[9:35] PuzLM: Solving Jigsaw Puzzles with Sequence-to-Sequence Language Models
[9:40] GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing
[9:45] On Test-Time Scaling for Vision-Language Models
[9:50] URoPE: Universal Relative Position Embedding across Geometric Spaces
[9:55] Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
[10:00] EmbedCopilot: Evaluating Vision-Language Models for Hardware-Aware Embedded System Development
[10:05] SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
(ends 10:30 AM)
10:30 a.m.
Posters 10:30-12:30
(ends 12:30 PM)
Break:
(ends 11:00 AM)
noon
Lunch:
(ends 1:30 PM)
1:30 p.m.
Orals 1:30-2:45
[1:30] Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
[1:45] OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
[2:00] OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
[2:15] Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
[2:30] FEEL (Force-Enhanced Egocentric Learning): A Dataset for Physical Action Understanding
(ends 3:00 PM)
Spotlights 1:30-2:30
[1:30] On the Reliability of Cue Conflict and Beyond
[1:35] One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models
[1:40] TR-MoE: Temporal Reliability-Aware Mixture-of-Experts for Robust Tracking
[1:45] Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer
[1:50] LoRC: Detecting AI-Generated Images via Low-Rank Collapse in the Semantic-Residual Space
[1:55] Proteus: Model Leakage-Induced Adversarial Attack in Federated Learning
[2:00] Beyond Filter Pruning: Top-K Spatial Selection for Efficient Neural Networks
[2:05] Spectral Gating via Damped Oscillations for Adaptive Implicit Neural Representations
[2:10] When Higher Order Hurts: Pre-Asymptotic Order Collapse in Generative ODE Sampling — A Theory of Discretization–Learning Interaction
[2:15] Structured-Noise Masked Modeling for Video, Audio and Beyond
[2:20] Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
[2:25] Multi-Anchor Distillation with Text-Guided Analytic Classifier for Continual Learning
(ends 3:00 PM)
Spotlights 1:30-2:30
[1:30] Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation
[1:35] Mimicking Radiologists: A Coarse-to-Fine Framework with Structural Sparse Tokens for Dual-LLM Computed Tomography Report Generation
[1:40] VesselTok: Tokenizing Vessel-like 3D Biomedical Graph Representations for Reconstruction and Generation
[1:45] RPM-Distill: Physiology-guided Adaptive Cross-modal Distillation for Robust Remote Physiological Measurement
[1:50] Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
[1:55] Ice Cloud Geometry Retrieval with Calibrated Uncertainty from Passive Satellite Imagery
[2:00] Event-based Sparse-view Background-Oriented Schlieren Tomography
[2:05] Zero-shot Depth from Defocus
[2:10] Bayesian Self-Attention with Local Pixel Correlations for Lightweight Denoising Transformers
[2:15] PolarAPP: Beyond Polarization Demosaicking for Polarimetric Applications
[2:20] Stokes-Informed Diffusion for Robust Linear Polarization Estimation
[2:25] Off the Planckian Locus: Using 2D Chromaticity to Improve In-Camera Color
(ends 3:00 PM)
3 p.m.
Keynote:
Kristen Grauman
(ends 4:00 PM)
4:30 p.m.
Break:
(ends 5:30 PM)
Posters 4:30-6:30
(ends 6:30 PM)

FRI 11 SEP
9 a.m.
Spotlights 9:00-10:15
[9:00] Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
[9:05] Mind-to-Face: Neural-Driven Photorealistic Avatar Synthesis via EEG Decoding
[9:10] GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild Videos
[9:15] OmniFace: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer
[9:20] OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data
[9:25] MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
[9:30] Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video
[9:35] Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
[9:40] EgoExoMoCap: Distributed Human Motion Capture via Ego- and Exocentric Body Tracking from Head-Mounted Devices
[9:45] Interaction-Aware 4D Gaussian Splatting for Dynamic Hand-Object Interaction Reconstruction
[9:50] OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
[9:55] CameraAnything: Refilming Videos with Arbitrary Camera Control
[10:00] RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
[10:05] A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models
[10:10] Grounding World Simulation Models in a Real-World Metropolis
(ends 10:30 AM)
Orals 9:00-10:30
[9:00] Steerable Vision Transformers
[9:15] Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
[9:30] World Knowledge in the Weights: Reading Concept Circuits of Vision Transformers
[9:45] What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
[10:00] Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
[10:15] Make Geometry Matter for Spatial Reasoning
(ends 10:30 AM)
Spotlights 9:00-10:15
[9:00] RoMa v2: Harder Better Faster Denser Feature Matching
[9:05] LoMa: Local Feature Matching Revisited
[9:10] Gravity-aware partially calibrated absolute pose estimation from affine- or rotation-covariant features
[9:15] Rolling Shutter Camera Self-Calibration
[9:20] Lost in the Tail: Addressing Geographic Imbalance in Urban Visual Place Recognition
[9:25] OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization
[9:30] E-MOTION: A Dataset for Event-Based Scene Flow Estimation with Independent Moving Objects
[9:35] SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers
[9:40] Training-free Controllable Motion Generation under Heterogeneous Constraints
[9:45] OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer
[9:50] VKSR: Scalable Kernel Surface Reconstruction Using Vecchia's Approximation
[9:55] RADmesh: Remesh-Aware Mesh Deformation
[10:00] Face Anything: 4D Face Reconstruction from Any Image Sequence
[10:05] HoloTetSphere: Unified TetSphere Mesh Reconstruction for Physical Simulations
[10:10] InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360° Image
(ends 10:30 AM)
10:30 a.m.
Posters 10:30-12:30
(ends 12:30 PM)
Break:
(ends 11:00 AM)
noon
Lunch:
(ends 1:30 PM)
1:30 p.m.
Orals 1:30-2:45
[1:30] Heat Kernel Textures -- the Geodesic Gaussians That Do Not Splat
[1:45] NURBS Splatting: A Unified Differentiable Rendering Framework for Vector Graphics
[2:00] TriFlow: Generating Artist-Like 3D Mesh Topology via Nearest-Vertex Vector Fields
[2:15] Repurposing Geometric Foundation Models for Multi-view Diffusion
[2:30] MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation
(ends 3:00 PM)
Spotlights 1:30-2:30
[1:30] Silhouette-based Gait Foundation Model
[1:35] CMCC-ReID: Cross-Modality Clothing-Change Person Re-Identification
[1:40] QVAM: Query-guided View-aware Adaptive Modulation for Aerial-Ground Person Re-Identification
[1:45] Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search
[1:50] PADFormer: Pose-agnostic Anomaly Detection from Sparse View Images
[1:55] RiO-DETR: DETR for Real-time Oriented Object Detection
[2:00] Align and Segment: Unsupervised Learning for Building Segmentation From Misaligned Labels
[2:05] Exclusivity-Guided Mask Learning for Semi-Supervised Crowd Instance Segmentation and Counting
[2:10] Instance Segmentation as Tracking: A New Paradigm for Multi-Small-Object Tracking with Event Cameras
[2:15] SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors
[2:20] PhenoLeaf-TS: A Time-Series Benchmark for Leaf Instance Segmentation, Tracking, and Growth Stage Classification
[2:25] Event-based Gaze Control Systems for Real-time Spin Estimation in Professional Ball Games
(ends 3:00 PM)
Spotlights 1:30-2:30
[1:30] UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer
[1:35] RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
[1:40] Scaling Multi-Reference Image Generation with Dynamic Reward Optimization
[1:45] AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
[1:50] Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
[1:55] EditHF-1M: A Million-Scale Rich Human Preference Feedback for Image Editing
[2:00] WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
[2:05] Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
[2:10] GIDE: Unlocking Diffusion LLMs for Precise Training-Free Image Editing
[2:15] Push–Pull Attentional Anchoring for Diffusion Concept Erasure
[2:20] When Rubrics Fail: Error Enumeration as Reward for Reference-Free RL Post-Training
[2:25] Drift-AR: Single-Step Visual Autoregressive Generation via Anti-Symmetric Drifting
(ends 3:00 PM)
3 p.m.
Keynote:
Yann LeCun
(ends 4:00 PM)
4 p.m.
Posters 4:00-6:00
(ends 6:00 PM)
6 p.m.
Reception:
(ends 9:00 PM)

SAT 12 SEP
9 a.m.
Keynote:
Jamie Shotton
(ends 10:00 AM)
10 a.m.
Panel:
(ends 10:30 AM)
10:30 a.m.
Break:
(ends 11:00 AM)
Posters 10:30-12:30
(ends 12:30 PM)
noon
Lunch:
(ends 1:30 PM)
1:30 p.m.
Spotlights 1:30-2:40
[1:30] SPEAR: A Simulator for Photorealistic Embodied AI Research
[1:35] Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
[1:40] Unordered Landmark Visual Navigation
[1:45] R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation
[1:50] Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum
[1:55] Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch
[2:00] Cooking beyond Frames: A Stereo Event Camera Dataset in the Kitchen
[2:05] TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception
[2:10] VOCA: Visual Odometry with Codec Awareness
[2:15] PriorEye: Geospatial Visual Priors for End-to-End Autonomous Driving
[2:20] Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior
[2:25] MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model
[2:30] CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization
[2:35] VIPS: Vehicle-Infrastructure Cooperative Planning Benchmark via Pseudo-Simulation
(ends 3:00 PM)
Spotlights 1:30-2:45
[1:30] DoCoG: Mask-based Multi-Type Grounded Chain-of-Thought for Document QA
[1:35] Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity
[1:40] SyntheticDoc: A Large Synthetic Dataset for Document Unwarping and Illumination Correction
[1:45] PanoRec: Spatially-Structured Sequence Modeling for Multi-Granularity Panoramic Retrieval
[1:50] NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
[1:55] Incentivizing Vision Language Models to Search for Long Video Question Answering
[2:00] Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
[2:05] HumanOmni-Speaker: Identifying Who said What and When
[2:10] UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation
[2:15] See & Sniff: Learning Visuo-Olfactory Representations
[2:20] Bridging Vision and Language Concepts through Optimal Transport Semantic Flow
[2:25] CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
[2:30] When Sinks Help or Hurt: Unified Framework for Attention Sink in MLLMs
[2:35] How to Teach Large Multimodal Models New Skills
[2:40] Event-Driven Video Generation
(ends 3:00 PM)
Orals 1:30-3:00
[1:30] Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
[1:45] Geometric Context Transformer for Streaming 3D Reconstruction
[2:00] LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
[2:15] GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
[2:30] Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
[2:45] Seeing Through the Weights: Privacy Leakage in Scene Coordinate Regression
(ends 3:00 PM)
3 p.m.
Posters 3:00-5:00
(ends 5:00 PM)