Skip to yearly menu bar
Skip to main content
Main Navigation
Select Year: (2026)
2026
2024
2022
My Stuff
Create Profile
Reset Password
Login
Getting Started
Schedule
Tutorials
Workshops
Main Conference
Keynotes and Panels
Orals
Spotlights
Papers
Paper Awards
Sponsors
Organizers
Help
Layout:
mini
compact
topic
detail
×
No topics available
No sessions available
title
author
topic
session
shuffle
by
serendipity
bookmarked first
visited first
not visited first
bookmarked but not visited
Loading...
Enable Javascript in your browser to see the papers page.
Reflection-Aware Reasoning for Non-Line-of-Sight Pedestrian Localization
Egocentric Procedure Parsing
Kirin: Animal Motion Generation from In-the-Wild Video
LinStereo: Linear-Complexity Global Attention for Multi-Scale Iterative Stereo Matching
MedCAGD: Context-Aware Gated Decoder for Robust Medical Image Segmentation
3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints
Implicit Neural Representation Facilitates Unified Universal Vision Encoding
InstrAct: Towards Action-Centric Understanding in Instructional Videos
Fine-Grained Text-to-Video Retrieval for Camera-Trap Data
Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE
Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
From Blobs to Spokes: High-Fidelity Surface Reconstruction via Oriented Gaussians
G-ZAP: A Generalizable Zero-Shot Framework for Arbitrary-Scale Pansharpening
GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction
CCFM: Collision-Constrained Flow Matching for Safety-Critical Scenario Generation
Molecular Identifier Visual Prompting and Verifiable Reinforcement Learning for Chemical Reaction Diagram Parsing
MonoArt: Progressive Structural Reasoning for Monocular Articulated 3D Reconstruction
Region-Aware Multimodal Interleaving for Animal Re-Identification
StratoSplat: Taming Layered Regularities for Sparse Aerial 3D Gaussian Splatting
Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration
Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation
Parsimonious Flow Matching for Efficient Image Generation
When W4A4 Breaks Camouflaged Object Detection: Token-Group Dual-Constraint Activation Quantization
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
Less is More: A Simple yet Effective Object-Centric Prompting Strategy for Vision-Language Reasoning in Autonomous Driving
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
Revisiting Avatar-As-Image: High-Fidelity Registration is All You Need
SCORE: SubDistribution-aware Collaborative Knowledge Reinforcing for Cloth-Hybrid Lifelong Person Re-Identification
STANCE: Controllable Video Generation for Structured Dynamics via Sparse-To-dense ANChored Encoding
StereoGS: Sparse-View 3D Gaussian Splatting via Stereo Priors
ZMIS-SAM: Segment Anything Model Enhanced With Wavelet Transform For Zooplankton Microscopy Image Instance Segmentation
Adaptive Neural Dynamics for Robust Geometric LiDAR-Inertial State Estimation on UAVs
AFFMAE: Scalable Vision Pre-Training for High-Resolution Microscopy Segmentation on Desktop Hardware
ESTANet: Efficient Online Error Detection in Procedural Videos via Prediction Inconsistency
If It's Not Efficient, It's Not Usable: Real-Time OOD Detection with Latent De-Biasing and High-Quality Negative Samples
OBBSeg: Irregular Lesion Segmentation under Oriented Bounding Box Annotations
How Far Are Video Models from True Multimodal Reasoning?
Scene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene Generation
RainODE: Continuous-Time Precipitation Forecasting with Latent Neural ODEs
ReShift: Aha-Moment-Driven Reasoning-Level Backdoor Attacks on Vision–Language Models
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
Human Mesh Modeling for Anny Body
PPTArena: A Benchmark for PowerPoint Editing
Tuning Real-World Image Restoration at Inference: A Test-Time Scaling Paradigm for Flow Matching Models
Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild
Why Linear Probing Works: Non-Vacuous Generalization Bounds via Effective Dimension
Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views
Different Changes Require Different Reasoning: Change-Type-Specialized Experts for Robust Change Captioning
Frequency Director: Learnable Mixture of Frequency Experts for Unified Concealed Scene Segmentation
MMAgent-R2: Learning to Rerank and Reject for Agentic mRAG
FaCT-GS: Fast and Scalable CT Reconstruction with Gaussian Splatting
FST-SAM3: Taming SAM~3 with Frequency-Spatio-Temporal Refinement for Video Polyp Segmentation
Music-to-Dance Generation via Atomic Movements
StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production–Living Simulations with Stardew Valley
Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning
UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception
Don’t Starve the Boundaries: Boundary-Constrained Label Propagation for Weakly Supervised 3D Segmentation
InfiniteDance: Scalable 3D Dance Generation Towards in-the-wild Generalization
Rethinking Adversary in Semantic Segmentation: An Out-of-Distribution Perspective
Rethinking Visual Privacy: A Compositional Privacy Risk Framework for Severity Assessment with VLMs
Self-Improving Diffusion Classifiers with Minority Preference Optimization
Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs
Sparsity-Inducing Divergence Losses for Biometric Verification
Coordinate Singularities Break Conformal Coverage for Gaze and Head Pose
Counting Trees from Satellite Imagery with Noisy Supervision
Training-free Controllable Motion Generation under Heterogeneous Constraints
Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples
MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
Adversarial Score Distillation for Stable One-Step Diffusion in Real-World Image Super-Resolution
Auto-Prompting: Layer-Specific Prompt Fusion Discovery via Differentiable Search
Geo-ID: Test-Time Geometric Consensus for Cross-View Consistent Intrinsics
Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation
UNet-Twice: A Simple Structured Reference-based Inpainting Framework
UniH3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration
MVFusion-GS: Motion-Variance Guided Temporal Attention for High-Quality Dynamic Gaussian Splatting
RePL: Pseudo-label Refinement for Semi-supervised LiDAR Semantic Segmentation
Rethinking Robust Adversarial Concept Erasure in Diffusion Models
ActiveStructure: Plane Scene Graph-Guided Active 3D Gaussian Splatting
GeoEdit: Geometry-Aware Object Editing via Dual-Branch Denoising
OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding
Efficient Camera Pose Augmentation for View Generalization in Robotic Policy Learning
TecoPrompt: Temporal-Conservative Prompt Learning for Vision-Language Models
Multi-modal Knowledge Preserving Adapter for Embedding Backward Compatibility
Wavelet-based Intra-video Counterfactual Reasoning for Video Question Grounding
MotionSplicer: Part-Based Motion Editing for 4D Volumetric Videos
Gripper-aware Vision Language Action Models
REALM: An RGB and Event Aligned Latent Manifold for Cross-Modal Perception
Retrieving and Refining Winning Noise Tickets for Diffusion-Based Motion Generation
HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration
GraphVid: Interactive Graph-Controllable Video Generation
Monocular Avatar Reconstruction via Cascaded Diffusion Priors and UV-Space Differentiable Shading
RegHead: Non-Humanoid Head Blendshapes via Feed-Forward Registration
Geometry-Anchored Transport Framework for Exemplar-Free Class-Incremental Learning
LayerVerse: Finding the Sweet Spot for KV-Injection in Training-Free Image Editing
NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation
Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning
HumanOmni-Speaker: Identifying Who said What and When
UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving
Same Pool, Different Answer: Stable Best-of-N Selection for Vision-Language Models
HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding
Compact Low-Cost Hyperspectral Imaging via Angular-to-Spectral Diversity Conversion
Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation
Gradient sparsity regularization for training unlearning-compatible models
GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces
The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models
Weather-Conditioned Depth Anything
Latent Fusion: Decoding Consolidated 3D Geometry from Feed-forward Geometry Transformer Latents
Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs
DRS-VPT: Directly Re-localizing in Scenes using a Vision and Point Transformer
Gender Bias in Vision-Language In-Context Learning
Rethinking Temporal Modeling in Visual Object Tracking via Decoupled Auxiliary Supervision
Image Warping for Image-to-Image Translation
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection
Align and Segment: Unsupervised Learning for Building Segmentation From Misaligned Labels
Maximum Spanning Tree Guided Confidence and Sparse Graph for Robust Noisy Label Learning
Purify then Guide: Rethinking Domain Generalization for Multimodal Face Anti-Spoofing
Do Vision Language Models Recognize Visual Ambiguity?
Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration
Kiroshi: An Agentic Perception System for High-Accuracy Image Parsing
Semantic-Anchored Evidential Fusion for Domain-Robust Whole-Slide Survival Analysis
Knowledge-Centric Agents for Workflow Generation in ComfyUI
Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation
Tackling Misattribution in 3D Intrinsic Decomposition via Proximity Attention Point Rendering
P-CORE: Self-Supervised Surface Consistency for Point-Based Neural Editing
Indelible Backdoors: On the Limits of Post-Training Defenses
Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
Decompose, Compare, and Decide: Multimodal LLMs are Implicit Few-Shot Learners
TaxoMIL: Taxonomy-Constrained Learning for Hierarchical Whole Slide Image Analysis
JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising
Exo2EgoPolicy: Pose-Aligned Cross-View Policy Learning
Fixed Reality, Diffused Possibility: Disentangling Stochastic and Deterministic Latent for Cluttered Grasping
CrossFeat: Bridging Imaging Modalities in Feature Descriptor Space
LoRC: Detecting AI-Generated Images via Low-Rank Collapse in the Semantic-Residual Space
Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation
Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning
Learn to Rank: Visual Attribution by Learning Importance Ranking
ATP-Bench: Towards the Agentic Tool Planning for MLLM Interleaved Generation
Learning Ego-Centric BEV Representations from a Perspective-Privileged View: Cross-View Supervision for Online HD Map Construction
D²R²OSR: Degradation-Disentangled Representation for Real-World Omnidirectional Image Super-Resolution
Training-Free Multi-Concept Image Editing
Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion
OctWorld: Long-Range World-Consistent Video Generation with Octree-based 3D Mapping
SemCityLoc: Aerial 6DoF Localization Using Semantic 3D City Models
AnyView: Synthesizing Any Novel View in Dynamic Scenes
Trajectory Forcing: Structure-First Generation with Controllable Semantic Trajectories
LEGO: Leveled Language Gaussian Splatting
VPA-WM: Vision-Priors-Aligned World Models for Robust Visual Reinforcement Learning
Bootstrapping Articulated 3D Reconstruction from 2D Image Collections
World Models for Learning Dexterous Hand-Object Interactions from Human Videos
InstantHDR: Single-forward Gaussian Splatting for High Dynamic Range 3D Reconstruction
Rethinking Training and Inference for Trajectory Forecasting: Linking Winner-Take-All back to GMMs
When the City Teaches the Car: Label-Free 3D Perception from Infrastructure
Towards Video Anomaly Detection from Event Streams: A Baseline and Benchmark Datasets
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
Temporally Aware Densification for Dynamic 3D Gaussian Splatting
Holo360D: A Large-Scale Real-World Dataset with Continuous Trajectories for Advancing Panoramic 3D Reconstruction and Beyond
Caption Bottleneck Models
CtrlCoMo: Controllable Co-Speech Motion Generation with Gesture–Action Disentanglement
Following the Flow: Advection-Consistent Modeling for Event-based Small Object Detection
Don’t Mask Out the Background! Natural-Light Photometric Stereo via Illumination Reconstruction
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
SparseDriveV2: Scoring is All You Need for End-to-End Autonomous Driving
ShellMaker: Language-Guided Exterior Completion under Structural Constraints
MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation
Pointer-CAD v2: Plan-Then-Construct CAD Generation with Dimension-Aware Parametric Precision
Data-Free Client Contribution Estimation via Logit Maximization for Federated Learning
BAAF: Universal Transformation of One-Class Classifiers for Unsupervised Image Anomaly Detection
Unified and Efficient Point-Line Local Features
SupIR-GS: Thermal Infrared Super-Resolution Novel View Synthesis with Imaging-Calibrated 3D Gaussian Splatting
Reconstruction by Generation: 3D Multi-Object Scene Reconstruction from Sparse Observations
TanGO: Training-Free 3D Editing via Tangent-Space Guidance and Optimization
InSpace: Structure-Aware 3D Indoor Scene Generation from a Single 360° Image
Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging
StAR: Segment Anything Reasoner
Conversational Human Audio-visual Talking Dialogue Generation
NeSy-Route: A Neural-Symbolic Benchmark for Constrained Route Planning in Remote Sensing
TopoAgent: An Agentic Framework for Automated Topology Learning in Medical Imaging
Proto-Gaussian: MRI Modality Translation Based on Learnable Structural Prototypes and 2D Gaussian Splatting
Pixel-wise Geo-registration of Drone and Satellite Images
SubSplat: High-Resolution Pixel-aligned 3DGS via Sub-pixel Gaussian Reparameterization
GameWorlds: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
Discovering Geometric Biases in 3D Face Reconstruction: A Curvature-Aware Spectral Framework for Fairness Evaluation
Video Generative Models as Geometry Learner
ExPLoRe: Expert Patch-Level Loss Routing for Multi-Objective Masked Image Modeling
HoloTetSphere: Unified TetSphere Mesh Reconstruction for Physical Simulations
Mitigating Positional Leakage in 3D Masked Autoencoders for Robust Representation Learning
Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation
Leaving the City: A Large-Scale Aerial Dataset for Cross-Season Localization in Unstructured Environments
Event-based Gaze Control Systems for Real-time Spin Estimation in Professional Ball Games
InstantRetouch: Personalized Image Retouching without Test-time Fine-tuning
Spectral Consistent Flow for One-step 3D Medical Image Translation
UniQueR: Unified Query-based Feedforward 3D Reconstruction
Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation
DA-MergeLoRA: Hypernetwork-Based LoRA Merging for Few-Shot Test-Time Domain Adaptation
FreqOrtho-SR: Frequency-Guided Orthogonal Expert Learning for Real-World Image Super-Resolution
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
Iterative Perceptual Alignment for VLMs via Deterministic Reconstruction Feedback
Falcon: Functional Assembly and Language for Compositional Reasoning in X-ray
Teaching an Agent to Sketch One Part at a Time
Zero-Shot Image Personalization from Personas
VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction
FaceMoE: Mixture of Experts for Low-Resolution Face Recognition
Visual Prompt Discovery via Semantic Exploration
Visible Yet Unrecognizable: Frequency-Selective Facial Privacy via Attention
FontCopilot: Towards Generalist Multimodal Large Language Models for Holistic Chinese Font Engineering
Accurate Zero-shot Quantization via Hierarchical Teacher-Assistant Distillation
Diffusion Integrated Gradients: Controllable Path Generation for Flexible Feature Attribution
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
Generalizable Neural Reconstruction of High-Fidelity Surfaces via Sparse Volumetric Representations
DP-BOA: Dirichlet-Process Birth-or-Assign for On-the-Fly Category Discovery
ActionParty: Multi-Subject Action Binding in Generative Video Games
SPEAR: A Simulator for Photorealistic Embodied AI Research
Deep Noise Label Learning via Effective Rank Reduction
From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space
PhysConvex: Physics-Informed Dynamic Convex Fields for Reconstruction and Simulation
TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy
Saber: Anchoring Semantics to Scale-Aware Kinetic Salience for Zero-Shot Skeleton Action Recognition
Modeling and Compensating Phase Error in High-speed 3D Reconstruction
Fast and Flexible Robustness Certificates for Semantic Segmentation
What if? Emulative Simulation with World Models for Situated Reasoning
UniTeD: Unified Temporal Diffusion for Joint Perception and Planning in Autonomous Driving
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
Recognizing Co-Speech Gestures in-the-Wild
Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
MLP Splatting: Object-Centric Neural Fields
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models
Nexus-Vid: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
DualResPS: Dual-Resolution Photometric Stereo Using a Frame-Event Hybrid Camera
C3ASD: Multi-Level Consistency-Driven Representation Learning for Robust Active Speaker Detection
TopoFuse: Topology-Aware Tri-Planar Fusion for 3D Cryo-Electron Tomography Segmentation
BiSLW: Bi-Spectral Latent Watermarking for Generative Diffusion Models
MotionDreamer: Universal Skeletal Motion Generation for 3D Rigged Shapes
SSBP: Stage-Specialized Block Pruning for Video Diffusion Models
DIVA: Instruction-Aware Vision Token Pruning via Dual-Probe Attention Discrepancy
SIGNET: Motion-Level Knowledge Transfer for Cross-Language Sign Language Translation
PointGT: Simultaneous Geometric and Textural Editing for Point-Based Representations
Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation
LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
Filterless Snapshot Hyperspectral Imaging using Guided Patch Diffusion
SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation
Improving Sparse-View 3DGS Generalization via Flat Minima Optimization
CLDefocus: Physically Grounded Compound-Lens Defocus Blur Synthesis
Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models
PackForcing: Short Video Training Suffices for Long Video Sampling and Long Context Inference
UniReflect: Self-Reflection Tuning for Unified Multimodal Understanding and Generation
Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space
Training-Free Task Classification for Multi-Task Model Merging
iMED: A Multi-Endoscope Dataset for Surgical 3D Perception
Diffusion Models are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning
Pathryoshka: Compressing Pathology Foundation Models via Multi-Teacher Knowledge Distillation with Nested Embeddings
OCTOPUS: Multi‑Agentic Universal Compositional Visual Retrieval
OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents
Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations
DiffPro: Joint Timestep and Layer-Wise Precision Optimization for Efficient Diffusion Inference
STEP: Spatial Thinking and Egocentric Pointing for Embodied Instruction Following
Beyond Dense Futures: World Models as Structured Planners for Robotic Manipulation
MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes
HIDA: A Human-Intuition-Guided Depth-Aware Framework for Zero-Shot Amodal Segmentation
MixTTA: Low-Rank Cross-Channel Mixing for Reliable Test-Time Adaptation
Instant Expressive Gaussian Head Avatars at Over 100 FPS
DecoupleGS: Interactive 3D Gaussian Splatting for End-to-End Autonomous Driving Testing
Grounding World Simulation Models in a Real-World Metropolis
MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
Breaking Rigidity in Adversarial Patch Attacks
AutoSpeed: Annotation-Free Stage-Adaptive Motion Speed Learning for Robot Manipulation
Spatial Amsan: A Benchmark for Perception-Grounded Spatial Reasoning and Action Evaluation in Egocentric Manipulation
TRiGS: Temporal Rigid-Body Motion for Scalable 4D Gaussian Splatting
UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer
TriNLOS: Triplane Representations for Neural Non-Line-of-Sight Imaging
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
TPCNet: A Low-Light Image Enhancement Network Inspired by Triple Physical Constraints
VGGT-World: Transforming VGGT into an Autoregressive Geometry World Model
Composing Driving Worlds through Disentangled Control for Adversarial Scenario Generation
Going Deep: Deep Visual Prompting with LoTeP
Mind-to-Face: Neural-Driven Photorealistic Avatar Synthesis via EEG Decoding
Learning Structurally Consistent Representations for Multi-View Radar Semantic Segmentation
Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching
QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding
CellFluxRL: Biologically-Constrained Virtual Cell Modeling via Reinforcement Learning
DIVER: Disentangling Camera–Object and Active–Passive Motion for Video Generation
MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
ARMS: Anchor–Relational Motion Streaming for Seamless Solo-Social Motion Transitions
Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation
CountEx: Fine-Grained Counting via Exemplars and Exclusion
MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein
Infinite Gaze Generation for Videos with Autoregressive Diffusion
Relaxed Rigidity with Ray-based Grouping for Dynamic Gaussian Splatting
OmniRen: Neural Rendering wih Heterogeneous Scene Primitives
Semantic Browsing: Controllable Diversity for Image Generation
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
Test-Time Registers as Global Priors for Tokenized Image Generation
MemPose: Category-level Object Pose Estimation with Memory
FedLAS: Feature-Modulated Bidirectional Label Smoothing for Neural Network Calibration
CAM3R: Camera-Agnostic Model for 3D Reconstruction
VoxAnchor: Explicit Voxel-Semantic Grounding for Spatial Understanding in Videos
On Locality and Length-Generalization in Visual Reasoning
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
Trajectory-aware Cross-view Geo-Localization with Sequential Observations
Semantically Aligned Gradient-Driven Context-Preserving Image Editing
FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Arbitrary Images
HASSL: Hierarchy-Aware Self-Supervised Learning Framework for Single Cell Microscopy
WorldAgents: Can Foundation Image Models be Agents for 3D World Models?
TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On
FPicker: Topology-Guided Evolution for Filament Tracing in Low-SNR Microscopy
DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models
Progressively Spiral Mamba Fusion for Multimodal Tracking
LoGAN: Multilingual Font Localization with Generative Agents
Lessons and Open Questions from a Unified Study of Camera-Trap Species Recognition Over Time
SLAIR: Structured Latent Flow Matching for All-in-One Image Restoration
SkySplat-OV: Generalizable Language Gaussian Splatting for Open-Vocabulary Scene Understanding from Sparse Satellite Views
EvoVLA: Self-Evolving Vision-Language-Action Model
Make Geometry Matter for Spatial Reasoning
Rethinking Detection Calibration: A Coordinate Perspective
Invisible Shortcuts: Why Vision Encoders Know Your Camera
Motion-aware Sparse Pipeline for Lightweight Object Tracking
Active View Selection with Perturbed Gaussian Ensemble for Tomographic Reconstruction
A First Exploration of Neuromorphic OT-CFM for Multi-Speaker VSR
Solving Diffusion Inverse Problems with Restart Posterior Sampling
When Token Compression Breaks: Structural Pruning vs. Token Reduction for Robust ViT Segmentation under High Compression
LiteGS: a high-performance framework to train 3dgs in subminutes via system and algorithm codesign
SemGAN: A Semantic and Hierarchical Adversarial Network for 3D Human Pose Estimation
Interact3D: Compositional 3D Generation of Interactive Objects
WeatherReasonSeg: A Benchmark for Weather-Aware Reasoning Segmentation in Visual Language Models
GeoDetect: Geometric Adversarial Detection for VLPs
Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation
3DGS3: Joint Super Sampling and Frame Interpolation for Real-Time Large-Scale 3DGS Rendering
Learning Probabilistic Prompt for Continual Learning
IMMoE: Incomplete Multi-View Anomaly Detection via Mixture of View Experts Fusion
GAP-Track: Bridging the Resolution Gap for Cross-Resolution RGBT Tracking
Resolution-Agnostic Neural Operators for Multi-Rate Sparse-View CT
Unsupervised Semantic Segmentation Facilitates Model Understanding
Denoising-GS: Gaussian Splatting with Spatial-aware Denoising
Why Do Vision Language Models Struggle To Recognize Human Emotions?
Markov-Renewal Single-Photon LiDAR Simulator
WildWorld: A Large-Scale Dataset for Action-Conditioned World Modeling with Explicit State Annotations
Towards High-Resolution Visual Perception via Hierarchical Entity Exploration
MapDreamer: Aerial Imagery Conditioned Latent Diffusion For Lane Level Map Generation
One Demonstration Is Enough for Real-World Robotic Reinforcement Learning
InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars
HiReFF: High-Resolution Feedforward Human Reconstruction from Uncalibrated Sparse-View Video
PIC: Revisiting INR for Image Coding with Fast Encoding and Sub-Millisecond Decoding
Geometry-Propagated Gaussian Splatting for Aerial Sparse Novel View Synthesis
TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs
EditHF-1M: A Million-Scale Rich Human Preference Feedback for Image Editing
Seeing Fast and Slow: Learning the Flow of Time in Videos
Unveiling Transferability in Trajectory Prediction via Latent Scene Embeddings
Be Tangential to Manifold: Discovering Riemannian Metric for Diffusion Models
3D Gaussian Splatting Compression with Object Scalability
A second-order theory of texture for depth from focus
AracNet: Revealing Debiasing Signals across Layers with Shallow Monitors
Auto3R: Automated 3D Reconstruction and Scanning via Data-driven Uncertainty Quantification
Axolotl3D: a Unified Framework for Faithful 3D Shape Completion
Beyond Alignment: A Generative Matching Paradigm via Flow Matching for Zero-Shot Skeleton-Based Action Recognition
C2E: Boosting Ego-Only 3D Object Detection via Multi-Teacher Contrastive Knowledge Distillation
D-VLAM: Differential Vision and Language Mixing for Rehearsal Free Continual Learning
Efficient and High-Quality Depth Estimation via Pixel-Space Diffusion with Linear Attention
Fisheye3R: Adapting Unified 3D Feed-Forward Foundation Models to Fisheye Lenses
HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving
HairWeaver: Few-Shot Photorealistic Hair Motion Synthesis with Sim-to-Real Guided Video Diffusion
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning
Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting
Kinematics-Agnostic 3D Human Motion Prediction via Equivariant Latent Diffusion
Latent Visual Diffusion Reasoning with Monte Carlo Tree Search
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
Multi-THuMBS: Multi-person Tracking of 3D Human Meshes Beyond Video Shots
NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language Models
OmniPoser: Flexible Human Motion Recovery in the Wild with Masked Flow Matching
RT-SDGOD: Real-Time Single-Domain Generalized Object Detection
RUTaL: Residual Upcycling with Task Ladder for Efficient Multi-Task Learning
SAFE-EQA: Semantic-Aware Efficient Exploration for Embodied Question Answering
SiPhy: Single-Image Physical Property Reasoning
Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models
SpecEyes: Accelerating Agentic Multimodal LLM via Speculative Planning and Perception
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
Stochastic Optimal Control Sampling for Diffusion Inverse Problems
StreamGVE: Training-Free Video Editing via Few-Step Streaming Video Generation
TreeSRNF: Square-Root Normal Fields for Generative Modelling of the Geometric and Structural Variability in Tree-like 3D Objects
TriMotion: Modality-Agnostic Camera Control for Video Generation
WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation
Activation Quantization of Vision Encoders Needs Prefixing Registers
Asymmetric Anchoring: Opening the Black Box of MLLMs for Forgery Detection
Break Visual-Linguistic Asymmetry: Unleashing VLM's Cross-Modal Potential for General Face Forgery Detection
Capacity Overflow: A Blind Spot for Backdoor Attacks in Vision MoE
CerDETR: Cell-Prior Empowered DETR for Cervical Lesion Detection
DnA: Denoising Attention for Visual Tasks
Dynamic World Generation Made Efficient
ECoSim: Data Efficient Fine-Tuning for Controllable Traffic Simulation
Edit3r: Instant 3D Scene Editing from Sparse Unposed Images
Ego-Human Motion Prediction with 3D-Aware LLM
EMOTE: Expressive Motion and Shape Disentanglement for Human Animation
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance
Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition
Grounding Sim-to-Real Generalization in Dexterous Manipulation: An Empirical Study with Vision-Language-Action Models
InterCMDM: Block-Causal Diffusion for Autoregressive Human Interaction Generation
IP-SAM: Rethinking Prompt-Conditioned Segmentation for Prompt-Absent Deployment
Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention
NEOMAP: Novel-View Synthesis via Noise Initialization by Manifold Alternating Projection
Online Versatile Incremental Learning: Towards Class and Domain-Agnostic Adaptation at Any Time
PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
Rank-Aware Hyperbolic Alignment for Vision–Language Dataset Distillation
SFM: Taming State Space Models for Text-to-Motion via Spatial-Frequency Modeling
SkipGS: Post-Densification Backward Skipping for Efficient 3DGS Training
StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description
Stylized Video Generation via Decoupled Data Synthesis and Gated Style Token Injection
Tactile Modality Fusion for Vision-Language-Action Models
Tiled Prompts: Overcoming Prompt Misguidance in Image and Video Super-Resolution
TooBad: Backdoor Diffusion Models with Ultra-Low Poison Rate and Imperceptible Trigger
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
VKSR: Scalable Kernel Surface Reconstruction Using Vecchia's Approximation
What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
WildSplat: Feedforward Gaussian Splatting from Unposed In-the-Wild Images
WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild
4DGS360: 360° Gaussian Reconstruction of Dynamic Objects from a Single Video
Aggregating Cross-Domain Knowledge via Learnable Tokens for Multi-Teacher Distillation
Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming
BrepLLM: Enabling Large Language Models to Understand Boundary Representations
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
Context Blindness in DPO: Mitigating Object Hallucination in MLLMs via Context-Calibrated Preference Optimization
Context-Interactive Reasoning for Group Activity Detection
Delineating Knowledge Boundaries for Honest Large Vision-Language Models
GaussianLens: Localized High-Resolution Reconstruction via On-Demand Gaussian Densification
ChronusOmni: Improving Time Awareness of Omni-Modal Large Language Models
GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation
Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language Models
HitMem: Hierarchical Temporal 3D Memory with Multi-Modal Context-Aware Retrieval for Dynamic Environments
Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures
Learning to Attract and Repel: Dual Quality Margin Learning for Face Recognition (DQM-Face)
MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer
MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation
NoisEasier: Test-Time Noise Optimization for Text-to-Video Generation
OccDirector: Language-Guided Behavior and Interaction Generation in 4D Occupancy Space
One4D: Unified 4D Generation and Reconstruction via Decoupled LoRA Control
Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
Pix2NPHM: Learning to Regress NPHM Reconstructions From a Single Image
Pose Anything Anywhere: Model-free Object Poses from Arbitrary References
RADmesh: Remesh-Aware Mesh Deformation
RAGrasp: A Retrieval-Augmented Framework with Diversity-Aware Modeling for Dexterous Grasp Generation
RawGen: Learning Camera Raw Image Generation
Reconstructing Humans and Objects in Interaction using Large Reconstruction Models
Roam2Room: A Unified Floorplan-to-Furnished Framework for Controllable Indoor Scene Generation
Robust Trajectory Distillation: Hybrid Reweighting Meets Teacher-Inspired Targets
Scaling Dense Prediction with Latent Decoding
SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination
SCoT: Similarity-guided Conflict-aware Task Consolidation for Continual VQA
See & Sniff: Learning Visuo-Olfactory Representations
Stabilizing Ultra-Low-Bit Quantization of Multimodal LLMs via Global Bit Allocation
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning
The Label Imitation Game: Turing Test Network for Zero-Shot Pseudo-Label Pruning
WildDepth: A Multimodal Dataset for 3D Wildlife Perception and Depth Estimation
Abstract the Layout, Focus the Detail: A Dual-Granularity Representation Framework for Zero-Shot 3D Visual Grounding
AMCoNav: Asynchronous Multi-module Collaborative Framework for Embodied Visual Navigation
AquaStereo: Enabling Underwater Stereo Matching via Depth-Conditioned Diffusion and Geometry Self-Distillation
Articulat3D: Reconstructing Articulated Digital Twins From Monocular Videos with Geometric and Motion Constraints
CameraAnything: Refilming Videos with Arbitrary Camera Control
Color Pass-Through via Camera-Display Coupling
ConceptWeaver: Weaving Disentangled Concepts with Flow
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-visual Language Models
Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment
DriftScope: Measuring The Hidden Effects of Diffusion Model Fine-Tuning
E3VS-Bench: A Benchmark for Viewpoint-Dependent Active Perception in 3D Gaussian Splatting Scenes
Estimating Individual Tree Height and Species from UAV Imagery
FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors
FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation
GuideMe: Benchmarking Multi-Domain Task Guidance and Intervention in Streaming Video
InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
Leveraging Dark Knowledge for Intrinsic Multimodal Out-of-Distribution Detection
Leveraging Phase Information to Boost Unrolled Network Learning for Image Deblurring
MPO: Single-Stream Policy Optimization for Efficient Text-to-Image Alignment
On-Orbit Real-Time Wildfire Detection Under On-Board Constraints
OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization
OrthoTrack: Continuous 6-DoF UAV Trajectory Estimation Anchored in Public Orthophotos
PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided VLM
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
PuzLM: Solving Jigsaw Puzzles with Sequence-to-Sequence Language Models
Reconstructing 3D Human-Object Interaction via a Unified Triplane Space
Decoding Multimodal Causality: End-to-End Multimodal Mediation Pathways Inference
Segmentation-Guided Homography Estimation for Long-Term Planar Tracking
Simple Filtering Improves Masked Autoencoders
Stealthy Multi-task Adversarial Attacks
Structure Gaussian Splatting SLAM
Temporal and Cross-modal Alignment for Enhanced Audiovisual Video Captioning
Why Feature Magnitude Deceives OOD Detectors: An Angular Separation Perspective
Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods
AVSR-Diff: Scale-Agnostic Diffusion Priors for Temporally Consistent Arbitrary-Scale Video Super-Resolution
BEVOpen3D: Towards Open-World 3D Object Detection in Bird's-Eye-View
Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
Category-Level Articulated Object Pose Estimation via Pose–Shape Hypothesis Generation and Verification
CloSeR: Unified Relational Distillation from Closed-Set Teachers for Category Discovery
DeCoFlow: Structural Decomposition of Normalizing Flows for Continual Anomaly Detection
Dynamic Inverse Rendering for Enhanced Material-Lighting Decomposition
Every Dog Has Its Day, Probably: A Balanced Synthetic Benchmark and Probabilistic Modeling for 3D Dog Pose Estimation
Following Motion for Sequential Modeling in Video Frame Interpolation
FreeSwim: Revisiting Sliding-Window Attention Mechanisms for Training-Free Ultra-High-Resolution Video Generation
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
Improving Knowledge Distillation Under Unknown Covariate Shift Through Confidence-Guided Data Augmentation
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
Novel View Synthesis as Video Completion
Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints
Pathwise Test-Time Correction for Autoregressive Long Video Generation
PDF-Omni: Poincaré Dual Disk Distortion Field-based Recurrent Update for Omnidirectional Stereo Matching
Perceiving Better Moments: Cover Frame Reselection and Enhancement for Live Photos with the Live2K Dataset
PersonaDrive: Controllable Trajectory Prediction with Multi-Dimensional Driving Personas
Ranked Activation Shift for Post-hoc Out-of-Distribution Detection
SeeClear: Reliable Transparent Object Depth Estimation via Generative Opacification
SGQA: Semantic-Geometric Quality Alignment for Training-Free Few-Shot Instance Segmentation
SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks
SpecV: Specification Verification for Robust Unified Multimodal Evaluation
Structured Hyperedge Adaptation for Parameter-Efficient Fine-Tuning of Vision Transformers
Video Generation Models Are Inherent Lighting Estimators
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
When the Teacher Has More Bits: Self-Teacher Latent Distillation for Learned Image Compression
X-Stream: Benchmarking MLLMs as Multiplexers for Multi-Stream Understanding
Adapting MLLMs for Nuanced Video Retrieval
Adaptive Spectrum-Aware Feature Disentangled Network for Small Object Detection
AnchorGUI: Asymmetric Memory for Dual-Scale Learning in GUI Navigation
CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models
EgoExoMoCap: Distributed Human Motion Capture via Ego- and Exocentric Body Tracking from Head-Mounted Devices
Beyond Filter Pruning: Top-K Spatial Selection for Efficient Neural Networks
FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction
Flash-Refine: Frustum-Guided Local Incremental Learning for Efficient 3D Gaussian Splatting Completion
G3AFT: Glance Guided Gradient Aligned Fine-Tuning for Visual Autoregressive Models
GroundSet: A Cadastral-Grounded Dataset for Spatial Understanding with Vector Data
Hierarchical Style Aggregation for Versatile Chinese Handwriting Generation
HippoCamp: Benchmarking Contextual Agents on Personal Computers
MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
Mind2Cloud: EEG-to-Point Cloud Generation with Two-Granularity Diffusion Decoding
Moving Beyond More Views: Redundancy-Aware Ego–Exo Fusion for Proficiency Estimation
OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder
Pano3D: Unified 3D Reconstruction and Panoptic Segmentation
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
Physically Grounded Dual-Opacity Gaussian Splatting for Joint RGB-TIR Reconstruction
PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models
RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution
ReinDriveGen: Reinforcement Post-Training for Out-of-Distribution Driving Scene Generation
ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision–Language Models
Segmenting Visuals With Querying Words: Language Anchors For Semi-Supervised Image Segmentation
Short-to-Long Functional Connectivity Transfer via Structure-Aware Latent Diffusion
Solving Semi-Supervised Few-Shot Learning from an Auto-Annotation Perspective
Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution
TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation
Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
VSDiffusion: Taming Ill-Posed Shadow Generation via Visibility-Constrained Diffusion
X2SAM: Any Segmentation in Images and Videos
Rolling Shutter Camera Self-Calibration
ZTRS: Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer
AgentVLN: Towards Agentic Vision-and-Language Navigation
AiSCREAM: Absolute Target Localization with Language-Conditioned Cross-View Alignment for Autonomous Vehicles
Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection
AVSplat:Dense-View Feed-Forward 3D Gaussian Splatting with Assist-View Preconditioning
ComplexMimic: Human–Scene Interaction Imitation in Complex 3D Environments
CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization
Beyond Time Shifts: Adapting Omni-LLM as a Reference-Free Evaluator for Generative Audio-Visual Models
Correlation-Weighted Multi-Reward Optimization for Compositional Generation
CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning
Director: Instance-aware Gaussian Splatting for Dynamic Scene Modeling and Understanding
ECHO: Ego-centric Modeling of Human-Object Interactions
Efficient Document Tampering Localization with Multi-Level Discrepancy Features and Unified DCT–Quantization Embedding
Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models
A Scalable Vector Graphics Latent Space
Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation
Geometry-Aware Visual Representation for Remaining Useful Life Prediction
GlassGS: Geometry and Concept-Aware 3D Gaussian Splatting for Reflective Enclosures
HNDiff: Haze-Noise Diffusion for Image Dehazing
Human-like Object Grouping in Self-supervised Vision Transformers
InnoText: A Unified Model for Visual Text Generation and Editing
Layer-Aware Video Composition via Split-then-Merge
Multi-modality Image Fusion under Adverse Weather: Mask-Guided Feature Restoration and Interaction
NURBS Splatting: A Unified Differentiable Rendering Framework for Vector Graphics
PriSM: Parsing and Style-Mixed Consistency for Unsupervised Domain Adaptation in Facial Landmark Detection
OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer
Path-JEPA: Path Signature Based Predictive Learning for Skeleton Action Recognition
UniScale: Arbitrary-Scale Anomaly Generation
PhenoLIP: Phenotype Guided Medical Vision–Language Pretraining
Pixel Ignores, Superpixel Sees: Adverse Weather Image Restoration via Semantic-Center SSM
PolarAPP: Beyond Polarization Demosaicking for Polarimetric Applications
Prototype Normalization: Optimizing Prototype Separation for Heterogeneous Federated Learning
RAGA: Real Time Ray Traced Gaussian Shadow Casting for 3DGS Avatar-Scene Interaction
REDistill: Robust Estimator Distillation for Balancing Robustness and Efficiency
SD3.5-Flash: Distribution-Guided Distillation of Generative Flows
SHINE-PPG: Non-Lambertian Intrinsic Decomposition for Illumination-Robust rPPG
Structural Assessment for Understanding and Guiding Dataset Distillation in Discrete Token Space
Text-Conditioned Background Generation for Editable Multi-Layer Documents
Thermo-JEPA: Learning a Geometry-Grounded Thermal World Model via Cross-Modal Privileged Masking
TMI: Text-to-Image Meets Image-to-Image for Complementary Data Synthesis to Boost Long-Tailed Instance Segmentation
VIGA: View-Conditioned and Identity-Guided Adaptation for Aerial-Ground Person Re-Identification
Achieving Subcategorical Erasure in Text-to-Image Models
DART: Deformable Adaptive Reasoning with Temporal Queries for Online Skeleton-Based Action Recognition
Learning from Adversity: Semantic-Aware Mask Refinement through Adversarial Perturbation
Freqformer: Image-Demoiréing Transformer via Effective Frequency Decomposition
DDStereo: Efficient Dual Decoder Transformers for Stereo 3D Road Anomaly Detection
Decoupling Complexity from Scale in Latent Diffusion Model
DETRAM: End-to-end DEtection, Tracking and Recovery of HumAn Meshes
Diffusion Model as a Generalized Segmentation Learner
Enhancing Embodied Reasoning and Grounding by Novel View Synthesis
Fast and Compact 3D Gaussian Splatting with Polarized Opacity Prior
G2P: Gaussian-to-Point Attribute Alignment for Boundary-Aware 3D Segmentation
HAS: Highlight-guided Attention Steering for Multimodal LLM Video Summarization
LiFlow: Flow Matching for 3D LiDAR Scene Completion
LiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving
MASS: Motion-Aligned Selective Scan for Flow-Based Video Frame Interpolation
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
Momentum Guidance: Plug-and-Play Guidance for Flow Models
Neuromorphic X-ray Computed Tomography
Online Segment 3D Gaussians via Launching Virtual Drones
ParaFlow: Parallel Sampling for Flow Matching Models
Point Diffusion Mamba: Unified Diffusion-State-Space Modeling for Single-View 3D Reconstruction under Data Scarcity
Push–Pull Attentional Anchoring for Diffusion Concept Erasure
Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation
REFINE: Super-efficient Pruning for 3D Gaussian Splatting via Rendering-Free Primitive Importance
RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction
RoMan-4D: Learning Robot Arm Manipulation from 4D World Models
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
Small Vision-Language Models are Smart Compressors for Long Video Understanding
SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning
StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics
Straight-Path Flow Matching for Incomplete Multi-View Clustering
Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors
SyntheticDoc: A Large Synthetic Dataset for Document Unwarping and Illumination Correction
Unsafe by Reciprocity: How Generation–Understanding Coupling Undermines Safety in Unified Multimodal Models
VesselTok: Tokenizing Vessel-like 3D Biomedical Graph Representations for Reconstruction and Generation
ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning
ViTAL‑X: Video-Text Alignment with Cross‑Modal Temporal Edits
VoCa: Unified Autoregressive Modeling for Talking Audio-Video Generation
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation
Minute4D: Training High-Fidelity 4D Gaussian Splatting in One Minute
2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction
A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
Are Video Reasoning Models Ready to Go Outside?
ARVAR: Accelerating Visual Autoregressive Model via Attention Retrospect
Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench
Cast and Attached Shadow Detection via Iterative Light and Geometry Reasoning
Category-Level 3D Correspondence in Camera Space via Morphable Object Priors
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
Denoising-Enhanced Coarse-to-Fine Infrared Small Target Detection with Attention Prior-Guided Knowledge Distillation
On Geometric Understanding and Learned Priors in Feed-forward 3D Reconstruction Models
DG-Force: Disentangling and Gathering Forensic Cues is Needed for Image Manipulation Localization
DiffuPrompt: Adapting Video Foundation Models to 3D Medical Volumes via Latent Trajectory Priors
NeuralDMD: Interpretable Untrained Neural Network for Imaging from Sparse and Noisy Observations
DINO-SLAM: DINO-Informed RGB-D SLAM for Neural Implicit and Explicit Representations
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
Hierarchical and Holistic Open-Vocabulary Functional 3D Scene Graphs for Indoor Spaces
InstanceControl: Controllable Complex Image Generation without Instance Labeling
LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection
Learning Generatable Mutual Distance for Scene-Aware Human Motion Generation
LumiDepth: Stable Monocular Depth in Multi-Illumination Scenes
MiNQVIS: Mitigating Noisy Queries for Robust Online Video Instance Segmentation
mmIR: Frequency-Space Inverse Rendering for 3D Millimeter-Wave Radar ADC Synthesis
Multi-view Multi-vehicle Driving Dataset for Novel View Synthesis
No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs
NumColor: Precise Numeric Color Control in Text-to-Image Generation
On the Diffusibility of High-Dimensional Latents
Perceptual Projection Pruning: Diversity-Aware Video Token Pruning for Multimodal Large Language Models
PhysDrape: Learning Explicit Forces and Collision Constraints for Physically Realistic Garment Draping
PhysRVG: Physics-Aware Unified Reinforcement Learning for Video Generative Models
Robust 3DGS-based SLAM via Adaptive Kernel Smoothing
RoME: Robust Mixture of Low-Rank Experts against Multiple Adversarial Perturbations
SAFE-Pruner: Semantic Attention–Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
SE-DETR: Explicit Semantic Exploration for Generalizability and Distinguishability in Video Temporal Grounding
Sparse auto-regressive modeling for scene generation from multi-view images
StructPolicy: Structure-Guided Imitation Learning Robust to Visual Domain Shifts
SynFlow: Scaling Up LiDAR Scene Flow Estimation with Synthetic Data
TopoGS: Planar Reconstruction via Topology-Aware 3D Gaussian Splatting
WaterGen: Decoupling Scene and Medium in Underwater Image Generation
When Higher Order Hurts: Pre-Asymptotic Order Collapse in Generative ODE Sampling — A Theory of Discretization–Learning Interaction
JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation
When Sinks Help or Hurt: Unified Framework for Attention Sink in MLLMs
Efficient Quantization-Aware Adaptation for Visual Foundation Models
3D-Aware VLMs with Implicit and Explicit Geometries
AnchorSplat: Fast and Structure Consistent Detail Synthesis for Gaussian Splatting
CFM: Language-aligned Concept Foundation Model for Vision
DeCoPatch: Revealing Causal Latent Subspaces in Vision-Language Models for GUI Grounding
Consistent Monocular Depth Estimation with Contact Region Boundary-Aware Refinement
Boosting Correspondence Learning with Structure-Aware Estimator
CRD-Net: Frequency-Adaptive Feature Injection and Change Decoupling for Building Damage Assessment
D3F-IR: Dual-Domain Deterministic Flow Matching for Visible-to-Infrared Translation
DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues
FlowCIR: Semantic Transport via Flow Matching for Zero-Shot Composed Image Retrieval
ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance
GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild Videos
GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks
GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens
Hybrid Event–Frame Sensors: Modeling, Calibration, and Simulation
LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion
LGD-Net: Leader-Guided Cross-Modal Dynamics for Hyperspectral and Panchromatic Image Fusion
LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models
MANGO: Unleashing Image Generation Capability of Unified Multimodal Models
Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
On Test-Time Scaling for Vision-Language Models
OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization
ORBIT: Overcoming Hallucination Risks via Bi-manifold Interaction and Traction
PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models
PMGC-SimVP: Parametric Multi-scale Gated Convolution for Global Ionospheric TEC Prediction
PrimitiveUDF: Primitive-Based Unsigned Distance Fields for Surface Reconstruction from Point Clouds
PyraE2E: Enhancing End-to-End WSI Analysis via Cross-Scale Super-Resolution
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
Robust and Efficient Monocular 3D Gaussian SLAM for Kilometer-Scale Outdoor Scenes
Triangular Consistency as a Universal Constraint for Learning Optical Flow
See Only When Needed: Context-Aware Attention Intervention for Hallucination-Free LVLMs
SignRefine: Adapting Foundational Video Models for Sign Language Generation
SMG: Semantic Motion Graph for Monocular Dynamic Gaussian Splatting
SurvMILKD: A Weakly Supervised Survival Analysis Framework for Multi-Teacher Knowledge Distillation using Pathology Foundation Models
WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion
TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
TinyHistory: Lightweight Video History Embeddings via Two-Stage Context Learning
TR-MoE: Temporal Reliability-Aware Mixture-of-Experts for Robust Tracking
UC-VLM: Consistency-Driven Learning for AI-Generated Image Detection with Vision-Language Large Models
World Knowledge in the Weights: Reading Concept Circuits of Vision Transformers
XSemanticFlow: Cross Object Semantic Alignment for Zero-shot Manipulation
CaPCL: Caption-Preserved Continual Learning for Text-to-Image Retrieval
CMDer: Controllable Mode Decomposition-Based Single Motion Synthesis with Diffusion
CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation
CSS-BA: Gate Guided Column Space Search for Bundle Adjustment
D-Rex : Diffusion Rendering for Relightable Expressive Avatars
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
Differentiable Polarized Path Tracing
DIGS: Differentiable, Incremental, Global, Scalable Pruning for Language Models
Disentangling Rotation and Translation from SE(3)-Equivariant Features for Shape Assembly
Don’t Teach Instability, Teach Robustness: Selective Sensitivity Gating for Adversarial Robust Distillation
AVQ-Attention: Adaptive Vector-Quantized Attention
EGVLR: Evidence-Grounded Vision–Language Reinforcement for Anomaly Reasoning
Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection
FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation
FrozenDrive: Zero-Shot Text-Guided Driving Scene Generation and Data Augmentation with Parameter-Free Frozen Diffusion Model
GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM
Hi-DiT: Hybrid Latent-Pixel Diffusion Transformer for Image Generation
Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation
Learning to Deny: Action Denial in Multimodal Large Language Models
LIIFusion: Coarse-to-fine Framework for Generative MEF via Implicit Neural Representation
MotionChain: Fine-Grained Video Motion Understanding via Structured Decomposition
NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control
NeuralGarSim: Geometry-agnostic Garment Simulation with Neural Fields
Optimizing Mesh Animation from Video via Shape Flow Guidance
ORFC: Orthogonal Reparameterization for Low-Bitrate ViT Feature Coding
P3-SAM: Native 3D Part Segmentation
PADFormer: Pose-agnostic Anomaly Detection from Sparse View Images
PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
PhysPO: Physics-Aware Local Preference Optimization for Physically Consistent Video Diffusion
Progressive Pose-Guided 4D Animal Reconstruction from Monocular Video
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
PS-MOT: Cultivating Instance Awareness from Point Seeds for Multi-Object Tracking
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
SGMatch: Semantic-Guided Non-Rigid Shape Matching with Flow Regularization
SharpGS: Sharpness-Preserving 3D Gaussian Splatting with Differentiable Blur-Driven Density Control
SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning
Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility
SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model
Test-time Counterfactual Calibration for Hallucination-Resistant Temporal Grounding
LensStyle: Learning the Optical Aesthetics for Controllable Stylized Lens Effect Rendering
Thinking from the Robot’s View: The CoT-HRC Benchmark for Human Intent Reasoning in Embodied Collaboration
Towards In-Context Tone Style Transfer with A Large-Scale Triplet Dataset
Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels
TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction
Guide, Think, Act: Interactive Embodied Reasoning for Vision-Language-Action Model
URoPE: Universal Relative Position Embedding across Geometric Spaces
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling
What Matters in RL-Based Methods for Object-Goal Navigation? An Empirical Study and A Unified Framework
A Benchmark for Heterogeneous Stereo Deblurring with Physically- and Epipolar-constrained Cross Attention
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
AnyMatch: Supercharging Universal Multi-Modal Image Matching with Large-Scale Single-View Images
ARC-Loc: Leveraging Azimuthal Ray Convergence as a Geometric Cue for Direct Cross-View Localization
Beyond Atomic Layouts: Compositional Design Understanding with Vision-Language Models
Boba: Batched Simulation for Physics-Based Gaussian Digital Twins
Bridging Theory and Practice in Source-Free Domain Adaptation via Adversarial Proxy Perturbation
CasaMaestro: Multi-View Panoramas for House-Scale 3D Reconstruction
ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement
Clue Matters: Empower Video Reasoning with Brain-Inspired Latent Clue Learning
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection
Cube-Splat: High-Fidelity 360° Gaussian Splatting SLAM via Cubemap Factorization and Adjoint-Consistent Optimization
Discrete Diffusion Bridges for Spatiotemporally Aligned Image Translation and Generation
DLGStream: Dynamic Language-embedded Guassian Splatting for Open-vocabulary Enabled Free-viewpoint Video Streaming
DuoFlow: JVP-Free Finite-Difference Mean Flows for One-Step Image Generation
EmoteGPT: 3D Human Facial Expression from Natural Language Descriptions
ESCAPE: Episodic Spatial Memory and Adaptive Execution Policy for Long-Horizon Mobile Manipulation
Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs
FairSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation
FastSTAR: Spatiotemporal Token Pruning for Efficient Autoregressive Video Synthesis
Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
GradingBench: Evaluating End-to-End Compositional Reasoning of MLLMs for Automated Exam Grading
Learn to See the Unseen in Low-light Spike Streams
Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents
Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs
NaVLM-PVC: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
Neural Gate: Mitigating Privacy Risks in LVLMs via Neuron-Level Gradient Gating
NUN: Nested Unfolding Network for Real-World Concealed Object Segmentation
RAF: Reliability-Aware Fusion of Camera, LiDAR, and 4D RADAR for Robust 3D Object Detection in Adverse Weather
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
ReFP-AD: Rectified Flow Preconditioning for Energy-Based Anomaly Detection
GR-GRPO: Graph-Diffused Credit for Autoregressive Image RL Alignment
Robust Self-Supervised Cross-Modal Super-Resolution against Real-World Misaligned Observations
Seeing What Matters: Lesion-Aware High-Resolution Patch Discovery and Fusion for Chest X-ray Report Generation
Semantic Generative Tuning for Unified Multimodal Models
Natural Image Pretraining Improves Abstract Reasoning
Target-aware Image Editing via Cycle-consistent Constraints
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
STEP: Score-Based Temporal Energy for Human Pose Video Anomaly Detection
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
Taming Dynamic Clutter: Variance-Driven Adaptive Gain Control for Bio-inspired Small Target Detection
Triangle Splatting SLAM
UltraGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention
UniDynamics: Event-RGB Fusion for Unified Future 4D Dynamic Scene Generation
WAFT-Stereo: Warping-Alone Field Transforms for Stereo Matching
Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence
AGE: Agentic Gaussian Editing in 3D Scenarios
AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
Bayesian Uncertainty Attribution-Guided Fine-Tuning for Open-Set Action Recognition
BEV-GS: Feed-forward Gaussian Splatting in Bird’s-Eye-View for Road Reconstruction
Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives
Calibrated Harmonic Overlaid Implicit Neural Representations for Multi-Dimensional Data
CLIMP: Contrastive Language-Image Mamba Pretraining
Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging
CogSENet: Blind Image Deblurring with Blur-Conditioned Semantic Routing and Explicit Frequency Fusion
ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering
DAPS++: Rethinking Diffusion Inverse Problems with Decoupled Posterior Annealing
Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift
DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation
DreamLite: A Lightweight On-Device Unified Model for Image Generation and Editing
Generative Relightable Avatars
Inference-time Motion Calibration for Video Generation
K-Mask: Kinematic-Aware Masked Modeling for Controllable Text-to-Motion Synthesis
LaMP: Learning Vision-Language-Action Policies with 3D Scene Flow as Latent Motion Prior
Learning to Generate Rigid Body Interactions with Video Diffusion Models
LESV:Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding
LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation
QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding
MArFE: Multi-Contrast MRI Arbitrary Scale Super-Resolution with Fourier Enhancement
MaterialFlow: Attribute-Disentangled Material Transfer via Trajectory-Aware Velocity Modulation
Objects as Audio-Visual Modal Sound Fields
Optimization-Guided Diffusion for Interactive Scene Generation
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
Posterior Augmented Flow Matching
Posterior Samplings are Missing Modalities Generators for Medical Image Translation
Proteus: Model Leakage-Induced Adversarial Attack in Federated Learning
RegRet: Enhancing Region-Level Retrieval in Large Multimodal Models
Rethinking Token Reduction for Diffusion Models via Output-Similarity-Awareness
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video
OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation
SAMPLe: A Sharpness Aware Minimization based Optimizer for Prompt Learning in Vision-Language Models
ScenarioControl: Vision-language Controllable Vectorized Latent Scenario Generation
SeekFlow: Synergizing Radiology and Pathology Foundation Models for Precision Oncology via Knowledge-Guided Evidence Flow
Self-Evolving Just-In-Time Memory for Proactive Embodied Safety
SMART: When is it Actually Worth Expanding a Speculative Tree?
Decoding Children’s Gait Behavior
Spectral Evolution-Guided Token Pruning in Large Multimodal Models
Stabilizing Deep Reconstruction Operators with Contractive Anchoring
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
Stitched Embeddings: A Unified Latent Space for 3D Garments and 2D Patterns
Free‑CD: Probabilistically Decoupled Training-Free Open-Vocabulary Change Detection with Resolution-Invariant Feature Inversion
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
Rethinking Continual Anomaly Detection on the Edge: Benchmarking Under Realistic Industrial Conditions
Tac2Real: Reliable and GPU Visuotactile Simulation for Online Reinforcement Learning and Zero-shot Real-World Deployment
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
Two-Parameter Flow Map Learning for Continuous-Time Diffeomorphic Image Registration
Under One Sun: Multi-Object Generative Perception of Materials and Illumination
Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
Benchmarking Vision-Language Models for Microscopic Plant Image Understanding
BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations
Beyond Artifacts: Real-Centric Envelope Modeling for Reliable AI-Generated Image Detection
Beyond Script Family Boundaries: Towards Unified Open-Set Scene Text Recognition
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction
CMCC-ReID: Cross-Modality Clothing-Change Person Re-Identification
Co-Steer: Cross-Modal Collaborative Steering for Jailbreaking MLLMs
Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
Controlling Motion Transfer in Diffusion Transformers via Attention Heads
DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
Dive into the implicit biases of low-rank vision-language alignment
Enhancing Pretrained Model-based Continual Representation Learning via Guided Random Projection
Event-driven Motion Deblurring via Trajectory-based Kernel Reconstruction
Fast Spatial Memory with Scalable Elastic Test-Time Training
Fidelity- and Perception-Aware Local Implicit Attention for Arbitrary-Scale Image Super-Resolution
DeltaDeno: Zero-Shot Anomaly Generation via Delta-Denoising Attribution
FlowInOne: Unifying Multimodal Generation as Image-in, Image-out Flow Matching
From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation
Hyper-Network Neural Functional Maps for Unsupervised Robust 3D Shape Matching
Improved Immiscible Diffusion: Accelerating Diffusion Training by Reducing Miscibility
Learning from Reliable Negatives: Confidence-Anchored Test-Time Adaptation for GUI Grounding
Per‑Object IoU Forecasting for Deadline‑Aware Real‑Time Embedded Detection Control
Learning on the Manifold: Unlocking Standard Diffusion Transformers with Representation Encoders
LogiCo: A Unified Framework for Logical and Structural Anomaly Detection
MemRoPE: Training-Free Infinite Video Generation via Evolving Memory Tokens
SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs
PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates
MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations
ODONet: Online Dynamic Offset Network for Visual Object Tracking
Moiré Video Authentication: A Physical Signature Against AI Video Generation
OSVE: One Step Video Editing with One Step Diffusion Models
PosterCopilot: Toward Layout Reasoning and Controllable Editing for Professional Graphic Design
PWM-ArtGen: Part World Model for Articulated Object Generation
Evidence Triangulation for Multimodal Fact-Checking in the Wild
R2M: Real-Aware Residual Model Merging for Robust and Generalizable Deepfake Detection
RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning
SafeSAE-VLA: Interpreting OpenVLA Progress Dynamics with Sparse Feature Analysis
Seeing Touch from Motion: A Unified Modality-Aware Visuo-Tactile Policy with Tactile Motion Correlation
The Language of Visual Attention: Modeling Scanpaths via Autoregressive Token Prediction
FuDU: A Fuzzy Dual-dimension Uncertainty Framework for Streaming Active Learning in Industrial Defect Detection
MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learning
Domain Adaptive Object Detection via Dual-Stream Bilevel-Cycle Optimization
TouchAnything: Diffusion-Guided 3D Reconstruction from Sparse Robot Touches
WAPR: A Foundation Model for Wide-Angle Refinement in Unseen Object Pose Estimation
Local Spacing-Aware Hungarian Matching for Stable Point-Supervised Crowd Counting
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment
VOID: Video Object and Interaction Deletion
Robust Zero-shot Anomaly Detection under Limited Auxiliary Anomaly Priors
Art Beyond Semantics: Sheaf-Informed Contrastive Learning for Multi-Relational Representations
Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset
BeyondSight: Object Permanence for End-to-End Autonomous Driving
CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes via Geometric Pseudo Labels
SR-Edit: Region-Aware Image Editing via Self-Refinement
Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers
CrossView: Can Vision-Language Models Reason Across Cameras?
Diagnosing Aerial-View Object Detectors with Foundational Image Generative Models
DiffVP:Differential Visual Semantic Prompting for LLM-Based CT Report Generation
DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
DocLayout-VL: A Foundational Model for Hierarchical, Open-set, and Promptable Document Layout Segmentation
Dual Masked Generative Adversarial Transformer for Unsupervised Domain Adaptation
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
MSEditor: Toward Consistent Multi-Shot Video Editing
EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis
EventSpecPS: Photometric Stereo with Multispectral Reflectance Using an Event Camera
Face Anything: 4D Face Reconstruction from Any Image Sequence
Forecasting Animal Motion
QualiTeacher: Quality-Conditioned Pseudo-Labeling for Real-World Image Restoration
Learning to Corrupt for Better Restoration
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
HER-Count: Learning Hyper-Exemplar Representation for Generalized Zero-Shot Object Counting
HSDF-Lane: Height-Aligned Signed Distance Field with Semantic Lane Prior for 3D Lane Detection
Hyperbolic Hierarchical Clustering for Visual Representation Learning
Improving Image-to-Image Translation via a Rectified Flow Reformulation
Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation
Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI
Mitigate Modality-Asymmetric Forgetting via Stabilizing Visual Representations in CLIP-Based Class-Incremental Learning
SkelEM: Explicit Decoupling of Topology and Details for Self-supervised Axial Super-Resolution in Volume Microscopy
Multi-History-Step SDE Inversion for Image Editing with Superior Regional Awareness
OneHSI: A Unified Hyperspectral Foundation Model with Physical Consistency
PHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark
Racing in Volume with Flow Ensembles
InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem
RankT2I: A Submodular Framework for Discovering Interpretable and Diverse Semantics in Text-to-Image Models
Rosetum3D: A Large-Scale 3D Vision Dataset from Preharvest Roses
ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding
Sector-Level Cross-View Geo-Localization with Implicit Orientation via Azimuthal Scanning
SignBind-LLM: Multi-Stage Modality Fusion for Sign Language Translation
SK-Adapter: Skeleton-Based Structural Control for Native 3D Generation
UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
Nexels: Neurally-Textured Surfels for Real-Time Novel View Synthesis with Sparse Primitives
Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis
UniPart: Towards Zero-shot Language-Grounded 3D Part Segmentation for Embodied Interaction
VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
SAND: Stage-Aware Noise Decomposition for Training-Free Diffusion Guidance
Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision
Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs
ContextFlow: In-Context Flow Matching for Robot Manipulation
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
DeGuNet: Depth-Guided Ultra-Compact Backbones for Efficient LiDAR-Camera 3D Detection
DeLux: Cross-Modal Local Artifact Restoration in Video Using Neuromorphic Data
Distribution Matching Distillation Meets Reinforcement Learning
DualDiff3D: Dual Structure-Appearance Diffusion Priors for Reliability-Enhanced 3D Gaussian Splatting
HSFM: Hard-Set-Guided Feature-Space Meta-Learning for Robust Classification under Spurious Correlations
D2PO: Optimizing Diffusion Samplers via Dynamic Preference
Spectral Prior for Reducing Exposure Bias in Diffusion Models
ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving
Fair and Faithful: A Diffusion-Enhanced Dataset and Hybrid State-Space Mamba for Face Super-Resolution
Fast and Scalable LiDAR Data Generation for Autonomous Driving Simulation without Raycasting
FILT3R: Latent State Adaptive Kalman Filter for Streaming 3D Reconstruction
Harnessing SSL for Segmentation in 3D Microscopy with Noisy Labels and Hard Patches
Towards Consistent and Efficient Dataset Distillation via Diffusion-Driven Selection
Honey, I Shrunk the Arc de Triomphe!
Learning to Balance: Decoupled Siamese Diffusion Transformer for Reference-Based Remote Sensing Image Super-Resolution
MATCH: Flow Matching for Multi-View Anomaly Detection
Exposure Bias Can Alleviate Itself via Directional and Frequency Rectification in Flow Matching
MCPNet:Masked Coordinate Pooling-based Attention Network for Medical Landmark Detection
MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
MuCHeR: Multi-Person Camera-Centric Human Detection, Mesh Recovery and Tracking
MultiHaystack: Benchmarking Multimodal Retrieval and Reasoning over 40K Images, Videos, and Documents
Sparse-Aware Vector Quantization for Bandwidth-Efficient Collaborative 3D Semantic Occupancy Prediction
O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
Panoramic Affordance Prediction
Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
Progressive and Localized Super-Resolution of 3D Objects via Localized Latent Voxel Diffusion
RADIANCE: Relative Adaptive Denoising with IP-Adapter for Novel Concept Enhancement
SAM-MT: Real-Time Interactive Multi-Target Video Segmentation
SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models
StereoEdit: A Diffusion-Based Framework for Stereo-Consistent Image Editing
Surprise Forcing: What to Remember, When to Skip in Long Video Generation
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
One-Step Flow Policy: Self-Distillation for Fast Visuomotor Policies
Towards Long-Form Spatio-Temporal Video Grounding
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
UHD-MFF: Shattering Barriers in Multi-Focus Ultra-High-Definition Image Fusion via Learnable Lookup Tables
Unified Multi-plane Autoregressive Diffusion for 3D Multi-Contrast MRI Synthesis
UniREditBench: A Unified Reasoning-based Image Editing Benchmark
Selective Synergistic Learning for Video Object-Centric Learning
Unlocking Few-Shot Capabilities in LVLMs via Prompt Conditioning and Head Selection
SGP2: Coarse-to-Fine Controllable Multimodal Remote Sensing Image Generation
VisWordBench: Bridging the Gap in Cross-modal Reasoning for Multimodal Large Language Models
360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method
AMCI: Unlock the Potential of Large Multimodal Models for Fine-grained Open-world Classification via Adaptive Memory Context Injection
ReMoMask: Retrieval-Augmented Masked Motion Generation
Attention-based Vision-Language Memory for Spatial Reasoning
ColorFM: An Optimization-to-Learning Framework for Color Transfer via Flow Matching
Confidence-Based Mesh Extraction from 3D Gaussians
Distribution-Aware Feature Selection for Post-hoc Out-of-Distribution Detection
Consistent Feature Transport for Image Relighting
Degradation-Robust and Temporally Consistent Infrared–Visible Video Fusion via One-step Diffusion Framework
Demystifing Video Reasoning
DiffProxy: Multi-View Human Mesh Recovery via Diffusion-Generated Dense Proxies
Distill on a Diet: Efficient Knowledge Distillation via Learnable Data Pruning
DA-F2F: Domain-Adaptive Object Detection with Feature-to-Feature Modulation and Alignment
MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors
FD²: A Dedicated Framework for Fine-Grained Dataset Distillation
DoCoG: Mask-based Multi-Type Grounded Chain-of-Thought for Document QA
Safe Generalization: Mitigating Catastrophic Forgetting in Single-Source Multi-Organ Segmentation via Collaborative Causal Learning
EFlow: Fast Few-Step Video Generator Training from Scratch via Efficient Solution Flow
ELT: Elastic Looped Transformers for Visual Generation
Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks
Fragmented Text Is Insufficient for Image Representation: Fine-Grained Correspondence in Multimodal Dataset Distillation
FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs
Forge4D: Feed-Forward 4D Human Reconstruction and Interpolation from Uncalibrated Sparse-View Videos
Geometric Context Transformer for Streaming 3D Reconstruction
Prototype-Conditioned Imagination for Compositional Zero-Shot Learning
GLARE: Towards Generalizable Detection of Latent Diffusion Images with Global-Local Reconstruction Error
Habitat-GS: A High-Fidelity Navigation Simulator with Dynamic Gaussian Splatting
Heat Kernel Textures -- the Geodesic Gaussians That Do Not Splat
Preserving Knowledge across Space and Time for Continual Video Deepfake Detection
Hierarchical Prompt Injector for Domain Generalization Segmentation
HighlightBench: Benchmarking and Diagnosing Markup-Driven Table Reasoning in Scientific Documents
OP3DSG: Open-vocabulary Part-aware 3D Scene Graph Generation for Real-world Environments
Language-Guided Transformer Tokenizer for Human Motion Generation
Learning to Tessellate: Point Cloud Generation via Recursive Spectral Partitioning
LOOM: Weaving Geometry-Consistent Human-Object Interaction Videos via Progressive Curriculum Learning
MedRepBench: Benchmarking Structured Understanding of Medical Report Images
Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes
MixCompress: Mixture of Experts for Variable Rate Learned Image Compression
ASTAD: Asymmetric Style Transfer for Synthetic-to-Real Adaptation in Autonomous Driving
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
Rethinking Pseudo-Labels: Multi-Granularity Supervision for Domain Adaptive Object Detection
Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting
OCTA-SOT: Online Cross-Modal Trajectory Adjustment for RGBT Anti-UAV Single Object Tracking under Spatio-Temporal Misalignment
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
PoseImageNet: Pose Estimation for Extensive Classes Based on Rich Structure Prototypes
ProAct: Agentic Lookahead in Interactive Environments
ProtoMappingNet: Interpretable Hierarchical Prototypes through Relational Prototype Mappings
ReCamDriving: LiDAR-Free Camera-Controlled Video Synthesis for Novel Trajectories
ReconSplat: Generalizable 3D Scene Reconstruction Beyond Observed Views
ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration
Rethinking Cross-Spectral Image Generation via Shared-Specific Representation
SA-ResGS: Self-Augmented Residual 3D Gaussian Splatting for Next Best View Selection
Iterative Refinement of Semantic and Spatial Representations for Open-Vocabulary Camouflaged Object Segmentation
SAM2Matting: Generalized Image and Video Matting
Structured-Noise Masked Modeling for Video, Audio and Beyond
TopoGAT: Plug-and-Play Topological Graph Attention for Fine-Grained 3D Segmentation
Revisiting Deepfake Detection: BCNet for Robust Generalization Beyond Semantic Dependence
Virtual Category-Guided Continual Generalized Category Discovery
Label-Free Text Prototype Adaptation for Open Vocabulary Segmentation
Towards Real-World Wearable Motion Reconstruction
Unlocking Complex Image Editing via Natively Interleaved Visual Textual CoT with Deep Confidence Reasoning
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
VNC: A Scale-Space Foundation for Learnable 3D Surface Evolution
What Moves? Localized Motion Representations for Compositional Scene Control
When Specialists Meet Generalists: Segmenter-Coordinated Asymmetric Learning for Label-Deficient Concealed Object Segmentation
BrainRiem: Riemannian Prototype Learning for Source-Free Cross-Site Brain Network Diagnosis
NegROI: Click-Centric Uncertainty-Guided Refinement with Scene-Conditioned Negative Prompts for Robust Interactive 3D Segmentation
Bridge-UniPS: Bridging Calibrated Photometric Stereo toward Universal Photometric Stereo
Cambrian-P: Pose-Grounded Video Understanding
Domain Adaptation with Adaptive Imagination for Visual Reinforcement Learning under Limited Target Data
CollectionLoRA: Collecting 50 Effects in 1 LoRA for Deployment
PASR: Pattern-Aware Scene-Conditioned Reasoning for Camouflaged Object Detection
CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA
Deformable and Multi-view Gradient-Aligned Physical Adversarial Camouflage
Diffusion-based dual-view reflection removal
Direct Preference Optimization for Perceptual Alignment via Vision-Language Consistency
Driving like yourself: A Benchmark for Closed-Loop Personalized End-to-End Autonomous Driving
Event-based Sparse-view Background-Oriented Schlieren Tomography
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
DualCount: Structurally Consistent Density and Point Modeling for Zero-Shot Object Counting
GAINS: Gaussian-based Inverse Rendering from Sparse Multi-View Captures
GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence
H-SFP: Hierarchical Federated Learning with Decoupled Split-Model Prototyping
HVA-Fusion:Hierarchical Velocity-Aware 4D Radar-LiDAR Fusion for Robust 3D Object Detection
Geometry Grounding: Elevating Blind Distortion Correction with 3D Structural Priors
General Incomplete Multimodal Learning via Dynamic Quality Perception
InstaPano: Zero-shot Instance Layout Controlled Panorama Generation Via Global Attention Fusion
AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization
LibraGen: Playing a Balance Game in Subject-Driven Video Generation
Liquid Fusion of Heterogeneous Representations Towards General Salient Object Detection
LiteMatch: Lightweight Zero-Shot Stereo Matching via Cost Volume Stabilization
LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
Match-Any-Events: Zero-Shot Motion-Robust Feature Matching Across Wide Baselines for Event Cameras
MG-RWKV: Multi-Grained Context-Aware RWKV for Temporal Forgery Localization
MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane
S2-FracMix: Self-Saliency Fractal Mixup
FedNASP: Federated Vision-Language Navigation with Adaptive Step-wise Personalization
MotionEditGS: Editing Motion and Appearance of 4D Scenes from Monocular Video via Semantically Anchored Gaussians
Neural Harmonic Textures for High-Quality Primitive Based Neural Reconstruction
From Local Windows to Adaptive Candidates via Individualized Exploratory: Rethinking Attention for Image Super-Resolution
Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation
Residual-Guided Expert Specialization for Incomplete Multimodal Learning
Noise-Robust Facial Expression Recognition via Mamba-driven Neighbor Weight Refinement
OmniDS: Dual-Stream Context Fusion for Omnidirectional Depth from Fisheye Cameras
On the Plasticity Collapse in Continual Machine Unlearning
Open-Vocabulary 3D Object Detection with Co-Distillation Discovery and Dual Guidance Robust Training
OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models
RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models
Policy-Based Tuning of Autoregressive Image Models with Instance- and Distribution-Level Rewards
PolyLayout: Multi-room Manhattan Layout Estimation
Progressive Representation Learning for Multimodal Sentiment Analysis with Incomplete Modalities
RBE-Flow:Recurrent Bayesian Estimation on Feature Manifolds for Cross-Modal Registration
REON-NVS: Real-Time Online Novel-View Synthesis from Sparse-View Videos
RePer-360: Releasing Perspective Priors for 360° Depth Estimation via Self-Modulation
Task-Agnostic Incremental Vision-Language Object Detection via Prompt Augmentation and Distribution-Aware Fusion
Spanning the Visual Analogy Space with a Weight Basis of LoRAs
SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models
MED-LCDS: Multi-Expert-Domain CLIP Classification via Logit Calibration
SPARC: Scalable Path-Specific Counterfactual Fairness via Causal Conditional Independence
Structured SIR: Efficient and Expressive Importance-Weighted Inference for High-Dimensional Image Registration
STVFocus: Query-guided Spatio-Temporal Visual Focusing for Video LLMs
SwiftWA: An Efficient Action-Centered World-Action Model
Teaching Vision-Language-Action Models What to See and Where to Look
ThermoGS: Decoupling Physical Surface Attributes for Spatio-Temporal Thermal Field Emulation via 4D Gaussian Splatting
Topology-Weighted Effective Rank: A Zero-Cost Proxy for Training Dynamics Stability in Neural Architecture Search
DualTAP: A Dual-Task Adversarial Protector for Mobile MLLM Agents
Towards Sparsely Annotated Open World Object Detection
Fisher-Routed Mixture of Experts for Federated Class-Incremental Learning
Two Birds, One Projection: Harmonizing Safety and Utility in LVLMs via Inference-time Feature Projection
UniTemp: Unlocking Video Generation in Any Temporal Order via Autoregressive Distillation
UrbanAlign: Post-hoc Semantic Calibration for VLM-Human Preference Alignment
VectorReLoc: Reliable Vectorized SD Map Visual Re-localization with Contrastive Feature Alignment
VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision
World Reconstruction From Inconsistent Views
A Mechanism-Driven Theory of Phase Transitions in Active Learning
Accelerated Likelihood Maximization for Diffusion-based Versatile Content Generation
MorphGS: Morphology-Adaptive Articulated Motion Transfer from Videos
Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache
Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning
Better Call CineCrew: Consistent Ultra-Long Narrative-to-Film Generation
SPDA: Efficient Online Test-Time Adaptation for Promptable Medical Segmentation
One Trap to Block Them All: Defending Encoder Stealing via Isotropic Uniformity
Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation
BeyondMasks: Evaluating Causal and Physical Consistency in Video Object Removal
Deformable Triangle Splatting: Flexible Primitives for Real-Time Radiance Field Rendering
DETRPose: Real-Time End-to-End Multi-Person Pose Estimation via Modified Transformer Decoder and Novel Denoising Keypoints
Identifiable Gated Residual Personalization for Federated Parameter-Efficient Fine-Tuning
TSEmbed: Unlocking Task Scaling in Universal Multimodal Embeddings
DiffHDR: Re-Exposing LDR Videos with Video Diffusion Models
Diffusion-Based Immersive Visual Reasoning
PACO: Stabilizing Vision Embeddings along Local Paths for Robust Vision-Language Models
Diffusion-Based Material Regularization for Physics-Based Inverse Rendering
DreamWorld: Geometry-Grounded Video Diffusion for 3D-Consistent World Modeling
DriveVA: Video Action Models are Zero-Shot Drivers
EmbedCopilot: Evaluating Vision-Language Models for Hardware-Aware Embedded System Development
Event Stream-based Sign Language Translation: A High-Definition Benchmark Dataset and A Novel Baseline
Extreme Face Super-Resolution through Identity Fitting and Decoupling
FeVOS: Foresight Expression Video Object Segmentation
FillGS: Filling Observation Gaps in 4D Gaussian Splatting via Viewpoint-Time Selection and Generative Refinement
COVERT: Privacy-Preserving Covariant Obfuscation for VLMaaS via Exact Reparameterization and Tailored Tuning
VD-LoRA: Adaptive Reuse of Low-Rank Directions for Continual Learning
FLEG: Feed-Forward Language Embedded Gaussian Splatting from Any Views via Compact Semantic Representation
Gaussian Belief Propagation Network for Depth Completion
Gravity-aware partially calibrated absolute pose estimation from affine- or rotation-covariant features
Improving Reasoning in Vision-Language Models via Perception Verified Self-Training
InfraNet: Quality-Aware RGB Guidance for Infrared Object Detection
LDC-MTL: Balancing Multi-Task Learning through Scalable Loss Discrepancy Control
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
Local-to-global Cross-modal Coordination for Self-supervised RGB-T Tracking
Map2World: Segment Map Conditioned Text to 3D World Generation
ME-IQA: Memory-Enhanced Image Quality Assessment via Re-Ranking
Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering
From Local Geometry to Global Pseudo-Labeling for Robust Positive–Unlabeled Learning under Covariate Shift
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
Isotropic Embedding Perturbations for Robust Vision Language Encoders
OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents
OmniX: Any-view and Any-time 4D reconstruction via Feed-forward Trajectory Fields
OrthoTailor: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation
PRISM: Latent Composition Consistency for Single-Image Reflection Removal
Reflecting Process Expertise in Procedural Material Generation
Low-Rank Ternary Adaptation for Fine-Tuning Transformers
ReSWD: ReSTIR‘d, not shaken. Combining Reservoir Sampling and Sliced Wasserstein Distance for Variance Reduction.
Bottom-up modeling of repeated elements via single image analysis-by-synthesis
Revisiting the Volumetric Data of 4DME: Compression, Extension and Benchmarking for Micro-Expression Analysis
SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting
SARA: Structure-Aware Riemannian-Guided Alignment for Drone Image-Text Retrieval
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
OCA: ODE-Driven Cross-Attention for Image-to-Point-Cloud Registration
Steerable Vision Transformers
TerrainGraphNet: Terrain-Constrained Graph Reasoning for Landslide Segmentation
Test-Time Noise Guided Adaptation for Realistic Autoregressive Video Generation
Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models
DepWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors
Training-Free Refinement of Flow Matching with Divergence-based Sampling
RGB-Pointmap Pretraining for Unified 3D Scene Understanding
UniCSG: Unified High-Fidelity content-constrained style-driven generation via Staged Semantic and Frequency Disentanglement
Unified Prediction and Planning via Conflict-Aware Disjoint Parameter Training
Unmasking-Time Visual Calibration for Hallucination Mitigation in Multimodal Discrete Diffusion Language Models
VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension
ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images
VIGS-SLAM: Visual Inertial Gaussian Splatting SLAM
VLTR: Vision-Language Tool Reasoning for Instruction-Guided Image Editing
2D Features Are All You Need for 3D Shape Understanding
Agentic Collaborative Cognition for Zero-Shot 3D Understanding
Attention is Case-Sensitive
BWAFDA: Block-wise Weighted Attention Fusion with Detail-aware for No-Reference Image Quality Assessment
CoCo-IR: Conversational Composed Image Retrieval
RayMap3R: Inference-Time RayMap for Dynamic 3D Reconstruction
Constrained Rotation Optimization: Revisiting Crop-Based Gaze Estimation
Context-Aware Joint Alignment for Cross-Scene Hyperspectral Image Classification
Cross-Resolution Distribution Matching for Diffusion Distillation
Curvature-Guided Mixing for MLLM Adaptation
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
DINOv3D: 2D-3D Joint Optimization for Unified Spatial Understanding
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance
Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry
GaINeR: Geometry-Aware Implicit Neural Representation for Image Editing
SLAM-Former: Putting SLAM into One Transformer
TAQ: Static-Deployable Temporal-Aware Quantization for Real-World Video Super-Resolution
HiPolicy: Hierarchical Multi-Frequency Action Chunking for Policy Learning
Histogram-constrained Image Generation
IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video
JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
LANCE: Low Rank Activation Compression for Efficient On-Device Continual Learning
LEAP-VLA: Latent-Enhanced Action Prototyping via Continuous Residual Latent Spaces for Vision-Language-Action Models
ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents
Learning 1-Bit LiDAR-based Localization with Auxiliary Objective
SWIFT: Spatial-Window Integrated Frequency-aware Token Pruning for Efficient MLLMs on Edge Devices
Less is More: Reducing Complexity in Vision-Language-Action Systems
MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning
LASER: A Corrective Lens for LVLMs via Visual Attention Preservation and Sink Suppression
LISA: Locality-Informed Speculative Decoding for Accelerating Autoregressive Image Generation
OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning
MoCam: Unified Novel View Synthesis via Structured Denoising Dynamics
MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
MotionAnymesh: Physics-Grounded Articulation for Simulation-Ready Digital Twins
MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction
MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving
Noise-Robust Face Recognition via Non-target Similarity Distribution Guided Sample Selection
OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
Preventing Expert Collapse in MoE-dVLMs via Modality-Wise Norm Alignment
ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space
Enhancing Interpretability in CLIP with Optimal Transport-based Submodular Optimization for Ophthalmic Imaging
RefDiT: Local Attribute Guidance in Reference-Based Image Generation
ReSplat: Learning Recurrent Gaussian Splatting
PanoRec: Spatially-Structured Sequence Modeling for Multi-Granularity Panoramic Retrieval
Revisiting Weakly-Supervised Video Scene Graph Generation via Pair Affinity Learning
Reward Modeling for Computer-Using Agent from Video Execution
Robust onion: Peeling Open Vocab Object Detectors Under Noise
SALT: Self-Consistent Distribution Matching with Cache-Aware Training for Few-Step Video Generation
Single-Query Person-Centric Bimanual Hand-Object Interaction Detection
Toward Robust In-Context Segmentation via Concept Guidance
Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving
VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
μFlow: Leveraging Average Images for Improving Generalisation of Deepfake Faces Detectors
A Dual-space Patch-driven Complementary Learning Framework for Semi-supervised Multi-organ Segmentation
Adaptive Latent Trajectory Anchoring for Action Segmentation Dataset Condensation
IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment
Aligning Anything: Hierarchical Motion Estimation for Video Frame Interpolation
Mixture of Specialized Vision Experts: Unlocking Complementary Visual Insights for Faithful MLLM Reasoning
APT: Anchor-aligned Perturbations for Tamper Localization in Fully Regenerated Images
Attention-Logit Steering to Compositional Generalization for Continual VQA
Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Large Language Models and Reinforcement Learning
MindBlock: Probing Spatial Assembly and Structure in Unified Multimodal Models
Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
Beyond Random Sampling: Distribution-Aware Alignment for Semi-Supervised Medical Image Segmentation
BLOB-Q: Boosting Low Bit ViT Quantization via Global Optimization on Model Distortion
Bridging Vision and Language Concepts through Optimal Transport Semantic Flow
SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing
Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis
CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection
MotionAtlas: A High-Quality Dataset and Benchmark for Dense Motion Captioning
CLUE-VAD: Structured Semantic Clues for Understanding Explainable Events in Video Anomaly Detection
Co-evolving Representations in Joint Image-Feature Diffusion
CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing
Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations
Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection
CURE: Cumulative Knowledge Reuse for Efficient Device-Server Hybrid Inference in Vision-Language Models
DARE to Mitigate Hallucination: Dual-path Auto-Regressive-aware Editing
EffiDINO: Task-Specific Model Pruning via Gram Anchoring Subspace Consistency
Towards Temporal Compositional Reasoning in Long-Form Sports Videos
ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval
Environmental Change Detection for Real-World Change Analysis
Equivariant Symmetry-Aware Head Pose Estimation for Fetal MRI
EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders
EVEE: Event-Based Online Adaptation for Matching on Unknown Targets
Fast Dynamic Prototypes for Unsupervised Anomaly Detection and Localization
GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models
HUGE-Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks
Hybrid-LUT: Channel-Aware Hybrid Lookup Table and Filtering for Efficient Image Restoration
Layering Virtual Try-On
Learning Consistent Temporal Grounding between Related Tasks in Sports Coaching
ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
MeGAS: Thermomechanical Dynamic Gaussian Splatting for Thermophysical Scene Editing
Mimicking Radiologists: A Coarse-to-Fine Framework with Structural Sparse Tokens for Dual-LLM Computed Tomography Report Generation
From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
NavWM: A Unified Navigation World Model for Foresight-Driven Planning
Physics Meets Perception: A Reinforcement Learning Framework for Unpaired Real-World Image Dehazing
PrintAnything: Learning Geometric Plan Map for 3D Printing G-code Generation from Unoriented Point Clouds
Rethink Backdoor Robustness in Vision Transformers
SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation
FUSE: A Flow-based Mapping Between Shapes
SDUM: A Scalable Deep Unrolled Model for Universal Cardiac MRI Reconstruction
SGC-Lane: Monocular 3D Lane Detection with Standard-Definition Map Guidance and Lane Completion
SkyLume: A Large-Scale Multi-Illumination Aerial Benchmark for Urban Scene Reconstruction and Beyond
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
SMP-UWGS: Coupled Physics-Geometry Optimization for Scalable Multi-Partition Underwater 3D Reconstruction
SOVTrack: Open-Vocabulary Multi-Object Tracking with Self-Supervised Pseudo Labeling and Feature Distillation
OmniNWM: Unifying the State-Action-Reward Triad for Closed-Loop Panoramic Driving Navigation World Models
SplitHDR: Saturation-Aware HDR Recovery and Denoising for Real-Time Detection
Towards Generalizable Robotic Manipulation in Dynamic Environments
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
SyncVL: Synchronizing Vision ⟷ Language Using Unsupervised Adaptation
SynHMR: Synergistic Joint-Mesh Modeling for LiDAR-based Human Mesh Reconstruction
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents
A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention
Training-free Cross-domain Few-shot Segmentation via Robust Semantic Representation and Matching
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
CoMaTrack: Competitive Multi-Agent Game-Theoretic Tracking with Vision-Language-Action Models
Beyond Inpainting: Unleash 3D Understanding for Stable Camera-Controlled Video Re-rendering
VGEdit: Unlocking Video Generation Priors for Reasoning-Informed Image Editing
ANFI: Rethinking Neighbor Feature Interaction in Person Re-ID
RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics
BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension
CoDePose: Multi-View 3D Human Pose Estimation via Coupled 2D-3D Denoising Diffusion
COSY: Compositional 3DGS Synthesis for Disentangled Human Head Editing
CulinaryCut: A Physics-aware Vision-Language-Action Benchmark for Food Cutting via Material Point Method
Depth-guided Multi-view Exposure Bracketing for HDR Robot Vision
DETR is Secretly a Multispectral Detector: Zero-Parameter Adaptation via Semantic Alignment
DICT: Data Injection and Contrastive Trajectory Refinement for Conditional Image Generation with Diffusion Models
EgoPHI: Estimating 3D Hand-Object Contact and Force from Egocentric Vision
Stabilizing Real-World Visual Active Tracking with Action-Smooth Test-Time Adaptation
MobileManiBench: Simplifying Model Verification for Mobile Manipulation
Deconfounded Lifelong Learning for Autonomous Driving via Dynamic Knowledge Spaces
ELHINN: Unifying Dense Crowd Simulation Across Scales via Eulerian–Lagrangian Hydrodynamics
Explainability-aware Frustum Attack: Exposing Structural Vulnerabilities in LiDAR-Based 3D Object Detectors
Flexible Control of 3D CT Generation via Text and Semantically-Defined Segmentation Prompts
From Reconstruction to Decision: A Post-Encoder Plug-in Adapter for Curvilinear Segmentation
Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention
Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior
GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors
Geometry-Aware Spatio-Temporal Context Modeling for 4D Occupancy Forecasting
Granular Semantic Cognition for Visible-Infrared Person Re-Identification
GridVQA-X: A Diagnostic Framework for Evaluating Multimodal Explainability Methods
GrowFields: Compositional 4D Neural Fields for Topology-Changing Plant Growth
Incremental Online Scene Reconstruction by 3D Gaussian Triangulation
LACON: Training Text-to-Image Model from Uncurated Data
LaVPR: Benchmarking Language and Vision for Place Recognition
SP-TransientBench: A Real-Captured Single Photon Perception Benchmark
Learning to Suppress SPAD-based LiDAR Flare
Learning Video Dynamics with Predictive Differentiable Rendering
RESOLVE: A Multi-Resolution and Multi-Modal Dataset for Roadside Cooperative Perception
Lightweight Online Reinforcement Learning for Block Decomposition of CAD Models
OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence
TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception
Locality-Aware Continual Unlearning for Diffusion Models
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization
MagnetGS-Mesh: High-Quality Multi-Object Mesh Reconstruction via Adaptive Surface Optimization
MLVC: A Multi-platform Learned Video Codec for Real-World Deployment
MVI2V: Human Centric Image to Video Generation with Multiview Consistent Appearance
Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes
One-Shot Feed-Forward 360° Animatable Avatar via Inpainted UV-Space Gaussian Modeling
Open-Weather Robust 3D Detection via Dual-Critic Diffusion Alignment
PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videos
PKINet-v2: Towards Powerful and Efficient Poly-Kernel Remote Sensing Object Detection
PriorMaskMap: Robust Online Vectorized Map Construction with Biased Priors
Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation
ReInGS: Re-Initializing 3D Gaussians against Sparsity Discrepancy in Few-Shot Novel View Synthesis
Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution
Sequential Visual Place Recognition: Exploiting Trajectory Priors for Robust Localization
SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset
Social-Mamba: Socially-Aware Trajectory Forecasting with State-Space Models
Rethinking Attention Reallocation for Multimodal Emotion Recognition
NutriBench-Kitchen: Benchmarking Embodied AI for Nutrition Management
Stand Up and Move: Benchmarking Interactive Spatial Intelligence in WalkerBench
Sticking Information in Plain Sight: Encoding and Detecting Hidden Stickers in the Real World
SuperVoxelGPT: Adaptive and Ordered 3D Tokenization for Autoregressive Shape Generation
TerraDiT-Ω: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive
SUM: Unified Geometric Surgery on Spatio-Temporal Adaptation Vectors for Federated Class Incremental Learning
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
Thinking Ahead: Foresight Intelligence in MLLMs and World Model
Two-Way Street: Efficient VSLAM using Collaborative In-Sensor and Off-Sensor processing
Uncertainty-aware tree height change regression
Hi-Nav: Hierarchical Framework for Continuous Vision-Language Navigation via Map Guidance and Waypoint Reasoning
UECP: Uncertainty-Enhanced Collaborative Perception
Vision-TTT: Efficient and Expressive Visual Representation Learning with Test-Time Training
Twin-DAgger: Synergizing Digital Twins and Human Corrections for Efficient Robot Manipulation
CausalVAE as a Plug-in for World Models: Towards Reliable Counterfactual Dynamics
SAEdit: Token-Level Control for Continuous Image Editing via Sparse Autoencoder
WorldCache: Content-Aware Caching for Accelerated Video World Models
DH-VLM: Dual-Horizon Cooperative Latent Reasoning for Autonomous Driving
XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity
Zero-Shot Novel Depth Synthesis Using Foundation Models Scene Representations
AccelAes: Accelerating Diffusion Transformers for Training-Free Aesthetic-Enhanced Image Generation
AutoCompass: Accurate Visual Localization on Public Maps by Learning from Weak Labels
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
Beyond Categorical Matching: Intra-Class Graded Relevance Estimation for Cross-Modal 3D Retrieval
Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum
CMDS-AD: Cross-Modal Dual-Stream Decoupling for Few-Shot Anomaly Detection
ConTrack: Constrained Hand Motion Tracking with Adaptive Trade-off Control
Contrastive Conditional–Unconditional Alignment for Long-tailed Diffusion Model
Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning
CustomX: Unified Character, Action, and Scene Customization in Video World Models
DASAM3D: A Unified Foundation Model for Enhanced 3D Scene Reconstruction and Segmentation
Do Flat Minima Improve Sparse Novel View Synthesis?
DriveFine: Refining-Augmented Masked Diffusion VLA for Accurate and Robust Driving
Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout
Dual-End Consistency Model
EatVid-Bench: A Multimodal Fine-Grained Eating Behavior Video Dataset
ExpoMotion: A Large-Scale Benchmark and A Householder Projection Network for Multi-Exposure Fusion
From Predictions to Embeddings: Dual Knowledge Distillation for Instance-Dependent Partial Label Learning
FUSE: Filter-Free Unified Spatiotemporal Estimation of SpO2 via Wave-Transport Modeling
GraphCPD: Coherent Point Drift for Point Cloud Registration via Graph Signal Processing
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
GRE-Diff: Gaussian Room Embeddings for Structured Layout Diffusion
HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework
QST-SAM: Leveraging Cross-modal Instructions for Few-shot Referring Video Object Segmentation
How to Teach Large Multimodal Models New Skills
Semantic Line Diffusion: Character-Consistent Line Art from text-annotated Storyboards
HuCollisionField: Resolving Self-Collisions via Neural Fields for Human Prediction
Inductive Visual Logic for Few-Shot Out-Of-Distribution Adaptation in VLMs
Learning Spectral and Polarimetric Clues for One-to-Multimodal Novel View Synthesis
LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
LogicIR: Logic Gate Networks for Image Restoration
Memory-Supported Synergistic Adaptation for Training-Free Test-Time Medical Image Segmentation
MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model
PhenoLeaf-TS: A Time-Series Benchmark for Leaf Instance Segmentation, Tracking, and Growth Stage Classification
PRISM-VO: Scale-Aware Visual Odometry Using Photometric Plenoptic Bundle Adjustment
Proximity-Constrained Counterfactual Decoding for Hallucination-Robust Medical VQA
QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers
Rectified Embedding Flow Learning for Aerial Multi-view Geo-localization
Recurrent Cross-View Object Geo-Localization
RefracGS: Novel View Synthesis Through Refractive Water Surfaces with 3D Gaussian Ray Tracing
Revisiting Scene Graph Generation from the Perspective of Detector-Conditioned Reachability
Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs
Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training?
SFDATrack: Generalized Source-Free Domain Adaptive Tracking Under Adverse Weather Conditions
Sound-based Multi-Person 3D Pose Estimation
HiChor: Hierarchical Choreography Generation from Pop Music with Choreographic Primitives
SPAR: A Sequential Primacy and Attribution Ranking Framework for Skill Determination
SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
SVI360: Spherical Video Interpolation
UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization
Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?
Task-driven Processing with Coarse-to-Fine Glimpse-based Active Perception
The Path to Reconciling Quality and Safety Alignment in Text-to-Image Generation
The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
TiCRL: Textual Image Classification with Reinforcement Learning-Based Curriculum Learning
To Adapt or Not to Adapt? Selective Adaptation for Vision-Language Models
TRAM: Finetuning-Free Test-Time Adaptation for Generalized Face Anti-Spoofing with Only a Few Bonafide Samples
TriFlow: Generating Artist-Like 3D Mesh Topology via Nearest-Vertex Vector Fields
UEval: A Benchmark for Unified Multimodal Generation
Unpaired Geometry-Guided Sim2Real Translation for Autonomous Driving
Urban Boundaries, Social Barriers: A Benchmark and Vision-Centric Framework for Mapping Gated Communities and Equity Implications
Wan-R1: Verifiable-Reinforcement Learning for Generalizable Video Reasoning
Wavelet-Driven Cross-Domain Consistency for Mixed-Supervised 3D Tumor Segmentation
WildProp: Visual Estimation of Wildlife Body Proportions at Scale
XPos3R: Cross-Modal Transformer for Intraoperative 2D/3D Registration
A Classifier-Agnostic Zero-Shot Adversarial Attack Detection via CLIP
AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution
Any to Full: Prompting Depth Anything for Depth Completion in One Stage
ATOMIC: A Domain-Specific Vision-Language Model for Transmission Electron Microscopy
PRISM3D: Probabilistic Refinement and Robust Initialization for Physically Consistent Scene Modeling under Extreme Motion Blur
Beyond Isolated Scans: Cross-Phase Alignment of Structure and Topology for 3D Medical Pretraining
Boosting Text-Driven Video Segmentation via Geometry-Aware Distillation
BrepCoder: A Unified Multimodal Large Language Model for Multi-task B-rep Reasoning
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
Comprehensive language–image pre-training for 3D medical image understanding
DeMuS: Learning Decoupled Matching and Scoring for Batch Zero-Shot Industrial Anomaly Detection
Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts
DreamEdit3D: Personalization of Multi-View Diffusion Models for 3D Editing
FeDepth: Federated Learning for Depth Estimation under Robot Heterogeneity
NoiseTilt: Noise-Tilted Reverse Kernels for Diffusion Reward Alignment
DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation
EAGS: Error-Aware Gaussian Splatting with Dual-Confidence-Guided Modeling for Uncalibrated Driving Scenes
ECTraj: Enhanced Consistency Training for Multi-Agent Trajectory Prediction
EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents
Error-Driven Scene Editing for 3D Grounding in Large Language Models
Event-LiDAR: 3D Eventification for Efficient Point Cloud Processing
Fill2SR: Repurposing Inpainting Diffusion Transformers for Real-World Super-Resolution
Conditional Flow Matching for Visually-Guided Acoustic Highlighting
FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility
AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation
Natural Language Camera Movement Understanding
Fourier Splatting: Generalized Fourier encoded primitives for scalable radiance fields
Generalized Biomedicine Discovery
Articulated Object Reconstruction from Rest-State Observation
HHA: Hierarchical Hyperbolic Constraints for Imperceptible Point Cloud Attacks
Histopathology Multi-modal Embedding for Pathology Composed Retrieval
HLRAD: High-dimensional Latent Representation for Unified Anomaly Detection
Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions
LaxMotion: Rethinking Supervision Granularity for 3D Human Motion Generation
LEO-Fuse: A Modality- and Task-Agnostic Universal Framework for Multimodal Human Sensing
LINA: Learning INterventions Adaptively for Physical Alignment and Counterfactual Generation in Diffusion Models
Linear Fusion MultiDiffusion for Fast Training-Free Spherical Panorama Generation
MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
MirrorPPR: Exemplar-Based Portrait Photo Retouching
MMLoP: Multi-Modal Low-Rank Prompting for Efficient Vision-Language Adaptation
ModTrack: Sensor-Agnostic Multi-View Tracking via Identity-Informed PHD Filtering with Covariance Propagation
MomentSeg: Moment-Centric Sampling for Enhanced Referring Video Object Segmentation
Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing
MoScale: Autoregressive Next-Scale Prediction for Human Motion Generation and Editing
NearID: Identity Representation Learning via Near-identity Distractors
ORION: Ordinal Neural Collapse as a Representation Prior for Visual Navigation
RAU: Reference-based Anatomical Understanding with Vision-Language Models
P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling
PhysEdit: Physically Consistent Image Editing via Causal Enforcement
PixelU: A U-Shaped Transformer for Efficient End-to-End Pixel Diffusion
Point2Pose: Occlusion-Recovering 6D Pose Tracking and 3D Reconstruction for Multiple Unknown Objects Via 2D Point Trackers
NoPA: Non-Parametric Online 3D Scene Graph Generation
Prosthesis-Aware 3D Human Pose Estimation: A Dataset and Benchmark for RSP Users
Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction
ReliefSAM: A Geometry-Augmented Multi-Prior Adapter for Bas-Relief Segmentation
RGBT-GroundBench: Visual Grounding Beyond RGB in Complex Real-World Scenarios
Robustness Meets Uncertainty: Evidential Adversarial Training for Robust Selective Classification
Sen-Cap: Sensor-Flexible and Noise-Resilient Human Motion Capture via LiDAR-Camera Integration
Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery
STAT: Soft Tail-dropping for Adaptive Visual Tokenization
Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs
Towards Interactive Global Geolocation Assistant
Versatile Editing of Video Content, Actions, and Dynamics without Training
Video-Text Alignment Model for Sign Language Translation
HandSCS: Structural Coordinate Space for Animatable Hand Gaussian Splatting
3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
A Dual-Transformer Architecture with Cross-Attention for Multi-Camera View Recommendation
AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors
Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers
Beyond the Boundary: RL-Driven Solution Space Exploration for Blind Face Restoration
Boosting 6D Object Pose Estimation via Monocular Depth Cues
WiFi-JEPA: Self-supervised Learning for WiFi-CSI 3D Human Pose Estimation
Breaking the Model Forgetting Cycle in Long-Incremental 3D Object Detection
CAST3D: Customizing Arbitrary 2D Assets into 3D World
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
ChronoFlow Policy: Unifying Past-Future Interaction Flow in Visuomotor Policy Learning
CTEPM: Continuous-Time Event Process Memory for Long-Video Language Models
DANTE-W: Diffuse Albedo Neural Texturing in the Wild
Decoupled Illumination Priors for Spatially Controllable Multi-View Indoor Scene Relighting
PriSplat: Propagating Reliable Multi-view Information for Distractor-Free 3DGS
DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking
DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection
Dynamic-Robust Photometric–Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding
EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation
EventVGGT: Exploring Cross-Modal Distillation for Consistent Event-based Depth Estimation
EvoWorld: A World-Model-Centric Framework for Continuous Self-Evolution of Modular Embodied Skills
FlexiBrain: Resolution-Agnostic Voxel-Level Encoding for Native fMRI
Frames2Residual: Spatiotemporal Decoupling for Self-Supervised Video Denoising
Free-Range Gaussians: Non-Grid-Aligned Generative 3D Gaussian Reconstruction
Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation
From Open Loop to Closed Loop: A Test-Time Iterative Optimization Framework for Reference-Consistent Image Generation
Gaussians on Fire: High-Frequency Reconstruction of Flames
ZAP: Zero-Shot Assembly Planning with Large Language Models
Geometric Gradient Rectification for Safe Open-Set Semi-Supervised Learning
GIDE: Unlocking Diffusion LLMs for Precise Training-Free Image Editing
SiGMA: Sign-Guided Merging and Adaptation framework for Multimodal Continual Instruction Tuning
Learning Consistency in Reward Modeling for Multi-Modal Reasoning
HERO: Enhancing Multimodal Faithfulness via Dynamic Entropy-Aware Reinforcement Learning
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
Learning Probabilistic Embeddings for Unsupervised Action Segmentation
Learning Transferable Dynamics Priors from Action to World Modeling
Low-Level Dataset Distillation for Medical Image Enhancement
Multi-label Instance-level Generalised Visual Grounding in Agriculture
MuSViT: A Foundation Vision Model for Sheet Music Representation
RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes
Condensing Large-Scale Datasets Directly with Minimal Information Loss
MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing
On the Vulnerability of Parameter-Level Defenses to Model Merging
Guiding the Blind: Generalizing GUI Agents to Unseen Websites via Multimodal Tutorials
Personalized Reward Modeling for Text-to-Image Generation
Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation
Pro-Pose: Unpaired Full-Body Portrait Synthesis via Canonical UV Maps
Raw-JPEG Adapter: Efficient Raw Image Compression with JPEG
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
Reflection-aware generative novel view synthesis
REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion
Robustness Emerges Early in Training Dynamics, but Is Not Preserved
SegVGGT: Joint 3D Reconstruction and Instance Segmentation from Multi-View Images
Silhouette-based Gait Foundation Model
Spectral-Aware Analytic Class-Incremental Learning for Long-Tailed Distributions
SuperFlex: Deformable Superquadrics for Point Cloud Decomposition
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
SynLF: Zero-Shot Metric Depth from Light Field Cameras via Physics-Grounded Synthesis
Task Alignment: A simple and effective proxy for model merging in computer vision
Tesselating The Earth
The Devil Is in the Dark Pixels: Toward Brightness Bias-Robust Denoising
Towards Spatial Supersensing in the Wild
VIPS: Vehicle-Infrastructure Cooperative Planning Benchmark via Pseudo-Simulation
VLMSysTrojan: Stealthy System-Aware Backdoor Attacks Against Vision-Language Models
What Images Cannot Say: Language-Guided Olfactory Representation Learning
White Aggregation and Restoration for Few-shot 3D Point Cloud Semantic Segmentation
AV2T-Gen: Aerial Visible to Thermal Generation with Environment and Vehicle State Guidance
Beyond Common Sense: Grounding Logical Anomaly Detection in Inspection Criteria
Stable and Scalable Bundle Adjustment of Holistic 3D Structures
COLA: Continual Orthogonal Low-Rank Adaptation for Class-Incremental Learning
Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs
CUST : Clustered Unit-level Similarity Transformer for Lightweight Image Super-Resolution
Diffusion Image Generation with Explicitly Modeling of Data Manifold Geometry
Do Multimodal LLMs Understand Intraoral Dental Data? Dataset, Platform, and Baselines
Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models
ESC: Emotional Self-Correction for Reliable Vision-Language Models
FDM-MFVT: Few-step Sampling Diffusion Model for Mask-Free Virtual Try-On
FedDO: Dynamic Client Optimization for Adaptive Federated Learning
FlowPainter: Inpainting Optical Flow via Confidence-Guided Completion
GCMRD: Global Consistency Multi-teacher Robustness Distillation
Graph-GSReg: Leveraging 3D Scene Graphs for Gaussian Splatting Registration
H2SVC: Head-aware Heterogeneous Streaming Video Cache for Online Video Understanding
HOIMask: Towards Generative Masked Modeling for Human Object Interaction Generation
HomeGuard: VLM-based Embodied Safeguard for Identifying Contextual Risk in Household Task
IC-World: In-Context Generation for Shared World Modeling
In-context Region-based Drag: Drag Any Region to Any Shape
InFlux++: Real and Synthetic Data for Estimating Dynamic Camera Intrinsics
Learning Accurate Segmentation Purely from Self-Supervision
LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching
Mask-guided Semantic Alignment: Robust Learning with Noisy Labels via Temporal Attention Stability
Mitigating Sycophancy in Multimodal Chart Understanding via Vision-Grounded Verification
Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion
Online 3D Instance Segmentation at task-oriented granularity with Unposed Monocular Video
PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion
WARP: Wide Attention with Rich Projections for Image Super-Resolution
Point Ladder Tuning: Parameter-Efficient Hierarchical Adaptation for 3D Point Cloud Understanding
PUF: Plug-and-Play Uncertainty-Aware Fusion for Online 3D Scene Graph Generation
ScrollScape: Unlocking 32K Image Generation With Video Diffusion Priors
ReTarget: Representation Transformation via Adversarial Regularization for Geometric Misalignment
RAE-NWM: Navigation World Model in Dense Visual Representation Space
The 3D Mirage: Probing and Taming 3D Hallucinations
ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
RiO-DETR: DETR for Real-time Oriented Object Detection
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futures
SemLight: Distilled Semantic–Geometric Fusion for Efficient Local Feature Matching
SLER-IR: Spherical Layer-wise Expert Routing for All-in-One Image Restoration
Spectral Gating via Damped Oscillations for Adaptive Implicit Neural Representations
STARLINC: Satellite Trail Artifact Removal using Inter-Frame Correlation
The Map Is Not the Territory: Embedding-Coverage Blacklists for Safe Diffusion Steering
Unified Video Dense Prediction from Disjoint Data
Unsupervised Point Cloud Registration via Training-Time Semantic Guidance
URHead: A Unified UV-Space Representation for Joint Mesh–3DGS Optimization in Head Avatars
Human-Level Accuracy, Non-Human Strategies: Revealing Model-Human Divergence in Video Physical Reasoning
VarProtoAD: Variational Prototype-Conditioned Prompting for Zero-Shot Anomaly Detection
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
VQT: Vector Quantization Tuning for Efficient Fine-tuning and Compression of Pre-trained Vision Transformers
Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs
Walk through Paintings : Ego-centric World models from Internet Priors
WebEyeTrack: Scalable Eye-Tracking for the Browser via On-Device Few-Shot Personalization
Linguistically-Aligned and Visually-Grounded Preference Optimization for Clinically-Augmented Medical Report Generation
Where and What: Long-Term Object Tracking in Egocentric Videos
A Comprehensive Analysis about Unsupervised Outlier Detection for Images
Beyond Attention: Convolutional Global Context for Remote Sensing Change Detection
Capturing Spectral and Spatial Patterns for Federated Remote Sensing Segmentation
Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
DiscoVL: Unveiling Disentangled Cross-Modal Representation Learning via Orthogonal Adversarial Regularization for Vision-Language Models
DiTex4D: Direct Text-Driven 4D Generation with Structured Latent Diffusion
From Phase to Phenomenon: Self-Supervised Learning of Subsurface Scattering with Minimal Phase-shift Inputs
GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning
Global Pose Control for Generative View Synthesis in Normalized Object Coordinate Space
GTR: Guide-Then-Refine Token Compression for Training-Free Acceleration of Video-LLMs
Orthogonal Knowledge Refreshing for Domain-Incremental Object Detection
HieDG: A Hierarchical Discrete Geometry-Guided Framework for Multi-Animal Tracking
Holistic Optimal Label Selection for Robust Prompt Learning under Partial Labels
General Self-Calibration with Varying Intrinsics
IConE: Batch Independent Collapse Prevention for Self-Supervised Representation Learning
KineticGS: Momentum-driven Coherent 4D Gaussian Splatting for Monocular Dynamic Scene Reconstruction
Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments
CARA: Collision-Aware Resolution Adaptation for Multiresolution Hash Encoding Based Image Fitting
LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models
MobileSAM2: Lightweight Segment Anything in Images and Videos via Hypergraphical Knowledge Distillation
Modality-Aware Out-of-Distribution Detection for Multi-Modal Action Recognition
Occlusion-Resilient Category-Agnostic Pose Estimation with Conditional Flow Matching
One Slide, Many Views: Unifying Complementary Foundation Model Perspectives for WSI Analysis
Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
Practice Makes Perfect: From Explicit Decomposition to Reinforced Latent Planning in Text-to-Human Motion
Predictive Photometric Uncertainty in Gaussian Splatting for Novel View Synthesis
Probe, Anchor, and Amend: Active Test-Time Adaptation of Vision-Language Models
Q-REAL: Towards Naturalness and Distortion Evaluation for AI-Generated Content
QASA: Quality-Guided K-Adaptive Slot Attention for Unsupervised Object-Centric Learning
Quantile‑Adaptive Temperature Scaling for Confidence Calibration
NeuIDO: Neural Intrinsic Dynamics Operator for Physics-Informed 4D World Models
Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention
ReflectCAP: Detailed Image Captioning with Reflective Memory
ReGen3D: Generalizable Unified Representation Learning for 3D Understanding
Rethinking Real-World MRI Denoising: Learning from Physical Noise
RPM-Distill: Physiology-guided Adaptive Cross-modal Distillation for Robust Remote Physiological Measurement
Training-free Discriminative Patch Mining for Robust Few-Shot Recognition with CLIP
S2Gest: Split-Scan State Space Models for Dynamic Hand Gesture Recognition
S3-Prune: Stability-Aware Token Budgeting for Long-Form Video-Language Models
Scaling Multi-Reference Image Generation with Dynamic Reward Optimization
Shared LoRA Subspaces for almost Strict Continual Learning
Show Me Examples: Inferring Visual Concepts from Image Sets
Slim-DETR: Real-Time Tiny Object Detection with Efficient Interaction and Gaussian Query
Steering Diffusion Models via Class-Contrastive Influence for Few-Shot Classification
StyleFusion360: View-Consistent Head Stylization via Adaptive Style Modulation
TaskTok: Delving into Task Tokens for Task-driven Image Restoration
TDSR-VLA: Transition-aware Denoising Sequence Representations for Vision-Language-Action
TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection under Distribution Shifts
The Cost of Reasoning: Chain-of-Thought Induces Overconfidence in Vision-Language Models
TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
Towards Alias-Free 4D Gaussian Representations with Motion-Aware Filtering
Towards More Efficient Decoding for Autoregressive Vision-language-action Models
UniFlow: Zero-Shot LiDAR Scene Flow for Autonomous Driving
From Local to Global: A Progressive Reconstruction Network for Diffractive Snapshot Spectral Imaging
Render-FM: Feedforward Model for Real-time Photorealistic Volumetric Rendering
Universal Image Immunization against Diffusion-based Image Editing via Semantic Injection
VERITAS: A Multi-agent Co-scientist for Verifiable Image-Derived Hypothesis Testing
VISOR++ : VISUAL INPUT BASED STEERING FOR LARGE VISION LANGUAGE MODELS
VLOD-TTA: Test-Time Adaptation of Vision-Language Object Detectors
Vulnerability of Privacy-Preserving Visual Localization against Diffusion-based Attacks
XSurfer: Reconstructing surface meshes of cerebral and cerebellar cortex from diverse MRI data using untrained neural networks
CabinSI: Omni-Cabin Spatial Reasoning through Explicit Visual Cognitive Maps
YouTube-Occ: Learning Indoor 3D Semantic Occupancy Prediction from YouTube Videos
AdaThinking-E: One-Token Entropy Regulation for Adaptive Thinking
AffoGato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
Audio-Visual Continual Test-Time Adaptation without Forgetting
Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization
Boosting 3D Foundation Models with Featureless Pose Optimization
CascadeProto: Cascaded Cross-Modal Prototype Purification via Entropy-Aware Learning for Few-Shot 3D Point Cloud Segmentation
CGCC: Towards Generalizable Clothes-Changing Person Re-Identification
CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
Contrastive-Guided Self-Supervised Latent Visual Reasoning for Hallucination Mitigation
Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation
DPGS: A Diffusion-Prior Guided Framework for Large-Scale 3D Gaussian Splatting Reconstruction
Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs
EchoVLA: Robotic Vision-Language-Action Model with Synergistic Declarative Memory for Mobile Manipulation
EGGS: Explicitly Granular 3D Gaussian Splatting via Luma-Aware and Volume-Preserving Attribute Factorization
EoS-FM: Can an Ensemble of Specialist Models act as a Generalist Feature Extractor?
EruDiff: Refactoring Knowledge in Diffusion Models for Advanced Text-to-Image Synthesis
Exploiting Local Flatness for Efficient Out-of-Distribution Detection
Foundation-Guided Representation Alignment for Multimodal Medical Image Registration
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
Graph Coloring for Multi-Task Learning
HairOrbit: Multi-view Aware 3D Hair Modeling from Single Portraits
HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human–Scene Interactions
HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding
Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
Learning Geometry-Aware Embedding Fields for Intrinsic Riemannian Mappings
LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos
LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models
Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling
ETCH-X: Robustify Expressive Body Fitting to Clothed Humans with Composable Synthetic Data
Multi-Head Normalization for Wide Vision Transformers
MUSE: Unlocking Timestep as Native Task Steering for One-Step Dense Prediction
OPAL: Orthonormal Prototype Alignment Learning for Interpretable Image Classification
ViewSplat: View-Adaptive Dynamic Gaussian Splatting for Feed-Forward Synthesis
Parallax Portrait Matting
From smooth to sharp: Frequency-Decoupled Latent Optimization for Realistic Image Generation
PhysAlign: Learning Physical Priors for Dynamical Event-Driven Video Generation via Representation Alignment
PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving
SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training
PointLAM: Local Attentive Mamba for Efficient Point-based 3D Object Detection
Prompt2Effect: Training-Free LoRA Synthesis for Controllable Video Effects
LivingWorld: Interactive 4D World Generation with Environmental Dynamics
ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video
RelAfford6D: Relational 6D Affordance Graphs for Constraint-Driven Robotic Manipulation
F⁴Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting
Rethinking Garment Conditioning in Diffusion-based Virtual Try-On: Decouple, Don't Denoise
Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
SARIF: Segment Anything for Robust Image Forensics
Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing
HEM: a margin-based loss for visual categorisation tasks
SIGNER: Temporally Grounded Sign Language Generation via Time-Resolved Conditioning
SOMA: From Surface Observations to Muscle Anatomy
Sparse-View Surface Reconstruction using Gaussian Splatting through High-Confidence Depth Propagation with Normal Priors
Stokes-Informed Diffusion for Robust Linear Polarization Estimation
Toward Interpretable Analysis of Whole-slide Pathology Images via Large Language Model-based Agentic Reasoning
TRINITY: A Multi-Perspective Benchmark for Personal-Style Video Highlight Detection
Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models
UBone3D: Physics-Rectified Conditional Flow Matching for Anatomical 3D Shape Completion from Ultrasound
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
UniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth Alignment
UniTranslator: A Unified Multi-modal framework for End-to-end In-Image Machine Translation
Wat3R: Underwater 3D Geometry Learning without Underwater Annotations
3D-Layout-R1: Structured Reasoning for Language-Instructed Spatial Editing
AC3S: Adaptive Conditioning for 3D-Aware Synthetic Data Generation
ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs
Automatic Method Illustration Generation for AI Scientific Papers via Drawing Middleware Creation, Evolution, and Orchestration
Beyond Linear Shortcuts: Rectifying Diffusion Preference Optimization with Intrinsic Generative Geometry
CoT-PL: Chain-of-Thought Pseudo-Labeling for Open-Vocabulary Object Detection
CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
AutoWeather4D: Autonomous Driving Video Weather Conversion via G-Buffer Dual-Pass Editing
Cross-Species Animal Re-Identification with Semantic Consistency Learning
Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
DiTailed: Ensuring Visual Object Consistency in Text-Image-to-Image Flow Matching Models
ArcAD: Anomaly-Rectified Calibration for Cold-Start Supervised Anomaly Detection
Doe-2: 3D Representation World Model for Unified Driving Scene Forecasting
Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy
Dual-Output Multi-Exposure HDR Reconstruction via SDR Fusion and Gain Map Inverse Tone Mapping
Egocentric World Model for Photorealistic Hand Object Interaction Synthesis
3D-ReGen: A Unified 3D Geometry Regeneration Framework
EgoCogNav: Cognition-aware Human Egocentric Navigation
Entropy-Controlled Flow Matching
Fabric Image Demoiréing Benchmark from Synthesis to Restoration
FaceArmor: A Universal Facial Image Protection Against Diffusion-Based Manipulations
FingerCap: Fine-grained Finger-level Hand Motion Captioning
FlashBEV: Fast and Memory-Efficient Exact BEV Transformation with IO-Awareness
FlowLess: Controlling Abstract Image Generation
From Script to Shot: A Benchmark for Grounding Screenplays in Movies
GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
HolisticSemGes: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching
Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation
Incentivizing Vision Language Models to Search for Long Video Question Answering
JSON: Jigsaw Self-play Optimization for Normalizing Flows
KineBench: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding
LumiTokens: 3D Relighting via Token-Space Lighting Transformation
MessyKitchens: Contact-rich object-level 3D scene reconstruction
NanoGS: Training-Free and Lightweight Gaussian Splat Simplification
C3-Bench: A Context-Aware Change Captioning Benchmark
Narrative-Driven Paper-to-Slide Generation via ArcDeck
NGPS: Structure-Preserving Self-Supervised Denoising via Neighbor-Guided Patch Sampling
RoomPlanner: Reachability-Aware View Sampling for Text-to-Room 3D Gaussian Splatting
Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis
OmniFace: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer
On the Reliability of Cue Conflict and Beyond
OneVAE: Joint Discrete and Continuous Optimization Helps Discrete Video VAE Train Better
OREO: Fidelity Alignment in 3D Generation via On-the-fly Rendering-Editing Optimization
Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
Personalize Your Large Vision-language Models With In-context Prompt Tuning
PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation
Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective
SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation
SegFly: A 2D-3D-2D Paradigm for Aerial RGB-Thermal Semantic Segmentation at Scale
SEM-ROVER: Semantic Voxel-Guided Diffusion for Large-Scale Driving Scene Generation
SPLIT: Training-Free AI-Generated and Partially Edited Video Detection via Spatial Patch‑Level Incoherence and Temporal Roughness
Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
Towards Reliable Medical Large Vision-Language Models via Counterfactual Preference Optimization
Towards Reliable Multi-Label Classification via Conditional Dependency Modeling
Trust-Region Noise Search for Black-Box Alignment of Diffusion and Flow Models
UniBYD: A Unified Framework for Learning Robotic Manipulation Across Embodiments Beyond Imitation of Human Demonstrations
VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus
VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
XDen-1K: A Density Field Dataset of Real-World Objects
λSplit: Self-Supervised Content-Aware Spectral Unmixing for Fluorescence Microscopy
ActionPlan: Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning
Beyond Sequential Distance: Inter-Modal Distance Invariant Position Encoding
Calibrate Before Adapt: Training-Free Pseudo-Label Calibration for Semi-Supervised Cross-Domain Few-Shot Detection
CGCE: Classifier-Guided Concept Erasure in Generative Models
CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization
Concept-to-Pixel: Prompt-Free Universal Medical Image Segmentation
Hierarchical 3D Scene Graph Construction and Belief-based Planning for Semantic Navigation
Controllable Generative Reference for Stereo Image Compression via Reliability-Aware Gating
Controlling Embedding Spaces with Text-Conditioned Transformations
CoReLIN: Constraint-based Reasoning for Zero-shot Lifelong Interactive Navigation
Denoising the Deep Sky: Physics-Based CCD Noise Formation for Astronomical Imaging
DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion
Revisiting Autoregressive Models for Generative Image Classification
Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
Diversity-Aware View Partitioning for Scalable VGGT
EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance
Flash-DD: An Ultra Parameter-Efficient Approach to Dataset Distillation
FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation
Fourier Self-Supervision for Fine-Grained Generalized Category Discovery
From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models
Thinking in Streaming Video
From Illusion to Intention: Visual Rationale Learning for Reliable Evidence Acquisition
From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal
DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding
GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
Geometric Foundation Model Distillation for Efficient Lunar 3D Reconstruction
GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
Hierarchical Spatial and Channel Aggregation for Cross-domain Few-shot Segmentation
Information-Regularized Attention for Visual-Centric Reasoning
Intra-Class Consistency Guided Class-Agnostic Event Segmentation
Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding
M4-SAR: A Multi-Resolution, Multi-Polarization, Multi-Scene, Multi-Source Dataset and Benchmark for optical-SAR Object Detection
MedGSSR: Generalizable Medical Image Super-Resolution 3D Reconstruction via Hierarchical Feed-forward Gaussian Splatting
Meric: A Unified Framework for Multimodal Music Generation and Retrieval via Representation Space Anchoring
Multi-Hypothesis Test-Time Adaptation to Mitigate Underspecification
Noise is a Good Teacher: A Noise-Driven Framework for Robust Collaborative Perception
OmniX: From Unified Panoramic Generation and Perception To Graphics-Ready 3D Scenes
OrthoEraser: Coupled-Neuron Orthogonal Projection for Concept Erasure
P²Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation
PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models
ProSR: Semantic-Prototype-Guided Discrete Modeling for Physically Consistent SAR Super-Resolution
RealDyadic: Synthesizing Realistic Dyadic 3D Dialogue with Neural Appearance Priors
Rethinking Prototype-based Similarity Learning for Few-Shot Object Detection
RIGS: Radar-Informed Gaussian Splatting for Uncertainty-Aware 3D Occupancy and Motion Prediction
SEAR: Simple and Efficient Adaptation of Visual Geometric Transformers for RGB+Thermal 3D Reconstruction
StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning
StructSplat: Generalizable 3D Gaussian Splatting from Uncalibrated Sparse Views
SV-TAD: Native Sparse Convolutions for Efficient Temporal Action Detection
SWSL: Semantic-aware Weakly Supervised Learning for 3D Motion Generation using 2D Motion Data
TASE: Truncation-Aware Semantic Embeddings for 3D Scene Understanding and Editing
FreqPhys: Repurposing Implicit Physiological Frequency Prior for Robust Remote Photoplethysmography
Towards Unsupervised Multi-modal Semantic Segmentation
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
Identity-Preserving Human Reconstruction from a Single Image via 3D Token Inference
Video-Holmes: Can MLLM Think like Holmes for Complex Video Reasoning?
ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
Vision Bridge Transformer at Scale
HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
WorldFlow3D: Flowing Through 3D Distributions for Unbounded World Generation
OpenPanoD: Aligning Multimodal Prompts and Spherical Representations for Open-Vocabulary Panoramic Detection
Φeat: Physically-Grounded Material Feature Representation
340 FPS Reflection-free Video from Spikes Modulated by a Rapidly Rotating Polarizer
AdaptiveSplat: Texture Aware Controllable 3D Gaussian Allocation for Feed-Forward Reconstruction
AirSplat: Alignment and Rating for Robust Feed-Forward 3D Gaussian Splatting
ART-VSR: Adaptive Rectified Trajectories for One-Step Video Super-Resolution
BeTTER: Diagnose the Illusion of Embodied Reasoning in Vision-Language-Action Models
BiCE-HG: A Bi-Conditional Egocentric Hand Gesture Dataset for Intelligent Reality Systems
CoGoal3D: Collaborative 3D Object Detection with 3D-Aware Fusion and Refinement
Comprehensive Robustness Analysis of LiDAR-based 3D Object Detection in Autonomous Driving
EgoGVAE: Ego-body Mesh Reconstruction via Guided Variational Autoencoder
Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
EVAR: Edge Visual Autoregressive Models via Principled Pruning
FEEL (Force-Enhanced Egocentric Learning): A Dataset for Physical Action Understanding
Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints
FlowFace: Rectifying Identity Conditioning with Riemannian Geometry for Face Generation
Follow-Your-Mind: Towards Inversion-Free Brain-Driven Visual Context Synthesis and Editing
FoundYou: A Unified Model for Personalized Segmentation and Retrieval
FSD-Net: Foundation-Guided Spatiotemporal Distillation for Video Polyp Segmentation
Fully Rotation-Equivariant Spectral-Spatial Learning for Multispectral Object Detection
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
Generalization and Memorization in Rectified Flow
Generative Refinement Network for Visual Synthesis
GMODiff: One-Step Gain Map Refinement with Diffusion Priors for Efficient HDR Reconstruction
Group3D: MLLM-Guided Semantic Grouping for Open-Vocabulary 3D Object Detection
HRDiT: Training-Free High-Resolution Image Generation with Off-the-Shelf Diffusion Transformer Models
Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment
Know3D: Prompting 3D Generation with Knowledge from Vision-Language Models
LaGen: Towards Autoregressive LiDAR Scene Generation
Mapping the Concept Landscape: Structural Perception of Global Distributions for Transparent Data Pruning
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments
MVGS: Multi-view Regulated Gaussian Splatting for Novel View Synthesis
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
On-Policy Diffusion Reinforcement Learning Meets Off-Policy Quality Anchoring
Online Reasoning Video Object Segmentation
PASTEL: Panoramic Alignment for Monocular 4D Scene Reconstruction
Pol-CACTI: A System and dataset forHigh-Speed Polarized Video Compressive Imaging
Progression as Latent Drift: Generative Forecasting of Slow-Evolving Pathologies
Reinforcing Vision-Language Models for Image Quality Assessment with Grounding Process Rewards
ProtoFair: Fair Self-Supervised Contrastive Learning via Pseudo-Counterfactual Pairs
Puppet-CNN: Continuous Parameter Dynamics for Input-Adaptive Convolutional Networks
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection
ReynoldsFlow: Physics-Inspired Spatiotemporal Flow Representation for Video Understanding
Rotate Your Character: Revisiting Video Diffusion Models for High-Quality 3D Character Generation
Scalable Cross-embodiment Dexterous Grasping via Morphology-Prior Diffusion
SHIFT: Motion Alignment in Video Diffusion Models with Adversarial Hybrid Fine-Tuning
SIMON: SImultaneous Multi-Object Navigation
SRRA: Stable-Rank-Based Residual Adaptation for Generalizable Deepfake Detection
TETO: Tracking Events with Teacher Observation for Motion Estimation and Frame Interpolation
TraversRL: Traversable Pedestrian Pathway Generation With Reinforcement Learning
UniGeo: Unifying Geometric Constraints for Camera-Controllable Image Editing via Video Priors
UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation
VideoSfM: Exploiting Temporal Structure for Video-Based Structure-from-Motion
WildCity: A Real-World Dataset for City-Scale Rendering and Beyond
A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models
Evaluating and Understanding Model Editing for Medical Vision Language Models
Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO
Allo{SR}2: Rectifying One-Step Super-Resolution to Stay Real via Allomorphic Generative Flows
V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models
BitRIC: Efficient Neural Compression of LiDAR Range Images via Hierarchical Bitplanes
Ceptor: Vision-Language Model-Infused Diverse Guidance for Detecting Anything
Continuous Heart Rate Variability Estimation from Egocentric Systems for Skill Assessment
Continuous Speculative Decoding for Autoregressive Image Generation
Cross-View Yaw Estimation in Location Uncertainty with Line-Aligning Yaw Scoring
CubicSplat: Differentiable Vector Graphics via Error-Bounded Forward Relaxation
CURE: Contextual Debiasing and Unbiased Refinement for Training-Free Open-Vocabulary Semantic Segmentation
Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement
EM3M: An Electron Micrograph Dataset for Microstructural Segmentation and Generation
Exploring Efficient Reasoning Segmentation with Small Language Models
FAIR: Feature-Augmented Implicit Regularization for AI-generated Fake Image Detection
From Masks to Pixels and Meaning: A New Taxonomy, Benchmark and Metrics for VLM Image Tampering
Geometric Regularization for Long-Tailed Semi-Supervised Learning via Gaussian Feature Bridges
HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding
Hypothesis Graph Refinement: Hypothesis-Driven Exploration with Cascade Error Correction for Embodied Navigation
ICLAgent: Integrated Circuit Footprint Geometry Labeling via LMM-empowered Multi-Agent Framework
Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision
Lifting Ego World Models for Planning and Control
Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs
Manifold-Aware Spectral Compaction: A Graph Signal Processing Perspective on Online Gaussian Reduction for 3DGS SLAM
MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution
MG2-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation
MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models
MoE-KD: Your Teacher Model is Worth Mixture-of-Experts for Knowledge Distillation
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
Neutralizing Token Aggregation via Information Augmentation for Efficient Test-Time Adaptation
OmniPoint: Universal Monocular Metric Pointcloud from Any Camera
One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach
PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation
PixVOD: Pixel-Distributed Direct Visual Odometry and Depth Estimation
QuantV2X: A Fully Quantized Multi-Agent System for Cooperative Perception
RA-SOD: Reliability-Aware RGB-T Salient Object Detection under Modality Degradation
Reconstructing Dense Depth of Dark Scenes with Sparse LiDAR, Noisy Events, and Blurry RGB
Remembering Across Blocks: Topology-Conditioned Block-Progressive Memory for Skeleton-Based Action Recognition
RobustRDP: Advancing Reaction Diagram Parsing via Synthetic-to-Real Data Scaling and Robustness-Oriented Training
Consistent Video-to-Video Translation via Explicit Correspondences
SPECSIA: Stylization Dataset for Novel-View Enhancement in Drawing-based 3D Animation
StreamSpatial: A Benchmark and Framework for Streaming 3D Visual-Spatial Reasoning
StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring
The Telephone Game: Evaluating Semantic Drift in Unified Models
Through Van Gogh’s Eyes: Global Style Transfer with Diffusion Model
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration
Together, Then Apart: Balancing Alignment and Distinctiveness for Multimodal Survival Analysis
TOPA: Mitigating Concept Dominance in Diffusion Personalization via Target-Oriented Perturbation Augmentation
Towards Geometry-Grounded Dense Semantic Matching with VGGT Priors
Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models
Tri-Efficient Transfer Learning for Point Cloud Videos
Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild
VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection
VOCA: Visual Odometry with Codec Awareness
VVSim: A Large-Scale Aerial-Ground Dataset and Benchmark for Cooperative Perception
WALL-EVE: World Alignment with Rule Learning in Visual Environments
Weight Feedback Computes the Exact Jacobian Transpose in Modern Deep Networks
A Simple Baseline with Placement Prior for Point-Supervised Oriented Object Detection
AdaBridge-SR: Adaptive Bridge Matching for Real-World Image Super-Resolution
Amplify, Aggregate, and Adjust: VideoMAE-based Holistic-Subtle Aggregation for Micro-Action Recognition
AnaPFL: When Closed-Form Solutions Meet Generalizationand Personalization in Personalized Federated Learning
AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning
Bayesian Self-Attention with Local Pixel Correlations for Lightweight Denoising Transformers
ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation
CogniCred: A Dataset and Benchmark for Cognitive Credential Forgery Detection
DecepGPT: Schema-Driven Deception Detection with Multicultural Datasets and Robust Multimodal Learning
Delaunay Canopy: Building Wireframe Reconstruction from Airborne LiDAR Point Clouds via Delaunay Graph
Dual-Prior Guided Null-Space Learning with Mixture-of-Splines for Arbitrary Medical Slice Super-Resolution
Efficient RGB-T Object Detection via Sparse Cross-Modality Fusion
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
EgoSim: Egocentric World Simulator for Embodiment Interaction Generation
Ex-Sim(3)-Reg: 2D-3D Correspondence Pruning via Extended Sim(3) Registration
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation
Frozen CLIP Priors for Robust Self-Supervised Poisson Inverse Problems
FusionTrack: Collaborative Multi-Object Tracking with Arbitrary Multi-UAVs
G2FM: A Geodesic Flow Matching Framework with Geometric Prior for Category-Level 9-DoF Pose Estimation
Glance: Accelerating Diffusion Models with 1 Sample
H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks
HERO: Heterogeneous Evidential Robust Object-Level Collaborative Perception
Histocomponent-driven Universal Model for Virtual Immunohistochemistry Multiplex Staining via Joint Manifold Evolution
InceptionGS: Generative Bootstrapping for Large-Scale Gaussian Splatting under Unstructured View Sampling
InstaEdit: Instant Image Editing via Optimized Noise Prediction
InterEdit: Navigating Text-Guided Multi-Human 3D Motion Editing
Let ViT Speak: Generative Language-Image Pre-training
LUCE: Constrained Curve-Domain Guidance for Training-Free Low-Light Enhancement with Hue-Preserving Decoupling
NeuroRefiner: Morphology-Aware Multi-Agent Refinement for 3D Fluorescence Microscopy Neuron Segmentation
OARS: Process-Aware Online Alignment for Generative Real-World Image Super-Resolution
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Based Video Editing
Parametric SDF for Dynamic Surface Reconstruction
PointSplat: Compact Gaussian Splatting via Human-Centric Prediction
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
Recolour What Matters: Region-Aware Colour Editing via Token-Level Diffusion
Region-Aware Test-Time Scaling for Compositional Image Generation
Representation Alignment for Just Image Transformers is not Easier than You Think
SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection
Schroedinger’s Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics
Seeing as Humans Do: Learning from Motion to Segment Anything Without Supervision
SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking
Sim, Yet Same: Physics-Aligned Simulator as Zero-Shot Data Scaler in Deformable Worlds
SkillSpotter: Pose-Aware Multi-View Skilled Action Detection and Grading in Ego-Exo Videos
SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts
Synthetic Sub-Aperture Phase Augmentation for Demosaicing 2×2 Shared Microlens Sensors
TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers
To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion
TurboMPLE: Joint Infrared Turbulence Mitigation and Physical Fields Estimation via Mutual Progressive Layered Extraction
Unfold The World: Factorize 4D Properties in Reinforcing Spatial Understanding
Unleashing the Power of Large-Scale ViT in Zero-Shot SBIR: A Strong Baseline with Multi-Layer Feature Aggregation
Video-Oasis: Rethinking Evaluation of Video Understanding
VLA Knows Its Limits
Zero-Shot Quantization for Object Detectors using Off-the-Shelf Generative Models
Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time
Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
BioMTBee: Biologically Constrained Multi-View Template-Based 3D Reconstruction of Bumblebee
BLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular priors
Broadband Wide Field of View Imaging with Computational Mirrors
CL4D: Contrastive Language–4D Pretraining for Vision-Language Reasoning in Dynamic Scenes
Compact and Structurally Transparent Cervical Cytology with Geometry-Driven Features and Closed-Form Attention
CurveStream: Boosting Streaming Video Understanding in MLLMs via Curvature-Aware Hierarchical Visual Memory Management
Defending from GeoLocalization through Adversarial Road Trips
DeRA: Decoupled Representation Alignment for Video Tokenization
DisRM: Reward Modeling as Discriminative Prediction
Dynamic Cluster Data Sampling for Efficient and Long-Tail-Aware Vision-Language Pre-training
EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA
FeatTracker: Short- and Long-Range Temporal Feature Consistency for Robust Underwater Object Tracking
FedMental: Topology-Aware Federated Prototype Learning for Polymorphic Multimodal Psychiatry
Foundation Model Selection for Remote Sensing via a Constraint-Aware Agent
From Visual Primitives to Semantic Masks: Fine-Grained Visual-Linguistic Alignment for Open-Vocabulary Remote Sensing Image Segmentation
GEM: Generative Supervision Helps Embodied Intelligence
Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth
Geometry-Aware Style Transfer in 3D Gaussian Splatting
HIVE: Understanding Post Hallucination Reasoning in Vision Language Models
IACD: Iterative Adversarial Collaborative Detection via Dual-Perspective Blind Spot Discovery
JacobianAvatar: Temporally Consistent Semi-rigid Avatar Reconstruction from a Monocular Video
KATANA: Knowledge-Aligned Topology-Aware Neural Agents for RL-Driven Vision-Language Model Compression
LangLoc: “Tell Me What You See”
Learn Once, Edit Anywhere: Visual Direction Transfer for Diffusion Models
LVSPM: Long Sequence View Synthesis and Pose Estimation Model
MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
MMR-Bench: A Comprehensive Benchmark for Multimodal LLM Routing
MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training
MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving
NanoVSR: Towards Real-Time Video Super-Resolution on Edge Devices
Off the Planckian Locus: Using 2D Chromaticity to Improve In-Camera Color
Personalizing MLLMs via Reinforced Multimodal Reference Game
PLOT: Pseudo-Labeling via Object Tracking for Monocular 3D Object Detection
Proposal Score Realignment Guided by Semantic Completeness for Weakly Supervised Temporal Action Localization
Real-Time LiDAR Gaussian Splatting SLAM via Geometry-Aware Covariance Coupling
Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation
Reinforcing Video Reasoning with Focused Thinking
Repurposing Geometric Foundation Models for Multi-view Diffusion
RotateAttention : RoPE-Aware Rotation and Range Rectification for INT4 Quantized Attention in Video Generation
RT-RMOT: A Dataset and Framework for RGB-Thermal Referring Multi-Object Tracking
Scale3D: Autoregressive Modeling for Large Outdoor Scene Generation
SDSA: Shallow-Deep Squeezing Adapter for Vision-Language Models
SFKD: Spatial–Frequency Joint-Aware Heterogeneous Knowledge Distillation via Multi-Level Wavelet Spectral Interaction
SIMPLER: Efficient Foundation Model Adaptation via Similarity-Guided Layer Pruning for Earth Observation
SONIC: Spectral Optimization of Noise for Inpainting with Consistency
StreetForward: Perceiving Dynamic Street with Feedforward Causal Dynamics
SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation
Test Time Training for Long Videos via Frame Forgetting Network
Text-based Tactile Graphics Generation for the Visually Impaired
Towards Robustness against Typographic Attack with Training-free Concept Localization
VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition
WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment
Which Layer Causes Distribution Deviation? Entropy-Guided Adaptive Pruning for Diffusion and Flow Models
Who Does What and Where to Go: Orthogonal Alignment and Hierarchical Planning for Multi-Entity Trajectories
3D Field of Junctions: A Noise-Robust, Training-Free Structural Prior for Volumetric Inverse Problems
Ada-VNNs: Adaptive Equivariance for Vector Neural Networks
AnyGround3D: Towards Grounding Any 3D Object in the Wild via 2D-to-3D Lifting
Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis
BRepFacetGen: Reverse Engineering B-Reps By Generative Face Segmentation
Coarse-to-fine Contrast: A Hybrid Self-supervised Method for Non-rigid 3D Shape Matching
DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation
Debiased Textual Prompt Tuning for Enhancing Unknown Class Discovery
DefenseSplat: Enhancing the Robustness of 3D Gaussian Splatting via Frequency-Aware Filtering
Delayed Bidirectional Alignment via Disentangled Audio Semantics for Audio-Visual Segmentation
DiffRGD: An Inference-Time Diffusion Guidance Through Riemannian Gradient Descent
Divide and Align: Disentangled Vision-Language Learning for Text-Based Person Anomaly Search
Domain Generalization via Text-Anchored Information Bottleneck
Drift-AR: Single-Step Visual Autoregressive Generation via Anti-Symmetric Drifting
DroneFINE: Domain-Aware Parameter-Efficient Fine-Tuning of Vision-Language Detectors for Drone Images
EgoTraj: Real-World Egocentric Human Trajectory
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
Estimating Velocity and Spin of Spherical Objects from Rolling-Shutter Image(s)
EvDiff: High Quality Video with an Event Camera
Explicit Semantic–Spatial Alignment for Open-Vocabulary Object Detection
Extending a Large View Synthesis Model for Multi-view Panoptic Segmentation
From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion
FSDC-DETR: A Frequency-Spatial Domain Collaborative DETR for Small-Object Detection
Gaussian Volumetric Representation for Efficient Shear–Warp Visualization
GryphOne: Symbol-Aware Masked Diffusion for Structural Refinement in Offline Handwritten Mathematical Expression Recognition
HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization
i-Design: Step-by-Step Graphic Layout Design with Progressive Aesthetic Policy Optimization
Importance-Aware Low-Rank Distillation of Diffusion Transformers
Integrated Forward–Inverse Network for Reconstruction for Lensless Image Reconstruction
Jumping the Landing Phase: Noise Variance Matching Enables Accurate Few-Step Inversion
Learning Semantic-Robust Change Detection via Semantic-Invariant Self-Distillation
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving
MAVFusion: Efficient Infrared and Visible Video Fusion via Motion-Aware Sparse Interaction
Multi-dimensional Preference Alignment by Conditioning Reward Itself
MV-STRIDE: Enabling MLLMs to Master Multi-View Spatial Reasoning via Hierarchical Capability Modeling
PhysFlowNet: Learning Canonical Latent Manifolds via Spatio-Spectral Physics Priors for Underwater Object Detection
Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch
RadarGen: Automotive Radar Point Cloud Generation from Cameras
RefAlign: Representation Alignment for Reference-to-Video Generation
RePlan: Reasoning-Guided Region Planning for Complex Instruction-Based Image Editing
Rethinking IRSTD: Single-Point Supervision Guided Encoder-only Framework is Enough for Infrared Small Target Detection
Safe Responses Matter: Output-Aware Safety Guardrail Mitigate Over-Refusal in MLLMs
Self-Evolving MCP-GUI Agents via Automated Environment Generation and Experience Learning
SIGMA-Lane: Scale-pyramId Gated MAmba for Temporally Consistent Video Lane Detection
Spectral Gradient Orthogonalization Improves Differentially Private Training at Scale
Swap the Right Identity: Spatio-Temporal Preference Optimization for Identity Swapping
Towards Metric-Agnostic Trajectory Forecasting
Transferability Between Understanding and Generation in Unified Multimodal Models
Tricam-rPPG: A Multimodal Multispectral Dataset for remote Photoplethysmography
Unsupervised Pixel-Level Semantic Left-Right Understanding of In-the-Wild Images
V-HOLD: Stabilizing Flow Trajectories to Rethink the Edit–Preservation Trade-off
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting
When Rubrics Fail: Error Enumeration as Reward for Reference-Free RL Post-Training
AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision
AlphaRad: Grounded Zero-Shot Classification in Chest Radiology via α-Corrected Binary Cross Entropy and Factorized Latent Supervision
An Inverse-Adversarial and Difficulty-Adaptive Robust Vision-Language Model
Anomaly Factory 3D: A Modular Framework for Diverse Pseudo-Anomaly Synthesis in Unsupervised 3D Anomaly Detection
Less Tokens, Better Forecasts: Sparse Residual Routing for Efficient Weather Prediction
AutoPhyX: Automatic Text-Condition Physics Property Generation
Background Blurring Matters: Improving Visual Grounding by Merging Text-Irrelevant Tokens
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
CAR-MIL: Counterfactual Attention Regularization for Multiple Instance Learning
CaRe: Critical Parameter Rectification for Efficient Visual Modeling
Continuous Adversarial Flow Models
MegaFlow: Zero-Shot Large Displacement Optical Flow
Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency
Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision
DOGE: Differentiable Bézier Graph Optimization for Road Network Extraction
DTI: Dynamic Trajectory Initialization for Generative Face Video Super-Resolution
Early Estimation of Language to Latent Alignment in Diffusion Models
EcoVideo: Entropy-Orchestrated Video Generation Paradigm in Cloud-Edge Dynamics
Escaping the Low-Frequency Bias: Adversarial Frequency Perturbation for Generalisable Gaze Estimation
ESNE: Efficient Surface Normal Estimation for LiDAR Point Clouds with Sequential Modeling and Variability Guidance
Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control
GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine
Generative Manifold Distillation: Aligning Restoration Trajectories with the Natural Image Prior
Hierarchical Hyperbolic Representation Learning for Aerial-Ground Person Re-Identification
High-speed Imaging through Turbulence with Event-based Light Fields
InclusiveHuman-10K: Towards Inclusive Human Parsing Beyond the Intact-Limb Assumption
Inference-Time Scaling of Diffusion Models via Progressive Pruning Search
Instance Segmentation as Tracking: A New Paradigm for Multi-Small-Object Tracking with Event Cameras
Learning Sample-wise Rank-Aware Interpolation Weights for Composed Visual Data Retrieval
Learning to Mask: Cross-Modal Noise Modulation for Hallucination Mitigation in Multi-modal Large Language Models
MAC-Splat: Multi-Attribute Consistency for High-Fidelity Sparse-View Reconstruction
MobileOcc: A Human-Aware Semantic Occupancy Dataset for Mobile Robots
Scaling Whole-Slide Pathology Foundation Model Pretraining with Billions Off-the-Shelf Tokens
Obliviate: Erasing Concepts from Autoregressive Image Generation Models
Open Your Eyes: Benchmarking the Detection of Fabricated Realities and Weaponized Ethics in VLMs
PanoSAM2: Lightweight Distortion- and Memory-aware Adaptions of SAM2 for 360 Video Object Segmentation
PGCR: Pose–Geometry Coupled Reasoning for Image-to-Point Cloud Registration
PriorEye: Geospatial Visual Priors for End-to-End Autonomous Driving
R-ESC: Robustly Erasing Space Concepts via Stochastic Feature Remapping
R3RECON: Radiance-Field-Free Active Reconstruction via Renderability
RaPTGS: Render-Agnostic Post-Training Compression of 3D Gaussian Splatting
Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning
RoMa v2: Harder Better Faster Denser Feature Matching
Scaling Laws for Black-box Adversarial Attacks
SceneDiff: A Benchmark and Method for Multiview Object Change Detection
SegDiff: Segmented Trajectory Diffusion for Consistent and Adaptive Robot Manipulation
Self-supervised Garment Dynamics with Persistent Wrinkles
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
SplatCtrlA: Generalizable Single Image to Fully Controllable 3D Avatar
Streaming Dense Voxel Representations for 3D Occupancy Prediction
Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs
There and Back Again: A Flexible-Frame Transformer for Multi-Exposure Fusion
ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance
Towards Scalable Pre-training of Visual Tokenizers for Generation
UnderOneFacade: Worldwide Facade Semantic Segmentation Benchmark Dataset
VisCritic: Visual State Comparison as Process Reward for GUI Agents
VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context
When 3D Gaussian Splatting Recovers Real Surfaces
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens
360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
Benchmarking MLLMs on Mistake Recognition and Explanation in Single-Step Components of Cooking
Causal Yet Future-Aware: Dual-Path Temporal Modeling for Online Action Segmentation
CLIP-AUTT: Test-Time Personalization with Action Unit Prompting for Fine-Grained Video Emotion Recognition
CoIn: Comprehensive 2D-3D Inpainting with Gaussian Splatting Guidance
Ice Cloud Geometry Retrieval with Calibrated Uncertainty from Passive Satellite Imagery
CortexVideo: A Semantic-Spatial Dual-Anchor Framework for High-Fidelity fMRI-to-Video Reconstruction
DARL: Efficient Document-to-Markup Generation via Look-Ahead Diffusion Trajectory Sampling
DeCo: Zero-Shot Anomaly Generation through Decoupling and Recoupling
Difficulty-Conditioned Attribute-Specific Restoration for Low-Light Image Enhancement
Dual-Generalization-aware Minimization for Continual Fine-Tuning of Vision-Language Models
Exclusivity-Guided Mask Learning for Semi-Supervised Crowd Instance Segmentation and Counting
FreeGen: Feed-Forward Reconstruction–Generation Co-Training for Free-Viewpoint Driving Scene Synthesis
CLARITY: Medical World Model for Guiding Treatment Decisions by Simulating Context-Aware Disease Trajectories in Latent Space
GenHOI: Generalized Hand-Object Pose Estimation with Occlusion Awareness
GridFlow: Structured Latent Flow for Seamless City-Scale 3D Point Cloud Generation
Holo-Captioning: A Comprehensive Textual View of 3D Scenes
ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning
Interference-Aware Continual Vision–Language Learning via Instance-Level Expert Routing
IREU: Identity-Related Encoder-Only Unlearning for Customized Portrait Generation
KISS-GS: 3D Gaussian Splatting Compression Kept Simple
Learning from Primitive: Probing Visual Reasoning of LVLMs via Counting
Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks
MOOZY: A Patient-First Foundation Model for Computational Pathology
LightSTAR: Efficient Visual Document Retrieval via Lightweight Selection with Vision-Adaptive Refinement
Long-term Traffic Simulation via Structured Autoregressive Modeling
Zero-shot Depth from Defocus
HybridSim: A Physics–Learning Hybrid Digital Twin for mmWave Human Sensing
Masked BRep Autoencoder via Hierarchical Graph Transformer
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
ModuSeg: Decoupling Object Discovery and Semantic Retrieval for Training-Free Weakly Supervised Segmentation
Molmo-Point: Better Pointing for VLMs with Grounding Tokens
PARL-VLA: Pruning-Aware On-Policy Reinforcement Learning for Vision-Language-Action Model
Q-TriM: Question-Guided Tri-Modal Attention for Audio–Visual Question Answering
RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation
ReAL: Reference-to-Image (R2I) Aware Latent Diffusion for Image Super-Resolution
ReconDreamer-RL: Enhancing Reinforcement Learning via Diffusion-based Reconstruction
Reference-Free Quality Assessment for Virtual Try-On via Human Feedback
Reinforcement Learning for Multimodal Diffusion Language Models via Bidimensional Trajectory and Thought Optimization
SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows
SIMSplat: Language-Aligned 4D Gaussian Splatting for Driving Scenario Generation
SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
StrucTab: A Structured Optimization Framework for Table Parsing
TACO-Net: Topological Signatures Triumph in 3D Object Classification
Tuning-free Visual Effect Transfer across Videos
Unifying CNNs and ViTs for Learning-Efficient and Scalable Variational AutoEncoder
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
UniTriSplat: A Unified 3D Gaussian Splatting Framework with Uniform Spherical Rasterization for Universal Cameras
Vitality-Aware Compression for Efficient Image-to-Shape Diffusion Transformers
VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On
Wid3R: Wide Field-of-View 3D Reconstruction via Camera Model Conditioning
World-in-Loop: Online Correction via Event-Triggered World Models for Robust VLA Policies
ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device
Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints
E-M3RF: An Equivariant Multimodal 3D Re-assembly Framework
Improving Adversarial Robustness via Activation Amplification and Attenuation
InSeg: Interactive Refinement via Intent Propagation for Point Cloud Semantic Segmentation
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
MorphJEPA: Morphology-Aware Latent Prediction for Hyperspectral Images
Pixel-wise Planarity for High-Precision Monocular Plane Segmentation
Proximity-CLIP: Text-Guided Semantic Proximity Learning for Zero-Shot Anomaly Detection
Understanding Cross-Rig Generalization in Automotive Perception: a Multi-Rig Benchmark and Rig Variation Metrics
Wavelet-Guided Semantic Signal Compensation for Inversion-Free Image Editing
Learning to Recover Task Experts from a Multi-Task Merged Model
AlignMorph: Tuning-Free Diffusion Image Morphing via Explicit Semantic Transport
Physically Grounded Monocular Depth via Nanophotonic Wavefront Encoding
Physics-Guided Deep Learning for Linear Mueller Matrix Acquisition
Beyond Aesthetics: Quantifying Information Loss in Turbid Scenes
Mitigating Pose–Scale Discrepancy Bias and Reforming Multi-Support Reasoning for Few-Shot Semantic Segmentation
DSeq-JEPA: Discriminative Sequential Joint-Embedding Predictive Architecture
MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos
Anchor Forcing: Anchor Memory and Tri-Region RoPE for Interactive Streaming Video Diffusion
Finding Highlight Images In Your Albums:From Benchmark To MLLM
SA-V2V: Training-Free Subject-Aware Video-to-Video Personalization
BlenderFusion: 3D-Grounded Visual Editing and Generative Compositing
Matryoshka Gaussian Splatting
CS-TTA: Preserving Concept Sensitivity in Test-Time Adaptation
FED-Bench: A Cross-Granular Benchmark for Disentangled Evaluation of Facial Expression Editing
ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
TAIHRI: Task-Aware 3D Human Keypoints Localization for Close-Range Human-Robot Interaction
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors
Large-Scale Light Field Synthesis from Videos Enables Geometrically Consistent Bokeh Editing
TEX-Drive: Temporal Perception Meets Experience-Guided Mixture-of-Experts for End-to-End Autonomous Driving
Unified Multi-Layer Subspace Modeling for Cross-Domain OOD Detection
PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
Score-Based Matching with Target Guidance for Cryo-EM Denoising
PanoLess: Environment Reconstruction from Partial Reflective Views
Multi-View Foundation Models
MambaRaw: Selective State Space Modeling for Efficient 4K RAW Image Reconstruction
From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting
RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
This Looks Distinctly Like That: Grounding Interpretable Recognition in Stiefel Geometry against Neural Collapse
Adversarial Attack and Disturbance Detection by Hadamard-Coded Output Representations for Object Detection and Semantic Segmentation
Mitigating Radar-Inertial Calibration Ambiguities via SO(3) Manifold Steering
Incentive Noise and Structural Prior Infusion for Multi-Modal Object Re-Identification
What You Ask is What You Ground: Bridging Question Intent to Temporal Evidence for Grounded VideoQA
Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models
Same Person, Different Depiction: Counterfactual Evaluation of Vision-Language Models on Individuals with Limb Deficiencies
IndoorSplat: Enhanced Indoor Scene Reconstruction with Structured 2D Gaussian Splatting
VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Degradation-Agnostic Clarity Learning for Unpaired Image Dehazing
Disentangling and Reusing Interaction Cues for Zero-Shot HOI Detection
Vero: Open Reinforcement Learning Recipes for Visual Reasoning
Uncertainty-Driven Gaussian Sphere Propagation for 3D Semantic Segmentation
Y-diff: Structure-Texture Decoupled Diffusion Distillation for H&E-to-pCLE Translation
VisTa3D: A Dataset and Benchmark for Vision, Tactile, and 3D Point Clouds-based Thin Object Reconstruction
OmniFit: Multi-modal 3D Body Fitting via Scale-agnostic Dense Landmark Prediction
Grasp-Oriented Non-Prehensile Manipulation via Learning a Graspability Field
High-Throughput Event-Based Feature Detection and Tracking on an Embedded CPU
PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics
TimeWalker: Personalized Neural Space for Lifelong Head Avatars
Multiple Images Distract Large Multimodal Models via Attention Fragmentation
SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers
Penetration-Free Compositional 3D Generation via Gaussian Surface Offset
HorizonRelight: Relighting Long-horizon Videos Consistently via Diffusion Transformers
Learning Active Perception for Pixel-Space Reasoning via Visual-Intent Stratified GRPO
MCVL: Multi-Space Cross-View Learning for Aerial-Ground Person Re-Identification
GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction
AdaDexGrasp: Adaptive Dexterous Grasping via 3D Visuo-Tactile Representation Fusion
TextFace: Compositional Text-Guided Identity Preserving Face Synthesis for Face Recognition
Improving Adversarial Robustness by Mitigating Instability through Relearning
Visual Spatial Tuning
From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology
SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation
Towards Practical Lossless Neural Compression for LiDAR Point Clouds
MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
Exposing Implicit Vulnerabilities in Text-to-Image Models via Adversarial Agentic Probing
3D-LENS: A 3D Lifting-based Elevated Novel-view Synthesis method for Single-View Aerial-Ground Re-Identification
FitControler: Toward Fit-Aware Virtual Try-On
EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents
Closing the Capacity–Convergence Gap: Globally Optimal Configuration of Implicit Neural Representations
On the Faithfulness of Post-Hoc Concept Bottleneck Models
AViTS:Adaptive Spatiotemporal Token Selection for Efficient Dynamic-Resolution Generation
GKDT: General Keypoint Detection Transformer
TaxoGrasp: Taxonomy-Guided Human Grasp Synthesis with Sparse Contact Constraint
AdaBoosting Text Prompts for Vision-Language Models
Towards Robust Driving Perception: A Flexible Scale-Driven Family for Self-Supervised Monocular Depth Estimation
SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark
UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters
Measuring 3D Spatial Geometric Consistency in Dynamic Video Generation
CanoVerse: 3D Object Scalable Canonicalization and Dataset for Generation and Pose
Implicit Neural Representation for Spherical Harmonics Reconstruction of Motion-Corrupted Fetal Diffusion MRI
DiverseAD: A Large-Scale Driving Dataset with Diverse Atmospheric Conditions
DE2TR: Dual Evidence Detection Transformer for Video Temporal Grounding
Causal Intervention in Concept Bottleneck Models
ARGENT: Adaptive Hierarchical Image-Text Representations
DICE: Disentangled Instance-Class knowlEdge prompt tuning via SAE for Vision-Language Models
Verifying Cancer Segmentation in Vision Transformers via Internal Concepts
LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?
Source-Agnostic Image Translation Based on Latent Aware Adaptive Masking
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
Steering 3D Generations: Preference Alignment via Direct Reward and Preference Optimization
Mechanistic interventions for explainable digital pathology uncovers adversarial vulnerabilities
PaD-GS: Leveraging Distortion Map for Panoramic Gaussian Splatting
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding
TIDES: Time-Derivative Event Simulation via Deformable Reconstruction
Token-Based Affordance Grounding with Large Vision-Language Models
Editing Everything Everywhere All at Once
SpectralSplats: Robust Differentiable Tracking via Spectral Moment Supervision
Circuit-MLLM: Topological Logic-Guided Latent-Space Visual Reasoning for Circuit Schematic Understanding
ProGVC: Progressive-based Generative Video Compression via Auto-Regressive Context Modeling
LiDAR-EVS: Enhance Extrapolated View Synthesis for 3D Gaussian Splatting with Pseudo-LiDAR Supervision
LogFA: Efficient Feature-Space Data Augmentation for Egocentric Temporal Action Segmentation
REAL-OW: Rehearsal-free Open World Object Detection with Low-Rank Adaptation and Dual-Stage Objectness Modeling
Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?
Towards Reconfigurable Visual Feature Compression
QVAM: Query-guided View-aware Adaptive Modulation for Aerial-Ground Person Re-Identification
MoMCE: Mixture of Modality and Cue Experts for Multimodal Deception Detection
Benchmarking Federated Learning & Knowledge Distillation for Point Cloud Classification
LARY: A Latent Action Representation Yielding Benchmark
SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation
Towards Effective Long Video Understanding: Dynamic MAS Construction via Meta-Agent
Unbalanced Optimal Transport for Efficient Visual Document Retrieval
Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval
Leveraging Cross-Modal Knowledge Transfer for Knowledge-Aware Concept Customization
DisentangledTMR: Privacy-Preserving Skeleton Motion Retargeting via Factorized Transformers
MeanTalker: Efficient and Expressive Speech-Driven 3D Facial Animation via Geometric-Aware Mean Flow
City-Level 3D Surface Reconstruction with Viewpoint Orientation Partitioning and Scene Completion
Dense Video Understanding with Inter-tokenization Acceleration
RoboClaw: An Agentic Framework for Scalable Long-Horizon Robotic Tasks
OmniColor: A Unified Framework for Multi-modal Lineart Colorization
Geometry-Preserving in 3D Gaussian Splatting for LiDAR-Camera Extrinsic Calibration
ASSCG: Just-Right Gating over Chattering for Fast–Slow LLM Planning in Autonomous Driving
3D Gaussian Texture for Real-time Mesoscale Appearance Synthesis and Rendering
From Minimal Clinical Prompts to 3D: Spacing-Aware Prompt Propagation for Multimodal Prostate Lesion Segmentation in bpMRI
CHARTSTYLE-100K: A Large-Scale Dataset for Structured Visualization Style Transfer
Video Can Teach PAN-Sharpening: PSF-Aware Cross-Domain Supervision
What is the Right Embedding Space for Contrastive Learning in Referring Expression Counting?
Enlightening Photographic Style Transfer with a Self-Supervised Photographic Embedding
BackTranslation2.0 - A Linguistically Motivated Metric to Assess Sign Language Production
DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video-LLMs
Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction
PASDiff: Physics-Aware Semantic Guidance for Joint Real-world Low-Light Face Enhancement and Restoration
Control-DINO: Feature Space Conditioning for Controllable Video Diffusion
GroundingAnomaly: Spatially-Grounded Diffusion for Few-Shot Anomaly Synthesis
Reward Lightning: Fast Video Generation via Homologous Preference Distillation
AnyStyle: A Single LoRA is Sufficient for Image-Guided Style Transfer
Video Generation Models are General-Purpose Vision Learners
RCEdit-500K: Reference Completion for Image-Conditioned Image Editing
Weight-Space Mixture-of-Experts for Implicit Neural Representation Classification
Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
Last-Layer-Centric Feature Recombination: Unleashing 3D Geometric Knowledge in DINOv3 for Monocular Depth Estimation
SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning
Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes
GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models
EgoMAN: Interaction-Structured Reasoning for Egocentric 3D Hand Trajectory Prediction
Why Can Accurate Models Be Learned from Inaccurate Annotations?
Breaking High Confidence: Practical Face Impersonation under High-Security Thresholds
Bridging the Geometry Mismatch: Frequency-Aware Anisotropic Serialization for Thin-Structure SSMs
AIMold: An Autonomous AI-based Pipeline for Complex Mold Design
OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data
CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming
GhostPoint: Self-Supervised Representation Learning by Hallucinating Occluded LiDAR Structure
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
Trajectory-Level Continuous Action Representation for Robotic Manipulation
MoBa-GS: Learning a Spatially-Varying Motion Basis over a Dynamic Canonical Space for 4D Reconstruction
Data Circuit Breaker: Identifying Training, Test, and Generated Data in Image Generative Models
Seen2Scene: Completing Realistic 3D Scenes with Visibility-Guided Flow
RhymeFlow: Training Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling
MoAKE: Toward Unified All-in-One Action Quality Assessment via Mixture of Action Knowledge Experts
Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Generation
Linear Scaling Video VLMs for Long Video Understanding
SelfMOTR: Revisiting MOTR with Self-Generating Detection Priors
From Gaze to Meaning: An AI Agent for Unified Zero-Shot Grounding and Explanation
Embedding Rotation Invariance for Provable Multi-Oriented Scene Text Recognition
StructureGS: Structure-aware Gaussian Splatting for Articulated Object Reconstruction
Diffusion to Obfuscation: Time-Adaptive Synthesized Generation Against Gradient Leakage Attacks in Federated Learning
Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization
Defect-aware Hybrid Prompt Optimization for Zero-Shot Multi-type Anomaly Detection and Segmentation
Geometry-Preserving Image Generation for 6D Object Pose Estimation
SWAN: World-Aware Adaptive Multimodal Networks for Runtime Variations
Unwarping the Lens: A Physics-Grounded Approach to Video Glasses Removal
Trustworthy Image Authentication using Forensic Knowledge Graphs
Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task
Sentinel: Embodied Cooperative Spatial Reasoning and Planning
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
Multi-Channel Uncertainty-Weighted Score Matching for Conditional Diffusion in Medical UDA
InstGS: Shared-Template Gaussian Instancing for Object-Redundancy-Free Rendering
Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts
DSAR: Dual-Stream Autoregressive Modeling of Temporal Cloth Dynamics for Photorealistic Animatable Avatars
Provable and Robust Wavefront Sensing via Self-Reference Interferometry
STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation
Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation
Lost in the Tail: Addressing Geographic Imbalance in Urban Visual Place Recognition
Cross-token Guidance Transformer for Weakly Supervised Object Localization
SPHERE: From MRI Sampling Mechanisms to Spatial Priors for Generalizable Brain Tumor Segmentation
Enhanced Neural Video Representation Compression with High Scalability
Synesthesia via Direct Latent Augmentation: Bypassing the Decode-Encode Loop for Cross-Modal Distillation
SparkVSR: Interactive Video Super-Resolution via Sparse Keyframe Propagation
MMDiff: Extending Diffusion Transformers for Multi-Modal Generation
Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval
Real-Time Source-Free Object Detection
Adaptive Noise Covariance Scheduling under Riemannian Metrics for Diffusion Models
A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation
Combining Discrepancy-Confusion Uncertainty and Calibration Diversity for Active Fine-Grained Image Classification
Kilometer-Vision: A New Frontier for Large-Scale Spatial Awareness in VLMs
Table-MCR2TR: Merged-Cell-Aware Table Recognition via Reinforced Multimodal Language Models
Spatiotemporal Flux Probing for Single-Photon Videography
Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
YesTrack: Referring Multi-Object Tracking via MLLM-based Yes/No Reasoning
IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves
MonoSR: Open-Vocabulary Spatial Reasoning on Monocular Images
A²-Edit: Precise Reference-Guided Image Editing of Arbitrary Objects and Ambiguous Masks
BATQuant: Outlier-Resilient MXFP4 Quantization via Learnable Block-wise Optimization
Intrinsically Stable Spiking Neural Networks: Overcoming the Performance Barrier in the Absence of Batch Normalization
VC-VAE: Leveraging Video Codecs for Training-Efficient and High-Fidelity Video VAE
Rolling Shutter Relative Pose Estimation Made Practical
SCALE: Semantic-Calibrated Guidance Enhancement for Prompt-Faithful Diffusion
CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving
Accelerating Text-to-Video Generation with Calibrated Sparse Attention
Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute
Multi-Modal Controlled Coherent Motion Generation
Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
In-Context Sync-LoRA for Portrait Video Editing
Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models
Ray-Path-Aware Virtual Point Removal on 2D Layer-Wise Nearest Point Map
UniStitch: Unifying Semantic and Geometric Features for Image Stitching
Setting the Stage: Text-Driven Scene-Consistent Image Generation
DynEval: Holistic Evaluations of T2I Generative Models in the Wild
∂DIBR: Differentiable Depth Image-based Rendering for Fast Novel View Synthesis
ObjectForesight: Predicting 3D Object Trajectories from Human Videos
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
TEASR: Training-Efficient Any-Step Diffusion Transformer for Real-World Image Super-Resolution
FLM-Occ: Feed-forward Likelihood Maximization for Efficient Indoor Occupancy Prediction
ROAR-3D: Routing Arbitrary Views for High-Fidelity 3D Generation
Horizon3D: Sparse Radar-Camera Fusion for Long-Range 3D Perception in Autonomous Driving
NAPA: Natively Multimodal Autoregressive Perception Architecture
Learning Global Camera Poses from Noisy View-Graphs for Structure from Motion
BIP: Bi-level Information Transfer and Completion Prompting for Visual Recognition with Missing Modalities
When Distillation Breaks Motion Control: Restoring Generative Trajectories for Fast Video Generators
High-Resolution Artwork Outpainting with Global Blueprint Guidance and Layout Control
Pretrained Video Models as Differentiable Physics Simulators for Urban Wind Flows
Denoised Variance-Based Pruning with Optimal Brain Bias Compensation
Syn4D: A Multiview Synthetic 4D Dataset
TiltDiff: Tilted Weight-Space Diffusion for Neural Network Generation
CausalDrive: Real-time Causal World Models for Autonomous Driving
DiffUE: Enhancing Utility-Unlearnability Trade-off of Unlearnable Examples via Diffusion Autoencoders
E-MOTION: A Dataset for Event-Based Scene Flow Estimation with Independent Moving Objects
SCDL: Synergistic Confidence-Dispersion Learning for Semi-Supervised Video Polyp Segmentation
Zero-Shot Inference-Time Rectification for Real-World Arbitrary-Scale Super-Resolution
Dense Dynamic Scene Reconstruction and Camera Pose Estimation from Multi-View Videos
GeoMix: Descriptor-Free Visual Localization via Global Context and Multi-Detector Training
3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement
Disentangling Hallucinations: Orthogonal Semantic Projection for Robust Interpretability
ORACLE-3D: Open-world Region-aligned Cross-modal Learning for Label-efficient 3D Scene Understanding
Multi-Anchor Distillation with Text-Guided Analytic Classifier for Continual Learning
LUA: Latent Upscaling Adapter for Diffusion-Based Image Synthesis
PriorPose: Reference-Guided Joint Deformation and Alignment for Category-Level Object Pose Estimation
Q-BridgeNet: A Quantization Network for Cross-Lingual Sign Language Translation
DiNBV-Grasp: Real-Time Distance-Aware Two-Stage Next-Best-View for Robotic Grasping
Unified Removal of Raindrops and Reflections: A New Benchmark and A Novel Pipeline
LoMa: Local Feature Matching Revisited
h-Flow: Flexible Flow-based Image Editing via Doob's h-Transform
Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning
Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
LlamaSeg: Image Segmentation via Autoregressive Mask Generation
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA
SFD-Net: Sharp Feature Detection Network Based on Local Geometric Features
Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption
Plug-and-Play Attention Linearization for Pretrained Transformers
Walking in the Implicit: Interactive World Exploration via Neural Scene Representation
RTE-FM-Dehazer: Radiative Transfer Equation Inspired Flow Matching for Real-World Image Dehazing
Back-Tracking from Clarity: Self-Learning to See Text from Afar
PercepTax: Benchmarking Cross-Property Reasoning in Vision-Language Models
GoStop: Reinforcement Learning for Adaptive Temporal Aggregation in Event-Based Feature Tracking
Attribute Token Arithmetic: Disentangled and Continuous Semantic Control for Visual Autoregressive Models
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
DRPO: Disentangling Demographic Bias from Rewards for Fair Diffusion Alignment
When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition
Anchoring on Reality: Breaking the Pseudo-Target Ceiling in Makeup Transfer
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
Temporally Stable Generative Illumination with a One-Step Diffusion Model
ConfCtrl: Enabling Precise Camera Control in Video Diffusion via Confidence-Aware Interpolation
Compositional Non-Face Re-Identification Pressure under Cumulative Vision Releases
ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
GeoCFM: Positive-Only Conditional Flow Matching for Mineral Occurrence Sampling
Fast and Accurate Image Restoration with Rank Enhanced Linear Attention
From Noise to Events: Conditional Diffusion for Event Data Augmentation
Learning Manifolds in High-D Point Embedding for Anisotropic Surface Approximation from Unstructured Point Clouds
TRaM-VSR: Importance-Aware Token Routing and Merging for One-Step Diffusion Video Super-Resolution
GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision
Tempo-SAM3D: Monocular Video to 4D via Temporal Memory-Guided Generation
360Anything: Geometry-Free Lifting of Images and Videos to 360°
CMDR: Contextual Multimodal Document Retrieval
EDM:Event-guided Diffusion Model for Video Shadow Detection in Complex Dynamic Scenes
EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models
Prevention over Correction: Learning Aligned Representations in One-shot Federated Learning
FMA-Net++: Motion- and Exposure-Aware Joint Video Super-Resolution and Deblurring
X-SG2S: Safe and Generalizable Gaussian Splatting with X-dimensional Watermarks
IRIS: Intersection-aware Ray-based Implicit Editable Scenes
S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation
Reliability-Aware 3D Geometric Injection for Universal Person Re-identification
Geo-DPO: Aligning Semantic Intent with Geometry for 3D Affordance Segmentation
Masked Depth Modeling for Spatial Perception
Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment
OmniLife360: A Benchmark for 3D Reconstruction from In-the-Wild 360° Captures
DRIFT: Difficulty-aware Rectified Flows for Through-plane MRI Super-Resolution
UF0-6D: Unified Flow-based Zero-Shot 6D Object Pose Estimation without Refinement
Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
iSyncTab: Learning Cross-Modal Feature Sequencing for Image-Tabular Data via Neural Synchrony
MedQ-Deg: A Multidimensional Benchmark for Evaluating MLLMs Across Medical Image Quality Degradations
4D-VGGT: A SpatioTemporal Foundation Model for Dynamic Scene Geometry Estimation
Rdm: Re-conceptualizing Distribution Matching as a Reward for Diffusion Distillation
Taming LLMs for Codematic Indoor Scene Generation
DAP: Doppler-aware Point Network for Heterogeneous mmWave Action Recognition
Variational Patch Gating for Training-Free Few-Shot Classification
Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback
Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
FoundDP: Revisiting Weak Disparity Observability in Dual-Pixel Depth Estimation
Decoupling Moment from Event for Video Temporal Grounding
Bounding-Box Trajectories Matter for Video Anomaly Detection
Probing the 3D Object-Level Understanding of Pre-Trained Detection Transformers
RegVGGT: Sustainable Visual Geometry Grounding for Streaming via Regulated Memory
Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings
Open-Vocabulary Long Term Action Anticipation
Geometric Probing for Isotropic Optimization Manifold in Sparse-View 3D Gaussian Splatting
LineGraph2Road: Structural Graph Reasoning on Line Graphs for Road Network Extraction
Seeing Through the Weights: Privacy Leakage in Scene Coordinate Regression
DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos
Multi-Block-Attention-based Color Constancy
DR-GS: Physically-Based Deformable and Relightable 2D Gaussians
ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation
Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
Vector Scaffolding: Inter-Scale Orchestration for Differentiable Image Vectorization
Predictive Structure Improves Video Diffusion Dynamics
MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation
CORE-V: Chain-Of-thought REasoning for Image Editing with Visual Interaction
EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning
GARDEN: Gravity-Aligned Reconstruction of Disentangled ENvironments from RGB images
SAFER-Activities: A Dataset for Smart Assessment of Fall Events and Routine Activities
InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation
Unsupervised Source-Free Ranking of Biomedical Segmentation Models Under Distribution Shift
E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation
On the real-world generalisability of Optical Flow models
Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning
MetaPoint: Unlocking Precise Spatial Control in Visual Generation
One Video, One World: Turning Monocular Video into Physical 4D Scenes
HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training
Spanning Tree Autoregressive Visual Generation
Accelerating Multimodal Large Language Models with Prior-Corrected Token Reduction
BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure
HVGCD:Rethinking Generalized Category Discovery through Hypothesis–Verification
TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration
Learn2Fold: Structured Origami Generation with World Model Planning
TORA: Topological Representation Alignment for 3D Shape Assembly
Analytic Bayesian Uncertainty for LiDAR Segmentation: A Single-pass Generative Approach
PartCHOI: Part-Aware Guidance for Clothed Human-Object Interaction Generation
Seeing Isn't Orienting: A Cognitively Grounded Hierarchical Benchmark for Object Orientation in MLLMs
Prior-Conditioned Gaussian Discriminants for Generalizable AI-generated Image Detection
ELDiff: When Evidential Learning Meets Text-to-Image Diffusion
YeTI: You Only Need Two Noisy Images for Real-World sRGB Noise Generation
WiFlow: Estimating Optical Flow using WiFi Channel State Information
Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models
Drop-In Perceptual Optimization for 3D Gaussian Splatting
Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer
Physics-Grounded Disentangled Flow Modeling for Brain Disease Progression Trajectory
Don’t Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance
Blind to Position, Biased in Language: Probing Mid-Layer Representational Bias in Vision-Language Encoders for Zero-Shot Language-Grounded Spatial Understanding
From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP
ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion
Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Synthesis
Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
Poppy: Polarization-Based Plug-and-Play Guidance for Enhancing Surface Normal Estimation
Prefill-Time Interventions against Adversarial Attacks on Large Vision-Language Models
ID-PreFeR: ID-Preserving Face Restoration with Mixed Data Quality
What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility
3D FaceShell: Attribute Transfer in 3D Face Avatars as a VLM Defense Mechanism
Warp-free Cross-view Geo-localization via Feature-space Consensus Mining
SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers
Taming Camera-Controlled Video Generation with Verifiable Geometry Reward
NeLU3D: Neural Inverse Structured Light without Modeling the Projector
Mode-Conditioned Residual Calibration for Multi-Object Tracking
Seek to Segment: Active Perception for Panoramic Referring Segmentation
MemLearner: Learning to Query Context Memory for Video World Models
RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
GeoV2V: Geometry-Grounded Video Diffusion Model for Driving Scene Generation
OmniSch: A Multimodal PCB Schematic Benchmark For Structured Diagram Visual Reasoning
AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in the Real World
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
Unified Panoramic–Gaussian Representation for Monocular 4D Scene Synthesis
Quick ViTs: Speeding up Vision Transformers through Equivariance
CRISP: Calibration-Aware Visual State Space Duality for Remote Sensing Image Segmentation
VCBench: A Streaming Counting Benchmark for Spatial-Temporal State Maintenance in Long Videos
SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
PIAvatar: Physically Interactive Avatars via Deformation Gradient Decoupling
SlowBA: An efficiency backdoor attack towards VLM-based GUI agents
Pseudo-Stereo Inputs: A Solution to the Occlusion Challenge in Self-Supervised Stereo Matching
Learning Physics-based Forward Model Corrections in Unrolled Networks for Diffuser-based Imaging
VERTIGO: Visual Preference Optimization for Cinematic Camera Generation
Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models
Mapping Dark-Matter Clusters via Physics-Guided Diffusion Models
Analyzing and Improving Training-Free Fast Sampling of Text-to-Image Diffusion Models
EGM: Efficient Visual Grounding Language Models
SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
Head Avatars with Dynamic Explicit Hair
Fast Sam 3D Body: Accelerating SAM 3D Body for Real-Time Full-Body Human Mesh Recovery
RECO: Region-Aware Compensation for Extrinsic Perturbations in Roadside 3D Detection
M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models
Accelerating Diffusion Models via Equal-Risk Caching
RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild
SKEL-CF: Coarse-to-Fine Biomechanical Skeleton and Surface Mesh Recovery
The Sterkfontein Caves Dataset: A Novel View Rendering Challenge from the Cradle of Humankind
ECC: Encoder-Centric Corruption for Fine-Grained Vision in VLMs
Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning
GenSP: Consistent Spherical Parameterization via Learning Shape Generative Models
Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models
History-Aware Transformation of ReID Features for Multiple Object Tracking
Staying VIGILant: Mitigating Visual Laziness in MLLMs via Information-Theoretic Alignment
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
IoUCert: Robustness Verification for Anchor-based Object Detectors
Detect by Track: Making Detector-Free Matcher Trackable
Motion Style Slider: Endpoint-Supervised Continuous Style Control for Human Motion Diffusion
Neural Collapse-Inspired Multi-Label Federated Learning under Label-Distribution Skew
POET: Preference Optimization for Enhanced Text-to-Image Generation
GO-Renderer: Generative Object Rendering with 3D-aware Controllable Video Diffusion Models
TetraSDF: Analytic Isosurface Extraction with Multi-resolution Tetrahedral Grid
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer
SPICE: Simple Polysemantic feature Interpretation via Clustering-based Explanations
SyncFix: Multi-View Consistent Diffusion Refinement of 3D Reconstructions
EgoEverything: A Benchmark for Human Behavior–Inspired Long-Context Egocentric Video Understanding in AR Environment
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
Geometrically Consistent Multi-View Scene Generation from Freehand Sketches
The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting
Anchored, Not Graded: How Vision-Language Models Fail at Slant-from-Texture Perception
ROSE: Real-Time Open-World Scene Understanding from Monocular Video via Compact Multimodal 4D Scene Graphs
Interaction-Aware 4D Gaussian Splatting for Dynamic Hand-Object Interaction Reconstruction
LUNA: Learning Universal 3D Human Animation Beyond Skinning
Learning Implicit Constitutive Laws for Dynamic 3D Gaussian Splatting from Monocular Videos
FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control
Beyond Pixel Mimicry: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
LooseControlVideo: Directorial Video Control using Spatial Blocking
Complex-Valued 2D Gaussian Representation for Computer-Generated Holography
ReDesign: Recovering Editable Design Structures from Raster Images via Agentic Decomposition
Bridging Online and Offline Handwriting via Differentiable Physical Rendering
EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation
DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces
Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors
Learning to Stylize by Learning to Destylize: A Scalable Paradigm for Supervised Style Transfer
Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
Finite Difference Flow Optimization for RL Post-Training of Text-to-Image Models
SupGRPO: Enhancing GRPO with Matching-based Online SFT for Text Spotting
Making Partial-Label Datasets Easier: A Simple Yet Highly Effective Data Augmentation for Deep Partial-Label Learning
Rapidly Deploying On-Device Eye Tracking by Distilling Visual Foundation Models
Direct Autoregressive Diffusion Distillation via Error-aware Causal Pretraining
Expert Weaving: Marrying Masked AutoRegressive and Diffusion Models for Unified Image Restoration
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
Enhancing prompt-image alignment evaluations via cyclic mutual information maximization
Learning Structured Visual Compositional Representations for Weakly Supervised Referring Expression Comprehension
Can Vision Models Truly Forget? Mirage: Representation-Level Certification of Visual Unlearning
Beyond Disjoint Tasks: Towards More Natural Continual Learning for Vision-Language Models
SHReg: Strictly Rotation-Equivariant Point Cloud Registration via Spherical Harmonics
Evidence-Backed Video Question Answering
ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering
Experts-Guided Unbalanced Optimal Transport for ISP Learning from Unpaired and/or Paired Data
SEERBench: A Spatial Ego-Exo Reasoning Benchmark for MLLMs with a Simple Yet Effective Baseline
Event-Driven Video Generation
Plug-and-Play Traffic Element Awareness for End-to-End Autonomous Driving
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
Unordered Landmark Visual Navigation
R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation
Cooking beyond Frames: A Stereo Event Camera Dataset in the Kitchen
Boxer: Robust Lifting of Open-World 2D Bounding Boxes to 3D
DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation
Agent-OBJ: Prompt-Driven 3D Adversaries for Multi-Modal Perception
AeroVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control
Keep Your Friends Close, and the Right Neighbours Closer: Disaster-Conditioned Kernel-Regularized Graph Attention for Building Damage Classification
We use cookies to store which papers have been visited.
I agree
Successful Page Load
ECCV uses cookies for essential functions only. We do not sell your personal information.
Our Privacy Policy »
Accept
We use cookies to store which papers have been visited.
I agree