Tutorials
The Fifth Hands-on Egocentric Research Tutorial with Project Aria from Meta
James Fort
View full details
Efficient MLLM Inference via Approximated and Exact Computing
Wanghuan Wang
View full details
From Perception to Simulation: The Emergence of World Models in Multi-modal Reasoning
Yujun Cai
View full details
Post-Training Diffusion Models: Enhancing Capabilities, Control, and Alignment
Sayak Paul ⋅ Linoy Tsaban ⋅ Hila Chefer
Pre-trained generative models are built using massive, typically unlabeled corpora, enabling them to capture broad, generic knowledge across diverse domains. However, at inference time, we often aim to adjust and customize these models — to exert control, enhance specific capabilities, and align their behavior with user intent and preferences. Post-training techniques have therefore emerged as both a practical necessity and an accessible means of adapting these powerful, yet static, models. This tutorial surveys the state-of-the-art in post-training methods for diffusion models, analyzing their strengths, limitations, and areas of application. We conclude with a critical discussion on the boundaries of post-training — asking whether fundamental semantic malfunctions can truly be resolved without revisiting the pretraining process.
Show more
Building Large Video Generation Models: Data Processing, Architectural Insights, Optimization Methods and Evaluation Strategies
Vasilev Viacheslav
View full details
From Hallucination Detection to Adversarial Defense: A Unified Framework for LVLM and LLM Safety
Gianni Franchi
View full details
Successful Page Load