From Machine Perception to Human Expertise
Abstract
Computer vision has made remarkable progress in understanding the visual world. Yet these advances have focused on expanding the capabilities of the machine as an observer. The next frontier is to move from AI that observes human activity to AI that enables people to build their own expertise.
Imagine learning a new skill---whether physical, creative, or practical---with an AI guide that has learned from vast amounts of video, understands precisely what you are doing, anticipates the effects of your actions, and most importantly, can identify what you should change to improve. Such an AI guide would meet the learner where they are, offering guidance tailored to their abilities and goals.
To realize this vision, we need advances in two fundamental areas: fine-grained understanding of human activity and the generation of personalized, actionable guidance. Toward that end, I will present our recent progress on models that ground skilled activity in the body and the 3D world to anticipate outcomes and plan toward goals, as well as models that translate this understanding into concrete language and visual feedback. I will also show how these ideas make visual instruction more accessible, including AI guides on smartglasses that transform how-to videos into proactive guidance for blind and low-vision learners. Together, these advances bring us closer to AI that does more than make us more productive—it empowers us to learn and improve.
Speaker