ZTRS: Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer
Abstract
Human demonstrations are widely considered the corner-stone of end-to-end (E2E) autonomous driving despite human demon-stration’s scarcity for long-tail and safety-critical scenarios. Nonetheless,current E2E autonomous driving (AD) training paradigms continue torely on human demonstrations. Imitation learning (IL) requires humandemonstrations for training, whereas reinforcement learning (RL) hasemerged as a promising alternative to reduce this dependency. How-ever, most existing RL methods for E2E AD still rely implicitly on hu-man demonstrations. A pure rewards-based RL method can overcomethe need for human demonstrations, but general RL policy gradientmethods suffer from the cold-start problem. In this paper, we proposeZTRS (Zero-human demonstration end-to-end autonomous driving withTRajectory Scorer) — a complete RL-based E2E planning paradigmtrained solely on real-world images and rule-based rewards, entirely with-out human demonstration. Through our proposed Exhaustive PolicyOptimization (EPO), a policy gradient variant tailored for enumer-able trajectory actions and dense supervision, ZTRS enables the modelto generalize better to long-tail driving scenarios. We demonstrate thisgeneralization through our SOTA performance against IL approaches onboth long-tail Navhard and closed-loop HUGSIM datasets. Project page:https://zhenxinli.net/ZTRS/.