A Dual-Transformer Architecture with Cross-Attention for Multi-Camera View Recommendation
Josep Cabacas Maso ⋅ Carles Ventura ⋅ Ismael Benito-Altamirano
Abstract
for previously unobserved body regions; (ii) a two-stage canonical-to-pose-dependent architecture that bootstraps from sparse observations tofull pose-dependent Gaussian maps; (iii) a map-pose/LBS-pose decou-pling that absorbs multi-view inconsistencies from the generated data;(iv) a head/body split supervision strategy that preserves facial iden-tity. We evaluate on YouTube videos and on multi-view capture datawith significant occlusion and demonstrate state-of-the-art reconstruc-tion quality. We also demonstrate that the resulting avatars are robustenough to be animated with novel poses and composited into 3DGSscenes captured using cell-phone video. Our project page is available athttps://miraymen.github.io/ahoy/.
Successful Page Load