WiFi-JEPA: Self-supervised Learning for WiFi-CSI 3D Human Pose Estimation
Abstract
WiFi Channel State Information (CSI) enablesprivacy-preserving human pose sensing in camera-denied environments,but existing WiFi-based pose estimators often fail under environmentshifts and rely on costly camera-based annotation pipelines that limitscale. We propose WiFi-JEPA, a self-supervised framework that learnsCSI-native representations by predicting masked latent embeddings in-stead of reconstructing raw CSI signals that may contain hardware-specific artifacts. WiFi-JEPA makes three contributions: (i) CSI-specifictokenization and link masking tailored to the CSI tensor over channel,time, and link (C, T, L); masking entire Tx–Rx antenna links forces themodel to predict one spatial link view from others, capturing cross-linkcorrelations informative of 3D spatial structure. (ii) A ray-tracing CSIsimulation pipeline that generates diverse unlabeled CSI from random-ized geometric primitives, providing scalable pre-training data withoutpose annotations. (iii) State-of-the-art results on Person-in-WiFi-3D:WiFi-JEPA outperforms prior WiFi-CSI baselines on both single- andmulti-person 3D pose estimation under the same evaluation protocol.We also show that simulated CSI provides complementary pre-trainingsignal to real CSI, and that four vision-native SSL objectives degrade per-formance below training from scratch, whereas WiFi-JEPA consistentlyimproves downstream pose estimation.