
DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors
DeWorldSG generates spatio-temporally robust 3D semantic scene graphs from RGB-D sequences. It estimates instance-level geometric 3D Gaussian distributions through depth-guided filtering and represents each object as a probabilistic 3D node. It aggregates spatiotemporal evidence across object pairs and refines relations using contextual priors from a world model (V-JEPA 2). Improves triplet recall by 77.4% and predicate recall by 23.2% over prior SoTA.
Seok-Young Kim, Abdelrahman Elskhawy, Taewook HaJul 1, 2026
Scene graphsRGB-DV-JEPA 2Jul 1, 2026