CAS's dynamic topology perception empowers embodied navigation! DGNav: breaking through the granular rigidity problem of visual-language navigation

Posted: 2026/3/16 | Category: Artificial Intelligence | by Jiankun Peng, Jianyuan Guo, Ying Xu, Yue Liu, Jiashuang Yan, Xuanwei Ye, Houhua Li, Xiaoming Wang | with Institute of Aerospace Information Innovation, Chinese Academy of Sciences, School of Electronics, Electrical Engineering and Communication, University of Chinese Academy of Sciences, Department of Computer Science, City University of Hong Kong | Ref Link

The granular rigidity of the existing VLN-CE topology planning method is the core bottleneck. The fixed geometric threshold leads to a disconnect between map representation and environmental complexity, while the static geometric edge weights cause navigational myopia and reduce instruction fidelity. The proposed DGNav framework achieves dynamic regulation of graph granularity through a scene-aware adaptive strategy, addressing the contradiction between computational redundancy in simple areas and sparse nodes in complex areas. By fusing visual, linguistic, and geometric information through a multimodal dynamic graph Transformer to reconstruct edge weights, it effectively suppresses topological noise and enhances the alignment of path planning with linguistic instructions. A large number of experiments have verified the superiority of DGNav on the R2R-CE and RxR-CE datasets. Compared to end-to-end and traditional explicit map methods, it performs better in terms of navigation accuracy, efficiency, and generalization ability, and achieves an optimal trade-off between navigation efficiency and safe exploration. The dynamic topology-aware mechanism represents an effective approach to addressing the VLN-CE problem. By integrating explicit spatial memory with multimodal semantic reasoning, it provides an irreplaceable robustness guarantee for VLN-CE, even in the era of large models.