CoT-VLNBench: A Benchmark for Visual Chain-of-Thought Reasoning in Vision-Language-Navigation Robots
Chang Liu, Zheyuan Zhang, Xiao Zhao, Zhe Ren, Chufan Guo, Kuifeng Su, Linna Song, Luo Qingliang, Mingxu Zhu, Ruiteng Ji, YuHang Gao, Zhaolong Du
aaai
Research metadataShow detailsHide details
- Affiliations
- Not available
- Published
- 2026-03-17
- Processed
- 7/25/2026, 12:47:47 PM
- Analysis model
- gemini-2.5-flash
- Analysis status
- analyzed
- Local PDF artifact
- papers/pdf/2026/cot-vlnbench-a-benchmark-for-visual-chain-of-thought-reasoni.pdf
Summary
This paper introduces CoT-VLNBench, the first large-scale benchmark and dataset for visual chain-of-thought (CoT) reasoning in quadruped robot navigation. The benchmark features diverse indoor and outdoor scenes, multi-step navigation trajectories, and natural language instructions with fine-grained CoT reasoning traces. Alongside the benchmark, the authors propose CoT-VLN, a novel 7B Vision-Language-Navigation (VLN) model that integrates visual, linguistic, and reasoning modules. Empirical results demonstrate that CoT-VLN significantly outperforms existing non-VLM baselines and non-reasoning approaches on the new benchmark, highlighting the importance of CoT reasoning for interpretable and effective embodied navigation.
Problem
The paper identifies several bottlenecks in existing robot-centric datasets and benchmarks:
- Lack of support for vision-language navigation (VLN) and advanced reasoning: Most current robot-centric datasets, such as CODa, SIT, and JRDB, primarily focus on traditional 3D tasks like perception and prediction, lacking the necessary annotations and task structures for VLN.