SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery
Abstract
Urban uncrewed aerial vehicle (UAV) vision-language navigation (VLN) requires agents to follow instructions across extended urban spaces, inherently demanding long-term memory and geospatial grounding. However, scaling existing benchmarks remains difficult because of their reliance on costly reconstructed 3D assets, limiting geographic diversity and episode scale. To address this, we introduce SatNav, a scalable, long-horizon UAV VLN benchmark built from high-resolution satellite imagery. SatNav targets city-level navigation missions and uses satellite crops as approximations of UAV nadir views for visual observations. Through an automated cue-to-episode pipeline, SatNav constructs 118K episodes from 59 scenes across 18 cities, with an average trajectory length of 379 m. To stress-test long-horizon memory and geospatial reasoning, SatNav defines three task families: Boundary, Landmark, and Route, targeting loop progress tracking, landmark-based spatial grounding, and route following with counting cues. Benchmarking classical VLN agents and recent agents based on large vision-language models (LVLMs) on SatNav shows that city-scale navigation remains challenging. We further introduce SwiftVLN, a modular framework with switchable memory components, and conduct systematic memory-design ablations. Finally, satellite-to-UAV transfer experiments show that satellite-trained navigation models can operate on real-flight UAV observations, showing the practical relevance of SatNav. Our project page: https://eku127.github.io/SatNav/
Community
Hi everyone! We’re excited to share SatNav, our benchmark for long-horizon UAV vision-language navigation built from high-resolution satellite imagery.
Highlights:
- 118K navigation episodes across 59 scenes in 18 cities, with an average trajectory length of 379 m.
- Three task families—Boundary, Landmark, and Route—evaluating long-term memory and geospatial reasoning.
- Evaluations of classical VLN and large vision-language model agents, alongside SwiftVLN, our modular framework for studying memory designs.
- Satellite-to-real-UAV transfer experiments showing that satellite-trained models can operate on real-flight observations.
Code and datasets are available through the links below. We’d love to hear your thoughts on long-horizon navigation, memory design, and satellite-to-UAV transfer!
🤗 Episodes: https://huggingface.co/datasets/Eku127/SatNav-Episodes-v0.1
🤗 Satellite scenes: https://huggingface.co/datasets/Eku127/SatNav-Scenes-v0.1
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN (2026)
- Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation (2026)
- RiverVLN: Phase-Grounded Temporal Vision--Language Navigation for Unmanned Surface Vehicles (2026)
- AirGroundVLN: A Large-Scale Benchmark for Goal-Oriented Air-Ground Collaborative Vision-and-Language Navigation (2026)
- ARIES-Mission2: A Zero-Shot Vision-Language-Action Framework for Fast Large-Scale Aerial Mission Generation (2026)
- AirAnchor: Bridging Local and Global Spatial Information for Zero-Shot Aerial Vision-and-Language Navigation (2026)
- ForeFly: A Dual-Horizon World Action Model for Aerial Vision-Language Navigation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.31507 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 2
Eku127/SatNav-Scenes-v0.1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper