Skip to main navigation Skip to search Skip to main content

Words to Wheels: Vision-Based Autonomous Driving Understanding Human Language Instructions Using Foundation Models

  • Chanhoe Ryu
  • , Hyunki Seong
  • , Daegyu Lee
  • , Seongwoo Moon
  • , Sungjae Min
  • , D. Hyunchul Shim*
  • *Corresponding author for this work
  • Korea Advanced Institute of Science and Technology
  • Electronics and Telecommunications Research Institute

Research output: Contribution to conferenceConference paperpeer-review

Abstract

This paper introduces an innovative application of foundation models, enabling Unmanned Ground Vehicles (UGVs) equipped with an RGB-D camera to navigate to designated destinations based on human language instructions. Unlike learning-based methods, this approach does not require prior training but instead leverages existing foundation models, thus facilitating generalization to novel environments. Upon receiving human language instructions, these are transformed into a 'cognitive route description' using a large language model (LLM)-a detailed navigation route expressed in human language. The vehicle then decomposes this description into landmarks and navigation maneuvers. The vehicle also determines elevation costs and identifies navigability levels of different regions through a terrain segmentation model, GANav, trained on open datasets. Semantic elevation costs, which take both elevation and navigability levels into account, are estimated and provided to the Model Predictive Path Integral (MPPI) planner, responsible for local path planning. Concurrently, the vehicle searches for target landmarks using foundation models, including YOLO-World and EfficientViT-SAM. Ultimately, the vehicle executes the navigation commands to reach the designated destination, the final landmark. Our experiments demonstrate that this application successfully guides UGVs to their destinations following human language instructions in novel environments, such as unfamiliar terrain or urban settings.

Original languageEnglish
Title of host publicationIV 2025 - 36th IEEE Intelligent Vehicles Symposium
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages2200-2207
Number of pages8
ISBN (Electronic)9798331538033
DOIs
StatePublished - 2025
Event36th IEEE Intelligent Vehicles Symposium, IV 2025 - Cluj-Napoca, Romania
Duration: 2025.06.222025.06.25

Publication series

NameIEEE Intelligent Vehicles Symposium, Proceedings
ISSN (Print)1931-0587
ISSN (Electronic)2642-7214

Conference

Conference36th IEEE Intelligent Vehicles Symposium, IV 2025
Country/TerritoryRomania
CityCluj-Napoca
Period25.06.2225.06.25

Fingerprint

Dive into the research topics of 'Words to Wheels: Vision-Based Autonomous Driving Understanding Human Language Instructions Using Foundation Models'. Together they form a unique fingerprint.

Cite this