Theses and Dissertations

Date of Award

5-1-2026

Document Type

Thesis

Degree Name

Master of Science (MS)

Department

Computer Science

First Advisor

Dongchul Kim

Second Advisor

Erik Enriquez

Third Advisor

Haoteng Tang

Abstract

Large Language Models and their multimodal variants, Vision-Language Models (VLMs), can interpret natural language instructions and visual context. Reinforcement Learning (RL) methods such as Proximal Policy Optimization (PPO) have been shown to successfully learn stable quadruped locomotion. This work investigates whether a VLM integrated with an RL controller can control a quadruped robot to perform instruction-following navigation to a predefined target using text instructions, egocentric vision, and proprioceptive state.

Most VLM-based robotics work directly target prediction of robot actions from multimodal inputs. This approach often depends on large amounts of expert demonstration data, generated by controllers such as trained RL policies. Exploration of a VLM+RL approach, where the VLM provides high-level navigation guidance and the RL policy follows and handles low-level locomotion, remains limited.

We evaluated a VLM-based navigator that outputs heading adjustments and forward/stop commands to guide an RL controller for quadruped navigation. Results show that this hierarchical VLM+RL approach achieves high success rates, with failure rates below 10\% in unobstructed environments. In more complex environments, performance declined to a 66% navigation success rate, though results still demonstrated the potential of VLM-based navigation. Overall, this work demonstrates that a hierarchical VLM+RL framework can effectively enable quadruped navigation using egocentric visual input while identifying key areas for future improvement in obstacle avoidance reasoning and stop prediction.

Comments

Copyright 2026 Chen Yuan Wang. All Rights Reserved. https://proquest.com/docview/3371129015

Share

COinS