Hierarchical Floorplan-Guided Vision-Language Exploration for Embodied Question Answering

Autonomous Robots Lab, Norwegian University of Science and Technology (NTNU)

Given the question “What is the microwave's color?”, the robot must explore the unknown environment to find the required visual information. Using online RGB-D observations, the HLEX-EQA incrementally builds a scene graph and reasons over the visual memory and a topological floorplan to select the appropriate exploration strategy. After locating the lunch room and inspecting the microwave, it answers the question with high confidence.

Abstract

Embodied Question Answering (EQA) requires an agent to explore a previously unseen environment, gather relevant information, and answer questions about the scene. Recent approaches leverage Vision-Language Models (VLMs) together with semantic maps or scene graphs to guide exploration. However, exploration is typically driven only by local observations, while structural priors about the environment remain largely unused. We propose HFLEX-EQA, a hierarchical EQA framework that combines online scene graph construction, VLM-based planning, semantic frontier exploration, and floorplan priors. The system incrementally builds a hierarchical scene graph and an open-vocabulary occupancy map from RGB-D observations, enabling a VLM to jointly reason over the scene graph, task-relevant visual observations, exploration history, and an estimated topological floorplan. Furthermore, we introduce a room-discovery strategy that leverages the floorplan and open-vocabulary frontier semantics to guide exploration toward semantically relevant yet currently unobserved room types. We evaluate HFLEX-EQA on the OpenEQA and ExploreEQA benchmarks and demonstrate deployment on a quadruped robot in real indoor environments. Our results demonstrate the benefit of combining VLM-based hierarchical planning with structural floorplan priors for the EQA task.

HFLEX-EQA Overview

Online RGB-D observations are used to incrementally construct a hierarchical scene graph containing rooms, semantically-enhanced frontiers, a navigational graph, objects and a metric-semantic mesh. The question, optional answer choices, scene graph, relevant visual memory, action history, and floorplan are fed to a VLM-based high-level planner, which determines whether the question can be answered or further exploration is required. For exploration, the planner selects one of three modes: (i) go_to_objects, (ii) explore_room, or (iii) find_room. A low-level semantic planner then selects frontier or viewpoint targets according to the mode and generates a path.

Results

Quantitative Results

Across all evaluated VLMs, HFLEX-EQA consistently achieves the highest success rate on both datasets. The largest improvements are obtained with the stronger language models, where HFLEX-EQA reaches up to 75.0% success rate on OpenEQA and 63.2% on ExploreEQA. These improvements demonstrate that combining a scene graph representation with hierarchical planning, leveraging a structural floorplan graph, enables us to achieve state-of-the-art (SOTA) results on the EQA task. In contrast, ExploreEQA generally requires substantially more planning iterations and longer trajectories, while GraphEQA typically terminates after fewer planning steps but at the cost of lower success rates. Although GraphEQA often needs fewer planning iterations, this is largely because it finishes exploration earlier. The higher success rates achieved by HFLEX-EQA indicate that the additional planning iterations are spent gathering more question-relevant information.

Qualitative Results

Real-world deployment of the proposed framework on a quadruped robot in two EQA episodes. The top figures show the constructed occupancy maps and scene graphs, together with the selected exploration mode and goal at each high-level planning iteration. The relevant room images are used to answer the question when the system has gathered enough information.

Video

BibTeX


@inproceedings{puigjaner2026hflexeqa,
    title={Hierarchical Floorplan-Guided Vision-Language Exploration for Embodied Question Answering},
    author={Gassol Puigjaner, Albert and Alexis, Kostas},
    booktitle={arXiv preprint}, 
    year={2026}
}