Across encoders, TrackGraph ViT-H achieves the highest synonym frequency on Replica, is second to HOV-SG on ScanNet++, and remains competitive on HM3D, where OVI-MAP with SigLIP-L performs best. Using the same VL encoder, TrackGraph achieves equal synonym scores as OVI-MAP’s ViT-H variant on HM3D.
Abstract
Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Vision-Language (VL) inference. We present TrackGraph, an online open-vocabulary system that maintains short-term 2D mask identity directly in the image stream before fusing segments into 3D. FastSAM masks and CLIP features are computed at sparse keyframes, while dense DINOv3 features are used to propagate masks at a high rate in between. The resulting tracked masks are fused into a class-agnostic 3D segment layer within a hierarchical scene graph, with 3D association handling tracking interruptions and long-term revisits. Compact multi-view CLIP embeddings enable open-vocabulary retrieval. Across Replica, ScanNet++, and HM3D, TrackGraph achieves competitive open-vocabulary segmentation and retrieval against state-of-the-art mapping methods, including the highest synonym frequency on Replica (0.50). On the same NVIDIA A100, it is 1.7× faster and uses 3.3× less GPU memory than ViT-H OVI-MAP. Real-world quadruped deployments demonstrate onboard scene graph construction and object search at 7.5Hz, while recorded drone data is used to test the method under aerial viewpoints.
Overview
FastSAM masks and CLIP features are computed at sparse keyframes, while dense DINOv3 features are used to propagate the masks and their identities through the image stream in between. We refer to each 2D temporally tracked mask identity as a source track. Source track masks are continually fused into a TSDF and mesh, where repeated observations reinforce consistent track identities that are then used to build persistent 3D segments.
Interrupted tracks and long-term revisits are reconciled via 3D geometric and visual association, while CLIP features are averaged per source track and stored as separate views in each 3D segment’s feature gallery. A text or image query is embedded into the same CLIP space as the stored features. The query interface returns a ranked list of segment nodes.
Results
We evaluate open-vocabulary segmentation and object retrieval on Replica, ScanNet++, and HM3D following OpenLex3D. Runtime and resource usage are evaluated on Replica, whose scenes have comparable durations and memory footprints and for which baseline measurements are available.
Quantitative Results
Open-Vocabulary 3D Object Retrieval
Among the ViT-H variants, TrackGraph leads in AP50 and AP25 on Replica but trails OVI-MAP in AP, and outperforms ConceptGraphs and HOV-SG in AP and AP50 on ScanNet++. On HM3D, TrackGraph trails all baseline methods with matched encoder on AP and AP50 despite competitive point-wise semantics and AP25.
Runtime and Resource Usage
On the same NVIDIA A100, TrackGraph is 1.7× faster and uses 3.3× less GPU memory than ViT-H OVI-MAP. The broader efficiency comparison includes ConceptGraphs and HOV-SG measurements on an NVIDIA RTX 3090, while TrackGraph, OVI-MAP, and FindAnything are evaluated on an NVIDIA A100. These efficiency results are not all measured under identical hardware and input-frame settings and should therefore be interpreted as indicative comparisons.
TrackGraph CPU memory depends strongly on voxel resolution: increasing the voxel size from 0.03 m to 0.05 m reduces peak RAM usage by more than one-third, with little change in synonym frequency but lower retrieval AP.
Real-World Experiments
Quadruped Exploration and Object Search
We evaluate the complete system on a quadruped robot in three large-scale indoor environments. Each experiment comprises two autonomous phases: exploration using a volumetric planner while constructing the open-vocabulary scene graph, followed by object search using text or reference-image queries. Across all three environments, all queried targets were successfully retrieved.
Aerial Mapping
We further demonstrate TrackGraph on recorded drone RGB-D data from three indoor environments at a search-and-rescue training site, complementing standard ground-based benchmarks with uncommon aerial viewpoints. The figure shows qualitative results from one environment.
Video
BibTeX
@article{hellesylt2026trackgraph,
title={{TRACKGRAPH}: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking},
author={Hellesylt, Peder Borge and Gassol Puigjaner, Albert and Alexis, Kostas and Stahl, Annette},
journal={arXiv preprint arXiv:2609.31005},
year={2026},
url={https://arxiv.org/abs/2609.31005}
}