TrackGraph: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking

Norwegian University of Science and Technology (NTNU)

A quadruped builds a TrackGraph scene graph during exploration, then retrieves a soda dispenser, a backpack, and a printer using text and image queries.
TrackGraph constructs an online open-vocabulary 3D scene graph during autonomous robot exploration. 2D source tracks are fused into persistent 3D segments, enabling post-exploration natural-language retrieval and navigation without offline reconstruction or post-processing.

Abstract

Open-vocabulary 3D maps enable robots to reason about previously unknown environments using natural language. However, existing systems typically segment every incoming image, associate detections with persistent 3D segments, and frequently perform costly Vision-Language (VL) inference. We present TrackGraph, an online open-vocabulary system that maintains short-term 2D mask identity directly in the image stream before fusing segments into 3D. FastSAM masks and CLIP features are computed at sparse keyframes, while dense DINOv3 features are used to propagate masks at a high rate in between. The resulting tracked masks are fused into a class-agnostic 3D segment layer within a hierarchical scene graph, with 3D association handling tracking interruptions and long-term revisits. Compact multi-view CLIP embeddings enable open-vocabulary retrieval. Across Replica, ScanNet++, and HM3D, TrackGraph achieves competitive open-vocabulary segmentation and retrieval against state-of-the-art mapping methods, including the highest synonym frequency on Replica (0.50). On the same NVIDIA A100, it is 1.7× faster and uses 3.3× less GPU memory than ViT-H OVI-MAP. Real-world quadruped deployments demonstrate onboard scene graph construction and object search at 7.5Hz, while recorded drone data is used to test the method under aerial viewpoints.

Overview

TrackGraph pipeline: sparse FastSAM and CLIP keyframes, DINOv3 mask propagation, source track fusion into 3D segments, and geometric and visual association across interrupted tracks and revisits.

FastSAM masks and CLIP features are computed at sparse keyframes, while dense DINOv3 features are used to propagate the masks and their identities through the image stream in between. We refer to each 2D temporally tracked mask identity as a source track. Source track masks are continually fused into a TSDF and mesh, where repeated observations reinforce consistent track identities that are then used to build persistent 3D segments.

Interrupted tracks and long-term revisits are reconciled via 3D geometric and visual association, while CLIP features are averaged per source track and stored as separate views in each 3D segment’s feature gallery. A text or image query is embedded into the same CLIP space as the stored features. The query interface returns a ranked list of segment nodes.

Results

We evaluate open-vocabulary segmentation and object retrieval on Replica, ScanNet++, and HM3D following OpenLex3D. Runtime and resource usage are evaluated on Replica, whose scenes have comparable durations and memory footprints and for which baseline measurements are available.

Quantitative Results

Open-Vocabulary Semantic Segmentation

Across encoders, TrackGraph ViT-H achieves the highest synonym frequency on Replica, is second to HOV-SG on ScanNet++, and remains competitive on HM3D, where OVI-MAP with SigLIP-L performs best. Using the same VL encoder, TrackGraph achieves equal synonym scores as OVI-MAP’s ViT-H variant on HM3D.

Table I. OpenLex3D Top-5 segmentation frequencies on Replica, ScanNet++, and HM3D. TrackGraph ViT-H synonym frequency is 0.50, 0.35, and 0.32, respectively. Categories are synonyms, depictions, visually similar, clutter, missing, and incorrect.

Open-Vocabulary 3D Object Retrieval

Among the ViT-H variants, TrackGraph leads in AP50 and AP25 on Replica but trails OVI-MAP in AP, and outperforms ConceptGraphs and HOV-SG in AP and AP50 on ScanNet++. On HM3D, TrackGraph trails all baseline methods with matched encoder on AP and AP50 despite competitive point-wise semantics and AP25.

Table II. Open-vocabulary 3D object retrieval on Replica, ScanNet++, and HM3D, evaluated using AP, AP50, and AP25. TrackGraph ViT-H AP is 6.00, 2.21, and 2.16, respectively.

Runtime and Resource Usage

On the same NVIDIA A100, TrackGraph is 1.7× faster and uses 3.3× less GPU memory than ViT-H OVI-MAP. The broader efficiency comparison includes ConceptGraphs and HOV-SG measurements on an NVIDIA RTX 3090, while TrackGraph, OVI-MAP, and FindAnything are evaluated on an NVIDIA A100. These efficiency results are not all measured under identical hardware and input-frame settings and should therefore be interpreted as indicative comparisons.

TrackGraph CPU memory depends strongly on voxel resolution: increasing the voxel size from 0.03 m to 0.05 m reduces peak RAM usage by more than one-third, with little change in synonym frequency but lower retrieval AP.

Table III. Runtime, GPU memory, and CPU memory on Replica. Results include different encoders, voxel sizes, and hardware; the caption identifies the comparison settings.

Real-World Experiments

Quadruped Exploration and Object Search

We evaluate the complete system on a quadruped robot in three large-scale indoor environments. Each experiment comprises two autonomous phases: exploration using a volumetric planner while constructing the open-vocabulary scene graph, followed by object search using text or reference-image queries. Across all three environments, all queried targets were successfully retrieved.

Two indoor quadruped experiments showing online scene graphs, robot paths, and navigation to segments retrieved with text and reference-image queries.
For two large real-world environments, a) shows the open-vocabulary scene graph constructed online by TrackGraph onboard a quadruped robot, while b) highlights retrieval with both image- and text-based queries, with successful navigation to all segments in the indicated goal order.

Aerial Mapping

We further demonstrate TrackGraph on recorded drone RGB-D data from three indoor environments at a search-and-rescue training site, complementing standard ground-based benchmarks with uncommon aerial viewpoints. The figure shows qualitative results from one environment.

Scene graph and retrieval results from recorded drone RGB-D data, including a barrel reference-image query and text queries for a red car and stairs.

Video

BibTeX

@article{hellesylt2026trackgraph,
  title={{TRACKGRAPH}: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking},
  author={Hellesylt, Peder Borge and Gassol Puigjaner, Albert and Alexis, Kostas and Stahl, Annette},
  journal={arXiv preprint arXiv:2609.31005},
  year={2026},
  url={https://arxiv.org/abs/2609.31005}
}