HIGS: Hierarchical Implicit Grids for Joint Geometric and Semantic Scene Understanding

Hanwen Cao1†, Wenqiang Wu1,4*, Kuang-Ting Tu1*, Mathias Otnes1,5*, Jeffrey Delmerico2, Rui Wang2, Yulun Tian3 and Nikolay Atanasov1

1 University of California San Diego 2 Microsoft Spatial AI Lab 3 University of Michigan

4 Southern University of Science and Technology 5 Norwegian University of Science and Technology

† Work done partially during an internship at Microsoft Research. * Equal second author.

HIGS overview: fine- and coarse-level geometric and semantic features form joint features. An MLP decodes geometric features for signed distance estimation and mesh construction, while a transformer decodes features for open-vocabulary object grounding.

Abstract

Neural implicit representations have had a significant impact on scene reconstruction by enabling robots to build continuous, differentiable, and high-fidelity 3D maps. While most previous works focus on geometric reconstruction, we ask the question, can we have a unified representation that supports both geometric and semantic understanding? The answer is yes! We introduce HIGS, which leverages hierarchial implicit 3D grids and a unified query and decoding mechanism to support different kinds of features. We demonstrate our method on both SDF construction and open-vocabulary object grounding. On SDF construction, we achieve results better than or on par with state-of-the-art methods. On open-vocabulary object grounding, we show that our method can achieve more than 20x memory compression with small accuracy loss (2%-3%). We also explore the synergy between geometric and semantic features, and show they can enhance each other. To better handle the compute complexity as the map scale grows, we propose different techniques including hierarchical optimization, initial feature prediction, submap decomposition, and latent-space submap alignment. Our experiments show up to 30x speedup in scene construction and 6x speedup in submap alignment.

Applications

Scene construction
Open-volcabulary object grounding
Global consistent submap alignment

Hierarchical Implicit Grids

Hierarchical implicit grids method overview

HIGS represents scenes with hierarchical implicit 3D grids that explicitly disentangles coarse and fine information. A universal query and decoding process connects the latent grid features with raw input data. More specifically, at each query point, grid features are trilinearly interpolated, concatenated across levels, and decoded into either signed distances or vision-language features via the pretrained decoders. The 3D grids support either dense feature mapping by storing per-grid features or sparse feature mapping by leveraging a hash function.

Object Grounding

SDF-guided scene encoder method overview
SDF-guided scene encoder
Training-free inference method overview
Training-free inference

The object grounding task requirese accurate 3D bounding box prediction. However, the 3D grids do not necessarily lie on the object surfaces and are distributed evenly in space, thus losing detailed structure information. To mitigate this issue, we integrate the latent SDF features into the transformer attention blocks to enhance the structural awareness and greatly improve the accuracy of 3D bounding box predictions compared with no SDF guidance. Since our implicit grid features can be decoded into the original vision-language features, we can use existing object grounding models by mapping their corresponding features. In that sense, HiGS becomes a memory efficient representation.

Latent-space Submap Alignment

Coarse-level latent alignment method overview
Coarse-level latent alignment
Fine-level latent alignment method overview
Fine-level latent alignment

As the robot moves in the environment, the trajectory and mapping may accumulate errors over time. To get a globally consistent representation, HIGS corrects drift by aligning overlapping submaps directly in latent feature space and progressively incorporating finer grid levels without explicit point correspondences. Compared to previous works based on final decoded SDF, our latent-space alignment is more efficient and less sensitive to noise. An optional short refinement using decoded predictions can further improve accuracy. After alignment, features from overlapping submaps are averaged and decoded to produce a globally consistent field.

BibTeX

@article{cao2026higs,
    title = {HIGS: Hierarchical Implicit Grids for Joint Geometric and Semantic Scene Understanding},
    author = {Cao, Hanwen and Wu, Wenqiang and Tu, Kuang-Ting and Otnes, Mathias and 
    Delmerico, Jeffrey and Wang, Rui and Tian, Yulun and Atanasov, Nikolay},
    journal = {arXiv preprint arXiv:2609.38620},
    year = {2026}
}