OREN-X jointly maps signed distance, radiance and vision-language features of Replica room0 online with a single octree.

Abstract

To achieve general-purpose autonomy over long horizons, a robot needs to maintain spatial environment information that supports a variety of tasks: geometry for planning and control, radiance for rendering and relocalization, and vision-language features for open-vocabulary grounding. Existing methods represent and estimate each modality separately, multiplying memory and compute cost while forgoing potential synergy among the representations. We develop OREN-X, an online mapping method that uses an octree in 3D space as a shared data structure for indexing and storing a multi-modal field, capturing geometric, radiance, and vision-language information. OREN-X provides efficient unified storage and retrieval of these data in explicit/implicit and full/compressed form. Our unified representation yields cross-modality synergy: SDF estimates are sharpened by occupancy and radiance, while GPU-based ray-octree traversal and octree query enable real-time rendering. We also use online dictionary learning to compress the vision-language features, shrinking them 3.7× below full per-vertex storage while raising the query accuracy. On Replica, OREN-X maps in real time (80+ fps for SDF and 30+ fps for all four modalities), improves near-surface SDF accuracy by 33% over single-modality baselines, and improves mean open-vocabulary 3D mIoU by 71% and mean accuracy by 61% over the best prior method.

From a posed RGB-D stream, OREN-X estimates signed distance, occupancy, radiance and vision-language (VL) features of one multi-modal field in real time with a single octree.

  • Unified multi-modal mapping. The octree provides common storage, shared indexing and joint training for all modalities. Multi-pool indexing lets each modality share the octree at its own spatial resolution. Every modality stores its parameters at the octree vertices as explicit values, latent features decoded by a small network, or both, in full or compressed form. All modalities use the same octree lookup to interpolate their parameters at any query point, and an active-vertex online update keeps the per-frame cost bounded.
  • Cross-modality synergy. Modalities supervise one another: radiance supervision through an SDF-derived density improves SDF accuracy, and an SDF-occupancy consistency loss improves robustness under strong depth noise. One octree descent, ray-octree traversal or in-frustum voxel filtering serves several modalities at once, and shared indexing enables both 2D and 3D open-vocabulary queries.
  • Memory and compute efficiency. The shared allocation, indexing and updates are paid once rather than once per modality. Since the VL features need the most memory, an online dictionary learned from the streaming features compresses them: each vertex stores a code over append-only atoms instead of the raw feature, and text queries are scored directly in the low-dimensional code space.

Online Mapping

Replica TUM RGB-D
Left column: input RGB, input depth, rendered RGB, rendered VL feature Top: mesh with camera trajectory, SDF slice Bottom: VL feature (PCA), relevancy of the query "pillow"

Real-World Demos

Left column: input RGB, input depth, rendered RGB, rendered VL feature Top: mesh with camera trajectory, SDF slice Bottom: VL feature (PCA), relevancy of the query "table"

Extractor-Agnostic Language Field

CLIP TIPS
CLIP TIPS

The same octree stores vision-language features from different backbones, here CLIP and TIPS. Drag the divider to compare the PCA of the learned VL features.

Open-Vocabulary 3D Query

Left method Right method

Text query relevancy rendered on the reconstructed scene, from low (blue) to high (red). Pick a method for each half and drag the divider to compare.

BibTeX

@misc{dai2026orenx,
  title         = {OREN-X: Octree Residual Network for Real-Time Multi-Modal Mapping},
  author        = {Zhirui Dai and Qihao Qian and Dinh Minh Nguyen and Quan-Dung Pham and Kiana Bronder
                   and Carlos Nieto-Granda and Yiyu Chen and Quan Nguyen and Nikolay Atanasov},
  year          = {2026},
  eprint        = {2609.29157},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2609.29157}
}

@inproceedings{Dai_OREN_IROS26,
  author    = {Zhirui Dai and Qihao Qian and Tianxing Fan and Nikolay Atanasov},
  title     = {{OREN: Octree Residual Network for Real-Time Euclidean Signed Distance Mapping}},
  booktitle = {IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2510.18999}
}