OREN-X jointly maps signed distance, radiance and vision-language features of Replica room0 online with a single octree.
To achieve general-purpose autonomy over long horizons, a robot needs to maintain spatial environment information that supports a variety of tasks: geometry for planning and control, radiance for rendering and relocalization, and vision-language features for open-vocabulary grounding. Existing methods represent and estimate each modality separately, multiplying memory and compute cost while forgoing potential synergy among the representations. We develop OREN-X, an online mapping method that uses an octree in 3D space as a shared data structure for indexing and storing a multi-modal field, capturing geometric, radiance, and vision-language information. OREN-X provides efficient unified storage and retrieval of these data in explicit/implicit and full/compressed form. Our unified representation yields cross-modality synergy: SDF estimates are sharpened by occupancy and radiance, while GPU-based ray-octree traversal and octree query enable real-time rendering. We also use online dictionary learning to compress the vision-language features, shrinking them 3.7× below full per-vertex storage while raising the query accuracy. On Replica, OREN-X maps in real time (80+ fps for SDF and 30+ fps for all four modalities), improves near-surface SDF accuracy by 33% over single-modality baselines, and improves mean open-vocabulary 3D mIoU by 71% and mean accuracy by 61% over the best prior method.
From a posed RGB-D stream, OREN-X estimates signed distance, occupancy, radiance and vision-language (VL) features of one multi-modal field in real time with a single octree.
The same octree stores vision-language features from different backbones, here CLIP and TIPS. Drag the divider to compare the PCA of the learned VL features.
Text query relevancy rendered on the reconstructed scene, from low (blue) to high (red). Pick a method for each half and drag the divider to compare.
@misc{dai2026orenx,
title = {OREN-X: Octree Residual Network for Real-Time Multi-Modal Mapping},
author = {Zhirui Dai and Qihao Qian and Dinh Minh Nguyen and Quan-Dung Pham and Kiana Bronder
and Carlos Nieto-Granda and Yiyu Chen and Quan Nguyen and Nikolay Atanasov},
year = {2026},
eprint = {2609.29157},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.29157}
}
@inproceedings{Dai_OREN_IROS26,
author = {Zhirui Dai and Qihao Qian and Tianxing Fan and Nikolay Atanasov},
title = {{OREN: Octree Residual Network for Real-Time Euclidean Signed Distance Mapping}},
booktitle = {IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
year = {2026},
url = {https://arxiv.org/abs/2510.18999}
}