Millimeter-wave radar-based scene reconstruction via centroid displacement encoding

Yuan ZHANG, Ziheng HE, Ye WU

Journal of Tsinghua University(Science and Technology) ›› 2026, Vol. 66 ›› Issue (9) : 1844-1853.

PDF(6082 KB)
PDF(6082 KB)
Journal of Tsinghua University(Science and Technology) ›› 2026, Vol. 66 ›› Issue (9) : 1844-1853. DOI: 10.16511/j.cnki.qhdxxb.2026.27.047
Public Safety

Millimeter-wave radar-based scene reconstruction via centroid displacement encoding

Author information +
History +

Abstract

Objective: As a critical environmental perception sensor, millimeter-wave radar offers notable advantages such as privacy protection and immunity to lighting conditions, making it well-suited to complex environments. However, accurately detecting static targets, such as furniture and walls, remains a substantial challenge because static objects lack relative motion, causing their Doppler shifts to approach zero so that they are easily filtered out as noise. Furthermore, traditional radar scene reconstruction methods, such as radar Simultaneous Localization and Mapping (SLAM), rely heavily on static point cloud accumulation, which provides only rough geometric outlines and fails to extract semantic categories or fine-grained three-dimensional (3D) bounding boxes. Existing vision-based scene reconstruction methods depend largely on RGB-D data and remain constrained by privacy and lighting issues. Although recent studies have attempted to use human–object interaction for scene reasoning, applying these vision-driven methods directly to millimeter-wave radar is difficult. Radar point clouds are inherently sparse and accompanied by high-frequency jitter, and directly using traditional absolute global coordinate features introduces cascading errors. Therefore, exploring how to effectively use the dynamic human posture information captured by millimeter-wave radar to accurately infer static indoor scene layouts is critically important. Methods: To address the problems of sparse point clouds and unstable absolute pose features, this paper proposed a novel millimeter-wave radar-driven indoor scene reconstruction method that indirectly infers static scene layouts from human motion patterns. The proposed model consisted of three core modules: spatiotemporal feature extraction, a scene-aware voting mechanism, and Gaussian mixture decoding. Initially, human activity point clouds were captured, and a 3D skeletal posture sequence was extracted. In the feature extraction stage, a centroid displacement encoding scheme was designed. Using the relative displacement trajectory of the human centroid rather than absolute joint coordinates, the model captures the macroscopic dynamic features of human–object interactions while effectively filtering out local high-frequency noise. Subsequently, a spatial graph attention mechanism and one-dimensional temporal convolutions were encapsulated into a stacked spatiotemporal residual module to extract deep interactive representations. In the scene-aware voting phase, the human centroid served as a seed position, and a learnable offset function calculated the center votes for potential interactive objects; these votes were then clustered and weighted into stable voting clusters. Finally, considering the inherent uncertainty of predicting scenes from single-frame postures, a hybrid prediction module was introduced. It used a Gaussian mixture distribution to model the 3D bounding box parameters (center, size, and orientation), generating diverse and plausible scene hypotheses that are jointly optimized by classification and Huber regression losses. Results: Extensive experiments were conducted on a self-constructed real-world millimeter-wave radar dataset encompassing 6 indoor spatial layouts, 12 interacting object categories, and 20,000 human posture sequences. The quantitative results demonstrated that the proposed method achieved an overall mean average precision (mAP) of 50.81% in 3D scene reconstruction. This performance surpassed that of mainstream visual scene reconstruction models, achieving a 4.50 percentage point improvement over the best baseline. Notably, for highly challenging categories such as "sofa," the mAP reached 78.17%. Comprehensive ablation studies confirmed the necessity of each component: introducing the centroid displacement encoding raised the mAP from 8.66% to 33.89%, while integrating spatiotemporal encoding and the hybrid prediction module further improved the accuracy to the final 50.81%. Furthermore, cross-room generalization evaluations using minimal matching distance and total mutual diversity metrics under different data-splitting strategies (S1 and S2) showed that the proposed multimodal decoding effectively balances prediction accuracy and scene generation diversity. Qualitative visualization results further indicated that the generated 3D bounding boxes were highly consistent with real environments in terms of structural configuration and spatial rationality, with no severe unnatural penetrations. Conclusions: The proposed scene reconstruction method based on centroid displacement encoding reduces dependence on absolute posture coordinates and overcomes the instability caused by the sparsity and local jitter of millimeter-wave radar point clouds. By capturing macroscopic motion trends and employing a hybrid multimodal decoding mechanism, the method significantly improves both reconstruction accuracy and prediction stability for multiple typical human–object interaction targets. In addition, it demonstrates strong cross-scene generalization capabilities. Although the system relies on the accuracy of frontend human posture extraction and may be affected by severe multipath effects in complex metallic environments, it establishes an effective framework and offers a novel perspective for intelligent indoor environmental perception using millimeter-wave radar. Future work will focus on multisensor fusion and robust feature extraction under severe multipath conditions to further enhance the system's engineering applicability.

Key words

millimeter-wave radar / scene reconstruction / centroid displacement encoding / human–object interaction / intelligent perception

Cite this article

Download Citations
Yuan ZHANG , Ziheng HE , Ye WU. Millimeter-wave radar-based scene reconstruction via centroid displacement encoding[J]. Journal of Tsinghua University(Science and Technology). 2026, 66(9): 1844-1853 https://doi.org/10.16511/j.cnki.qhdxxb.2026.27.047

References

1
WEI Z Q, ZHANG F K, CHANG S, et al. MmWave radar and vision fusion for object detection in autonomous driving: A review[J]. Sensors, 2022, 22(7): 2542.
2
PEARCE A, ZHANG J A, XU R, et al. Multi-object tracking with mmWave radar: A review[J]. Electronics, 2023, 12(2): 308.
3
SOUMYA A, KRISHNA MOHAN C, CENKERAMADDI L R. Recent advances in mmWave-radar-based sensing, its applications, and machine learning techniques: A review[J]. Sensors, 2023, 23(21): 8901.
4
KONG H, HUANG C, YU J D, et al. A survey of mmWave radar-based sensing in autonomous vehicles, smart homes and industry[J]. IEEE Communications Surveys & Tutorials, 2025, 27(1): 463- 508.
5
ABDU F J, ZHANG Y X, FU M Z, et al. Application of deep learning on millimeter-wave radar signals: A review[J]. Sensors, 2021, 21(6): 1951.
6
WANG S, MEI L Y, LIU R F, et al. Multi-modal fusion sensing: A comprehensive review of millimeter-wave radar and its integration with other modalities[J]. IEEE Communications Surveys & Tutorials, 2025, 27(1): 322- 352.
7
ZHOU T H, YANG M M, JIANG K, et al. MMW radar-based technologies in autonomous driving: A review[J]. Sensors, 2020, 20(24): 7283.
8
TUNG NG H, IBRAHIM H, RAJENDRAN P. A literature review on the usage of mmWave radar in UAV's detect-and-avoid applications[J]. Journal of Computer Science & Computational Mathematics, 2023, 13(2): 39- 45.
9
HARMER S, BOWRING N, ANDREWS D, et al. A review of nonimaging stand-off concealed threat detection with millimeter-wave radar[Application Notes][J]. IEEE Microwave Magazine, 2012, 13(1): 160- 167.
10
SINGH A D, SANDHA S S, GARCIA L, et al. RadHAR: Human activity recognition from point clouds generated through a millimeter-wave radar[C]//Proceedings of the 3rd ACM Workshop on Millimeter-Wave Networks and Sensing Systems. Los Cabos, Mexico: ACM, 2019: 51–56.
11
SENGUPTA A, JIN F, ZHANG R Y, et al. mm-Pose: Real-time human skeletal posture estimation using mmWave radars and CNNs[J]. IEEE Sensors Journal, 2020, 20(17): 10032- 10044.
12
XUE H F, JU Y, MIAO C L, et al. mmMesh: Towards 3D real-time dynamic human mesh construction using millimeter-wave[C]//Proceedings of the 19th Annual International Conference on Mobile Systems, Applications, and Services. New York, USA: ACM, 2021: 269–282.
13
PATOLE S M, TORLAK M, WANG D, et al. Automotive radars: A review of signal processing techniques[J]. IEEE Signal Processing Magazine, 2017, 34(2): 22- 35.
14
LU C X, ROSA S, ZHAO P J, et al. See through smoke: Robust indoor mapping with low-cost mmWave radar[C]//Proceedings of the 18th International Conference on Mobile Systems, Applications, and Services. Toronto, Canada: ACM, 2020: 14–27.
15
ABDUL-RASHID H, YUAN J F, LI B, et al. SHREC'18 track: 2D image-based 3D scene retrieval[J]. Training, 2018, 700(70): 2- 4.
16
CHEN Z X, WANG G C, LIU Z W. SceneDreamer: Unbounded 3D scene generation from 2D image collections[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(12): 15562- 15576.
17
SAVVA M, CHANG A X, HANRAHAN P, et al. PiGraphs: Learning interaction snapshots from observations[J]. ACM Transactions on Graphics, 2016, 35(4): 1- 12.
18
YI H W, HUANG C H P, TRIPATHI S, et al. MIME: Human-aware 3D scene generation[C]//Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver, Canada: IEEE, 2023: 12965–12976.
19
NIE Y Y, DAI A, HAN X G, et al. Pose2Room: Understanding 3D scenes from human activities[C]//Computer Vision - ECCV 2022: 17th European Conference on Computer Vision. Tel Aviv, Israel: ECCV, 2022: 425–443.
20
HONG Z Y, PETILLOT Y, WALLACE A, et al. RadarSLAM: A robust simultaneous localization and mapping system for all weather conditions[J]. The International Journal of Robotics Research, 2022, 41(5): 519- 542.
21
FANG H, ZHU X, GUREVYCH I. Preemptive detection and correction of misaligned actions in LLM agents[C]//Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Suzhou, China: EMNLP, 2025: 222–244.
22
ESLAMI S M A, HEESS N, WEBER T, et al. Attend, infer, repeat: Fast scene understanding with generative models[C]//Proceedings of the 30th International Conference on Neural Information Processing Systems. Barcelona, Spain: Curran Associates Inc., 2016: 3233–3241.
23
TANG J P, YIN Y Y, MARKHASIN L, et al. Diffuscene: Denoising diffusion models for generative indoor scene synthesis[C]//Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle, USA: IEEE, 2024: 20507–20518.

RIGHTS & PERMISSIONS

All rights reserved. Unauthorized reproduction is prohibited.
PDF(6082 KB)

Accesses

Citation

Detail

Sections
Recommended

/