基于重心位移编码的毫米波雷达场景重建方法

张远, 何子恒, 吴冶

清华大学学报(自然科学版) ›› 2026, Vol. 66 ›› Issue (9) : 1844-1853.

PDF(6082 KB)
PDF(6082 KB)
清华大学学报(自然科学版) ›› 2026, Vol. 66 ›› Issue (9) : 1844-1853. DOI: 10.16511/j.cnki.qhdxxb.2026.27.047
公共安全

基于重心位移编码的毫米波雷达场景重建方法

作者信息 +

Millimeter-wave radar-based scene reconstruction via centroid displacement encoding

Author information +
文章历史 +

摘要

毫米波雷达在室内感知中难以探测缺乏速度变化的静态目标。当前雷达技术多关注动态目标(如人体)检测,未充分利用人体动态信息的场景理解潜力。为此,本文提出一种基于重心位移编码的毫米波雷达场景重建方法。该方法以人体重心的相对位移轨迹刻画人-物交互的宏观动态特征,并结合时空图卷积、场景感知投票与混合预测模型,实现对静态物体类别与空间布局的联合推理。通过捕捉人体宏观运动趋势以有效过滤噪声,该方法成功克服了因雷达点云稀疏及局部抖动导致的绝对坐标姿态特征不稳定的问题。实验结果表明,该方法在场景重建中的整体精度达到50.81%,较视觉基准方法提升了4.50个百分点,且在多类典型交互物体上表现出更稳定的预测能力。该研究为毫米波雷达室内环境感知提供了新思路。

Abstract

Objective: As a critical environmental perception sensor, millimeter-wave radar offers notable advantages such as privacy protection and immunity to lighting conditions, making it well-suited to complex environments. However, accurately detecting static targets, such as furniture and walls, remains a substantial challenge because static objects lack relative motion, causing their Doppler shifts to approach zero so that they are easily filtered out as noise. Furthermore, traditional radar scene reconstruction methods, such as radar Simultaneous Localization and Mapping (SLAM), rely heavily on static point cloud accumulation, which provides only rough geometric outlines and fails to extract semantic categories or fine-grained three-dimensional (3D) bounding boxes. Existing vision-based scene reconstruction methods depend largely on RGB-D data and remain constrained by privacy and lighting issues. Although recent studies have attempted to use human–object interaction for scene reasoning, applying these vision-driven methods directly to millimeter-wave radar is difficult. Radar point clouds are inherently sparse and accompanied by high-frequency jitter, and directly using traditional absolute global coordinate features introduces cascading errors. Therefore, exploring how to effectively use the dynamic human posture information captured by millimeter-wave radar to accurately infer static indoor scene layouts is critically important. Methods: To address the problems of sparse point clouds and unstable absolute pose features, this paper proposed a novel millimeter-wave radar-driven indoor scene reconstruction method that indirectly infers static scene layouts from human motion patterns. The proposed model consisted of three core modules: spatiotemporal feature extraction, a scene-aware voting mechanism, and Gaussian mixture decoding. Initially, human activity point clouds were captured, and a 3D skeletal posture sequence was extracted. In the feature extraction stage, a centroid displacement encoding scheme was designed. Using the relative displacement trajectory of the human centroid rather than absolute joint coordinates, the model captures the macroscopic dynamic features of human–object interactions while effectively filtering out local high-frequency noise. Subsequently, a spatial graph attention mechanism and one-dimensional temporal convolutions were encapsulated into a stacked spatiotemporal residual module to extract deep interactive representations. In the scene-aware voting phase, the human centroid served as a seed position, and a learnable offset function calculated the center votes for potential interactive objects; these votes were then clustered and weighted into stable voting clusters. Finally, considering the inherent uncertainty of predicting scenes from single-frame postures, a hybrid prediction module was introduced. It used a Gaussian mixture distribution to model the 3D bounding box parameters (center, size, and orientation), generating diverse and plausible scene hypotheses that are jointly optimized by classification and Huber regression losses. Results: Extensive experiments were conducted on a self-constructed real-world millimeter-wave radar dataset encompassing 6 indoor spatial layouts, 12 interacting object categories, and 20,000 human posture sequences. The quantitative results demonstrated that the proposed method achieved an overall mean average precision (mAP) of 50.81% in 3D scene reconstruction. This performance surpassed that of mainstream visual scene reconstruction models, achieving a 4.50 percentage point improvement over the best baseline. Notably, for highly challenging categories such as "sofa," the mAP reached 78.17%. Comprehensive ablation studies confirmed the necessity of each component: introducing the centroid displacement encoding raised the mAP from 8.66% to 33.89%, while integrating spatiotemporal encoding and the hybrid prediction module further improved the accuracy to the final 50.81%. Furthermore, cross-room generalization evaluations using minimal matching distance and total mutual diversity metrics under different data-splitting strategies (S1 and S2) showed that the proposed multimodal decoding effectively balances prediction accuracy and scene generation diversity. Qualitative visualization results further indicated that the generated 3D bounding boxes were highly consistent with real environments in terms of structural configuration and spatial rationality, with no severe unnatural penetrations. Conclusions: The proposed scene reconstruction method based on centroid displacement encoding reduces dependence on absolute posture coordinates and overcomes the instability caused by the sparsity and local jitter of millimeter-wave radar point clouds. By capturing macroscopic motion trends and employing a hybrid multimodal decoding mechanism, the method significantly improves both reconstruction accuracy and prediction stability for multiple typical human–object interaction targets. In addition, it demonstrates strong cross-scene generalization capabilities. Although the system relies on the accuracy of frontend human posture extraction and may be affected by severe multipath effects in complex metallic environments, it establishes an effective framework and offers a novel perspective for intelligent indoor environmental perception using millimeter-wave radar. Future work will focus on multisensor fusion and robust feature extraction under severe multipath conditions to further enhance the system's engineering applicability.

关键词

毫米波雷达 / 场景重建 / 重心位移编码 / 人-物交互 / 智能感知

Key words

millimeter-wave radar / scene reconstruction / centroid displacement encoding / human–object interaction / intelligent perception

引用本文

导出引用
张远, 何子恒, 吴冶. 基于重心位移编码的毫米波雷达场景重建方法[J]. 清华大学学报(自然科学版). 2026, 66(9): 1844-1853 https://doi.org/10.16511/j.cnki.qhdxxb.2026.27.047
Yuan ZHANG, Ziheng HE, Ye WU. Millimeter-wave radar-based scene reconstruction via centroid displacement encoding[J]. Journal of Tsinghua University(Science and Technology). 2026, 66(9): 1844-1853 https://doi.org/10.16511/j.cnki.qhdxxb.2026.27.047
中图分类号: TN957   

参考文献

1
WEI Z Q, ZHANG F K, CHANG S, et al. MmWave radar and vision fusion for object detection in autonomous driving: A review[J]. Sensors, 2022, 22(7): 2542.
2
PEARCE A, ZHANG J A, XU R, et al. Multi-object tracking with mmWave radar: A review[J]. Electronics, 2023, 12(2): 308.
3
SOUMYA A, KRISHNA MOHAN C, CENKERAMADDI L R. Recent advances in mmWave-radar-based sensing, its applications, and machine learning techniques: A review[J]. Sensors, 2023, 23(21): 8901.
4
KONG H, HUANG C, YU J D, et al. A survey of mmWave radar-based sensing in autonomous vehicles, smart homes and industry[J]. IEEE Communications Surveys & Tutorials, 2025, 27(1): 463- 508.
5
ABDU F J, ZHANG Y X, FU M Z, et al. Application of deep learning on millimeter-wave radar signals: A review[J]. Sensors, 2021, 21(6): 1951.
6
WANG S, MEI L Y, LIU R F, et al. Multi-modal fusion sensing: A comprehensive review of millimeter-wave radar and its integration with other modalities[J]. IEEE Communications Surveys & Tutorials, 2025, 27(1): 322- 352.
7
ZHOU T H, YANG M M, JIANG K, et al. MMW radar-based technologies in autonomous driving: A review[J]. Sensors, 2020, 20(24): 7283.
8
TUNG NG H, IBRAHIM H, RAJENDRAN P. A literature review on the usage of mmWave radar in UAV's detect-and-avoid applications[J]. Journal of Computer Science & Computational Mathematics, 2023, 13(2): 39- 45.
9
HARMER S, BOWRING N, ANDREWS D, et al. A review of nonimaging stand-off concealed threat detection with millimeter-wave radar[Application Notes][J]. IEEE Microwave Magazine, 2012, 13(1): 160- 167.
10
SINGH A D, SANDHA S S, GARCIA L, et al. RadHAR: Human activity recognition from point clouds generated through a millimeter-wave radar[C]//Proceedings of the 3rd ACM Workshop on Millimeter-Wave Networks and Sensing Systems. Los Cabos, Mexico: ACM, 2019: 51–56.
11
SENGUPTA A, JIN F, ZHANG R Y, et al. mm-Pose: Real-time human skeletal posture estimation using mmWave radars and CNNs[J]. IEEE Sensors Journal, 2020, 20(17): 10032- 10044.
12
XUE H F, JU Y, MIAO C L, et al. mmMesh: Towards 3D real-time dynamic human mesh construction using millimeter-wave[C]//Proceedings of the 19th Annual International Conference on Mobile Systems, Applications, and Services. New York, USA: ACM, 2021: 269–282.
13
PATOLE S M, TORLAK M, WANG D, et al. Automotive radars: A review of signal processing techniques[J]. IEEE Signal Processing Magazine, 2017, 34(2): 22- 35.
14
LU C X, ROSA S, ZHAO P J, et al. See through smoke: Robust indoor mapping with low-cost mmWave radar[C]//Proceedings of the 18th International Conference on Mobile Systems, Applications, and Services. Toronto, Canada: ACM, 2020: 14–27.
15
ABDUL-RASHID H, YUAN J F, LI B, et al. SHREC'18 track: 2D image-based 3D scene retrieval[J]. Training, 2018, 700(70): 2- 4.
16
CHEN Z X, WANG G C, LIU Z W. SceneDreamer: Unbounded 3D scene generation from 2D image collections[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(12): 15562- 15576.
17
SAVVA M, CHANG A X, HANRAHAN P, et al. PiGraphs: Learning interaction snapshots from observations[J]. ACM Transactions on Graphics, 2016, 35(4): 1- 12.
18
YI H W, HUANG C H P, TRIPATHI S, et al. MIME: Human-aware 3D scene generation[C]//Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver, Canada: IEEE, 2023: 12965–12976.
19
NIE Y Y, DAI A, HAN X G, et al. Pose2Room: Understanding 3D scenes from human activities[C]//Computer Vision - ECCV 2022: 17th European Conference on Computer Vision. Tel Aviv, Israel: ECCV, 2022: 425–443.
20
HONG Z Y, PETILLOT Y, WALLACE A, et al. RadarSLAM: A robust simultaneous localization and mapping system for all weather conditions[J]. The International Journal of Robotics Research, 2022, 41(5): 519- 542.
21
FANG H, ZHU X, GUREVYCH I. Preemptive detection and correction of misaligned actions in LLM agents[C]//Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Suzhou, China: EMNLP, 2025: 222–244.
22
ESLAMI S M A, HEESS N, WEBER T, et al. Attend, infer, repeat: Fast scene understanding with generative models[C]//Proceedings of the 30th International Conference on Neural Information Processing Systems. Barcelona, Spain: Curran Associates Inc., 2016: 3233–3241.
23
TANG J P, YIN Y Y, MARKHASIN L, et al. Diffuscene: Denoising diffusion models for generative indoor scene synthesis[C]//Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle, USA: IEEE, 2024: 20507–20518.

版权

版权所有,未经授权,不得转载。
PDF(6082 KB)

Accesses

Citation

Detail

段落导航
相关文章

/