基于多模态数据融合分析的疲劳驾驶检测方法

曾琪, 王树祎, 刘奕

清华大学学报(自然科学版) ›› 2026, Vol. 66 ›› Issue (9) : 1873-1880.

PDF(2265 KB)
PDF(2265 KB)
清华大学学报(自然科学版) ›› 2026, Vol. 66 ›› Issue (9) : 1873-1880. DOI: 10.16511/j.cnki.qhdxxb.2026.27.050
公共安全

基于多模态数据融合分析的疲劳驾驶检测方法

作者信息 +

A fatigued driving detection method using multimodal data fusion analysis

Author information +
文章历史 +

摘要

疲劳驾驶是造成道路交通伤害的主要原因之一,高效准确的疲劳驾驶检测是保障交通安全的重要手段。但传统单一信号在捕捉疲劳状态时存在数据采集侵入性高、数据结构复杂、难以实时预测、准确率低等局限性,融合不同模态的信息成为新的发展方向。该文通过长期实车驾驶实验采集的心电信号、车辆轨迹数据、驾驶员脸部图像数据构建基于多模态信号的疲劳驾驶检测模型。该模型包含数据处理、特征提取、特征筛选和疲劳预测等模块,通过自注意力机制捕获长时间依赖关系并输出预测结果。该模型预测疲劳状态类别的准确率最高达97.89%,整体召回率和F1分数达80%以上,表明模型具有较好的预测准确性。将同时利用3个模态数据的检测模型设置为对照组,以仅使用单个模态数据或同时使用2个模态的数据为实验组,共6组实验组。试验发现同时利用3个模态的数据的疲劳驾驶检测模型在不同指标上均取得最优结果,表明所提出的模型成功融合了不同模态的信息,具有较高的准确率和适应性,能够提升驾驶安全的保障水平。

Abstract

Objective: As a primary cause of road traffic injuries, fatigued driving requires efficient and accurate detection to improve traffic safety. Traditional single-signal approaches face limitations in capturing fatigue states, including high data collection intrusiveness, complex data structures, difficulty in real-time prediction, and low accuracy. Integrating information from different modalities has emerged as a new direction for development. Methods: This study developed a multimodal fatigue detection model for drivers by using electrocardiogram signals, vehicle trajectory data, and driver facial video data collected during long-term real-vehicle driving experiments. For ECG signal processing, an improved two-step adaptive filtering method was adopted for denoising, followed by time-domain and frequency-domain analyses to extract the driver's ECG feature set. For vehicle trajectory data, the quartile method was first used to remove outliers, backward filling was applied to fill in missing values, and two-dimensional discrete wavelet analysis was then employed for data denoising. For facial image data, the LabelMe annotation tool was used to construct a personalized training dataset by manually annotating each driver's face with at least 300 images per subject. The YOLOv8 deep learning model was then fine-tuned and trained on this dataset, and the optimal model weights were saved. Next, the optimized model was used to automatically annotate the remaining unlabeled images, and the Hopenet algorithm was applied to extract head pose angles from the cropped facial regions. The model incorporated modules for data processing, feature extraction, feature selection, and fatigue prediction, utilizing a self-attention mechanism to capture long-term dependencies and generate predictive outputs. Results: The model achieved a maximum accuracy of 97.89% in predicting fatigue state categories, with overall recall and F1 scores exceeding 80%, demonstrating strong predictive accuracy. The detection model utilizing data from all three modalities served as the control group, while six experimental groups were formed using a single modality or a combination of two modalities. The experiments revealed that the fatigued driving detection model employing data from all three modalities achieved optimal performance across various metrics. Conclusions: This study demonstrates that the proposed model successfully integrates information from different modalities, exhibits high accuracy and adaptability, and enhances the assurance of driving safety. Specifically, this model outperforms all comparison models in terms of accuracy, precision, recall, and F1 score, achieving the best performance in fatigued driving detection tasks, and maintains stable detection performance even under complex data structures, diverse sources, large time spans, average-quality facial images, and individual driver differences. The model achieves a maximum accuracy of 97.89% in identifying fatigue states, indicating that it correctly learns and recognizes fatigue patterns. Furthermore, models using a single modality or a combination of two modalities yield lower evaluation metrics than the three-modality model. Moreover, two-modality models consistently outperform single-modality variants, confirming that ECG signals, trajectory data, and facial images contribute positively to accurate fatigued driving detection.

关键词

交通安全 / 疲劳驾驶检测 / 多模态数据融合 / 注意力机制

Key words

traffic safety / fatigued driving detection / multimodal data fusion / attention mechanism

引用本文

导出引用
曾琪, 王树祎, 刘奕. 基于多模态数据融合分析的疲劳驾驶检测方法[J]. 清华大学学报(自然科学版). 2026, 66(9): 1873-1880 https://doi.org/10.16511/j.cnki.qhdxxb.2026.27.050
Qi ZENG, Shuyi WANG, Yi LIU. A fatigued driving detection method using multimodal data fusion analysis[J]. Journal of Tsinghua University(Science and Technology). 2026, 66(9): 1873-1880 https://doi.org/10.16511/j.cnki.qhdxxb.2026.27.050
中图分类号: TP393.1   

参考文献

1
中华人民共和国国家统计局. 中国统计年鉴[M]. 北京: 中国统计出版社, 2024.
National Bureau of Statistics of China. China Statistical Yearbook[M]. Beijing: China Statistics Press, 2024.
2
世界卫生组织. 2023年全球道路安全状况报告[R]. 日内瓦: 世界卫生组织, 2024.
World Health Organization. Global status report on road safety 2023[R]. Geneva: World Health Organization, 2024.
3
PHILIP P F, LE B P, TAILLARD J, et al. Fatigue, alcohol, and serious road crashes in France: Factorial study of national data[J]. BMJ, 2001, 322(7290): 829- 830.
4
LI G, CHUNG W Y. Detection of driver drowsiness using wavelet analysis of heart rate variability and a support vector machine classifier[J]. Sensors, 2013, 13(12): 16494- 16511.
5
WANG M S, JEONG N T, KIM K S, et al. Drowsy behavior detection based on driving information[J]. International Journal of Automotive Technology, 2016, 17(1): 165- 173.
6
魏恒建. 基于多通道信息融合的驾驶员疲劳检测系统研究[D]. 大连: 大连交通大学, 2024.
WEI H J. Research on driver fatigue detection system based on multi-channel information fusion[D]. Dalian: Dalian Jiaotong University, 2024. (in Chinese)
7
李泰国, 张天策, 李超, 等. 基于面部倒立摆模型与信息熵的驾驶员疲劳检测[J]. 交通运输系统工程与信息, 2023, 23(5): 24- 32.
LI T G, ZHANG T C, LI C, et al. Driver fatigue detection based on facial inverted pendulum model and information entropy[J]. Journal of Transportation Systems Engineering and Information Technology, 2023, 23(5): 24- 32.
8
戴诗琪, 曾智勇. 基于深度学习的疲劳驾驶检测算法[J]. 计算机系统应用, 2018, 27(7): 113- 120.
DAI S Q, ZENG Z Y. Fatigue driving detection algorithm based on deep learning[J]. Computer Systems & Applications, 2018, 27(7): 113- 120.
9
周蒙. 考虑驾驶员个体差异的多特征融合疲劳驾驶检测算法研究[D]. 上海: 上海海洋大学, 2025.
ZHOU M. Research on fatigue driving detection algorithm based on multi-feature fusion considering individual driver differences[D]. Shanghai: Shanghai Ocean University, 2025. (in Chinese)
10
张伯辰, 施鑫杰, 霍梅梅. 基于OpenCV的树莓派人脸识别疲劳驾驶检测系统[J]. 现代计算机, 2021, 27(23): 129- 132.
ZHANG B C, SHI X J, HUO M M. Raspberry Pi face recognition fatigue driving detection system based on OpenCV[J]. Modern Computer, 2021, 27(23): 129- 132.
11
MURRAY B J. Subjective and objective assessment of hypersomnolence[J]. Sleep Medicine Clinics, 2017, 12(3): 313- 322.
12
龙伟峰, 胡江碧. 不同工作负荷下驾驶员的心生理特征研究[J]. 交通标准化, 2009(21): 59- 63.
LONG W F, HU J B. Study on driver's cardiac and physiological characteristics under different workloads[J]. Communications Standardization, 2009(21): 59- 63.
13
ROSA M M A, COSTA P Ü, COSTA E A C, et al. Design of a low power and robust VLSI power line interference canceler with optimized arithmetic operators[J]. Analog Integr. Circuits Signal Process, 2022, 112(2): 247- 261.
14
吕建行, 李玉榕, 陈建国, 等. 两步式自适应阈值法滤除心电信号中运动伪迹[J]. 电子学报, 2024, 52(10): 3493- 3506.
LYU J X, LI Y R, CHEN J G, et al. Two-step adaptive threshold method for removing motion artifacts in ECG signals[J]. Acta Electronica Sinica, 2024, 52(10): 3493- 3506.
15
胡枫. 基于马尔科夫模型的短时交通流预测研究[D]. 南京: 南京邮电大学, 2013.
HU F. Research on short-term traffic flow prediction based on Markov model[D]. Nanjing: Nanjing University of Posts and Telecommunications, 2013. (in Chinese)
16
张玺君, 袁占亭, 张红, 等. 交通轨迹大数据预处理方法研究[J]. 计算机工程, 2019, 45(6): 26- 31.
ZHANG X J, YUAN Z T, ZHANG H, et al. Research on preprocessing method of traffic trajectory big data[J]. Computer Engineering, 2019, 45(6): 26- 31.
17
ZHOU Y, BARNES C, LU J W, et al. On the continuity of rotation representations in neural networks[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach: IEEE, 2019: 5745–5753.
18
OROZCO J, GONG S, XIANG T. Head pose classification in crowded scenes[C]//Proceedings of the British Machine Vision Conference (BMVC). London: British Machine Vision Association, 2009: 1-11.
19
GUPTA A, THAKKAR K, GANDHI V, et al. Nose, eyes and ears: Head pose estimation by locating facial keypoints [C]//ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Brighton, UK: IEEE, 2019: 1977-1981.
20
LIU L Y, KE Z R, HUO J, et al. Head pose estimation through keypoints matching between reconstructed 3D face model and 2D image[J]. Sensors, 2021, 21(5): 1841.
21
ABATE A F, BISOGNI C, CASTIGLIONE A, et al. Head pose estimation: An extensive survey on recent techniques and applications[J]. Pattern Recognition, 2022, 127, 108591.
22
RUIZ N, CHONG E, REHG J M. Fine-grained head pose estimation without keypoints [C]//Proceedings of the IEEE conference on computer vision and pattern recognition workshops, 2018: 2074–2083.
23
VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[J/OL]. arXiv, 2017. DOI: 10.48550/arXiv. 1706.03762. [2025-09-23]. https://arxiv.org/abs/1706.03762.

基金

国家重点研发计划项目(2024YFC3017000)
国家自然科学基金面上项目(72334003)
国家自然科学基金面上项目(72174102)

版权

版权所有,未经授权,不得转载。
PDF(2265 KB)

Accesses

Citation

Detail

段落导航
相关文章

/