A fatigued driving detection method using multimodal data fusion analysis

Qi ZENG, Shuyi WANG, Yi LIU

Journal of Tsinghua University(Science and Technology) ›› 2026, Vol. 66 ›› Issue (9) : 1873-1880.

PDF(2265 KB)
PDF(2265 KB)
Journal of Tsinghua University(Science and Technology) ›› 2026, Vol. 66 ›› Issue (9) : 1873-1880. DOI: 10.16511/j.cnki.qhdxxb.2026.27.050
Public Safety

A fatigued driving detection method using multimodal data fusion analysis

Author information +
History +

Abstract

Objective: As a primary cause of road traffic injuries, fatigued driving requires efficient and accurate detection to improve traffic safety. Traditional single-signal approaches face limitations in capturing fatigue states, including high data collection intrusiveness, complex data structures, difficulty in real-time prediction, and low accuracy. Integrating information from different modalities has emerged as a new direction for development. Methods: This study developed a multimodal fatigue detection model for drivers by using electrocardiogram signals, vehicle trajectory data, and driver facial video data collected during long-term real-vehicle driving experiments. For ECG signal processing, an improved two-step adaptive filtering method was adopted for denoising, followed by time-domain and frequency-domain analyses to extract the driver's ECG feature set. For vehicle trajectory data, the quartile method was first used to remove outliers, backward filling was applied to fill in missing values, and two-dimensional discrete wavelet analysis was then employed for data denoising. For facial image data, the LabelMe annotation tool was used to construct a personalized training dataset by manually annotating each driver's face with at least 300 images per subject. The YOLOv8 deep learning model was then fine-tuned and trained on this dataset, and the optimal model weights were saved. Next, the optimized model was used to automatically annotate the remaining unlabeled images, and the Hopenet algorithm was applied to extract head pose angles from the cropped facial regions. The model incorporated modules for data processing, feature extraction, feature selection, and fatigue prediction, utilizing a self-attention mechanism to capture long-term dependencies and generate predictive outputs. Results: The model achieved a maximum accuracy of 97.89% in predicting fatigue state categories, with overall recall and F1 scores exceeding 80%, demonstrating strong predictive accuracy. The detection model utilizing data from all three modalities served as the control group, while six experimental groups were formed using a single modality or a combination of two modalities. The experiments revealed that the fatigued driving detection model employing data from all three modalities achieved optimal performance across various metrics. Conclusions: This study demonstrates that the proposed model successfully integrates information from different modalities, exhibits high accuracy and adaptability, and enhances the assurance of driving safety. Specifically, this model outperforms all comparison models in terms of accuracy, precision, recall, and F1 score, achieving the best performance in fatigued driving detection tasks, and maintains stable detection performance even under complex data structures, diverse sources, large time spans, average-quality facial images, and individual driver differences. The model achieves a maximum accuracy of 97.89% in identifying fatigue states, indicating that it correctly learns and recognizes fatigue patterns. Furthermore, models using a single modality or a combination of two modalities yield lower evaluation metrics than the three-modality model. Moreover, two-modality models consistently outperform single-modality variants, confirming that ECG signals, trajectory data, and facial images contribute positively to accurate fatigued driving detection.

Key words

traffic safety / fatigued driving detection / multimodal data fusion / attention mechanism

Cite this article

Download Citations
Qi ZENG , Shuyi WANG , Yi LIU. A fatigued driving detection method using multimodal data fusion analysis[J]. Journal of Tsinghua University(Science and Technology). 2026, 66(9): 1873-1880 https://doi.org/10.16511/j.cnki.qhdxxb.2026.27.050

References

1
中华人民共和国国家统计局. 中国统计年鉴[M]. 北京: 中国统计出版社, 2024.
National Bureau of Statistics of China. China Statistical Yearbook[M]. Beijing: China Statistics Press, 2024.
2
世界卫生组织. 2023年全球道路安全状况报告[R]. 日内瓦: 世界卫生组织, 2024.
World Health Organization. Global status report on road safety 2023[R]. Geneva: World Health Organization, 2024.
3
PHILIP P F, LE B P, TAILLARD J, et al. Fatigue, alcohol, and serious road crashes in France: Factorial study of national data[J]. BMJ, 2001, 322(7290): 829- 830.
4
LI G, CHUNG W Y. Detection of driver drowsiness using wavelet analysis of heart rate variability and a support vector machine classifier[J]. Sensors, 2013, 13(12): 16494- 16511.
5
WANG M S, JEONG N T, KIM K S, et al. Drowsy behavior detection based on driving information[J]. International Journal of Automotive Technology, 2016, 17(1): 165- 173.
6
魏恒建. 基于多通道信息融合的驾驶员疲劳检测系统研究[D]. 大连: 大连交通大学, 2024.
WEI H J. Research on driver fatigue detection system based on multi-channel information fusion[D]. Dalian: Dalian Jiaotong University, 2024. (in Chinese)
7
李泰国, 张天策, 李超, 等. 基于面部倒立摆模型与信息熵的驾驶员疲劳检测[J]. 交通运输系统工程与信息, 2023, 23(5): 24- 32.
LI T G, ZHANG T C, LI C, et al. Driver fatigue detection based on facial inverted pendulum model and information entropy[J]. Journal of Transportation Systems Engineering and Information Technology, 2023, 23(5): 24- 32.
8
戴诗琪, 曾智勇. 基于深度学习的疲劳驾驶检测算法[J]. 计算机系统应用, 2018, 27(7): 113- 120.
DAI S Q, ZENG Z Y. Fatigue driving detection algorithm based on deep learning[J]. Computer Systems & Applications, 2018, 27(7): 113- 120.
9
周蒙. 考虑驾驶员个体差异的多特征融合疲劳驾驶检测算法研究[D]. 上海: 上海海洋大学, 2025.
ZHOU M. Research on fatigue driving detection algorithm based on multi-feature fusion considering individual driver differences[D]. Shanghai: Shanghai Ocean University, 2025. (in Chinese)
10
张伯辰, 施鑫杰, 霍梅梅. 基于OpenCV的树莓派人脸识别疲劳驾驶检测系统[J]. 现代计算机, 2021, 27(23): 129- 132.
ZHANG B C, SHI X J, HUO M M. Raspberry Pi face recognition fatigue driving detection system based on OpenCV[J]. Modern Computer, 2021, 27(23): 129- 132.
11
MURRAY B J. Subjective and objective assessment of hypersomnolence[J]. Sleep Medicine Clinics, 2017, 12(3): 313- 322.
12
龙伟峰, 胡江碧. 不同工作负荷下驾驶员的心生理特征研究[J]. 交通标准化, 2009(21): 59- 63.
LONG W F, HU J B. Study on driver's cardiac and physiological characteristics under different workloads[J]. Communications Standardization, 2009(21): 59- 63.
13
ROSA M M A, COSTA P Ü, COSTA E A C, et al. Design of a low power and robust VLSI power line interference canceler with optimized arithmetic operators[J]. Analog Integr. Circuits Signal Process, 2022, 112(2): 247- 261.
14
吕建行, 李玉榕, 陈建国, 等. 两步式自适应阈值法滤除心电信号中运动伪迹[J]. 电子学报, 2024, 52(10): 3493- 3506.
LYU J X, LI Y R, CHEN J G, et al. Two-step adaptive threshold method for removing motion artifacts in ECG signals[J]. Acta Electronica Sinica, 2024, 52(10): 3493- 3506.
15
胡枫. 基于马尔科夫模型的短时交通流预测研究[D]. 南京: 南京邮电大学, 2013.
HU F. Research on short-term traffic flow prediction based on Markov model[D]. Nanjing: Nanjing University of Posts and Telecommunications, 2013. (in Chinese)
16
张玺君, 袁占亭, 张红, 等. 交通轨迹大数据预处理方法研究[J]. 计算机工程, 2019, 45(6): 26- 31.
ZHANG X J, YUAN Z T, ZHANG H, et al. Research on preprocessing method of traffic trajectory big data[J]. Computer Engineering, 2019, 45(6): 26- 31.
17
ZHOU Y, BARNES C, LU J W, et al. On the continuity of rotation representations in neural networks[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Long Beach: IEEE, 2019: 5745–5753.
18
OROZCO J, GONG S, XIANG T. Head pose classification in crowded scenes[C]//Proceedings of the British Machine Vision Conference (BMVC). London: British Machine Vision Association, 2009: 1-11.
19
GUPTA A, THAKKAR K, GANDHI V, et al. Nose, eyes and ears: Head pose estimation by locating facial keypoints [C]//ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Brighton, UK: IEEE, 2019: 1977-1981.
20
LIU L Y, KE Z R, HUO J, et al. Head pose estimation through keypoints matching between reconstructed 3D face model and 2D image[J]. Sensors, 2021, 21(5): 1841.
21
ABATE A F, BISOGNI C, CASTIGLIONE A, et al. Head pose estimation: An extensive survey on recent techniques and applications[J]. Pattern Recognition, 2022, 127, 108591.
22
RUIZ N, CHONG E, REHG J M. Fine-grained head pose estimation without keypoints [C]//Proceedings of the IEEE conference on computer vision and pattern recognition workshops, 2018: 2074–2083.
23
VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[J/OL]. arXiv, 2017. DOI: 10.48550/arXiv. 1706.03762. [2025-09-23]. https://arxiv.org/abs/1706.03762.

RIGHTS & PERMISSIONS

All rights reserved. Unauthorized reproduction is prohibited.
PDF(2265 KB)

Accesses

Citation

Detail

Sections
Recommended

/