Understanding where drivers direct their visual attention during driving, known as gaze behavior, is important for developing next-generation advanced driver-assistance systems and improving road safety. This paper addressed this challenge as a semantic identification task from road scenes captured by a vehicle's front-view camera. Specifically, the collocation of gaze points with object semantics was investigated using three distinct vision-based approaches: direct object detection, segmentation-assisted classification, and query-based vision-language models. The results demonstrate that the direct object detection and vision language model–based approach outperformed other approaches, achieving macro F1-scores exceeding 0.84. The larger vision-language model, in particular, exhibited better robustness and performance for identifying small, safety-critical objects such as traffic lights, especially in nighttime conditions. Conversely, the segmentation-assisted paradigm suffers from a part-versus-whole semantic gap, which led to low recall. The results reveal a trade-off between the real-time efficiency of traditional detectors and the contextual understanding and robustness offered by large vision-language models. These findings provide guidance for the design of future human-aware driver monitoring systems.
Citation: Penghao Deng, Jidong J. Yang, Jiachen Bian. Cross-paradigm evaluation of gaze-based object identification for human-centered vehicle interfaces[J]. Applied Computing and Intelligence, 2026, 6(2): 115-137. doi: 10.3934/aci.2026007
Understanding where drivers direct their visual attention during driving, known as gaze behavior, is important for developing next-generation advanced driver-assistance systems and improving road safety. This paper addressed this challenge as a semantic identification task from road scenes captured by a vehicle's front-view camera. Specifically, the collocation of gaze points with object semantics was investigated using three distinct vision-based approaches: direct object detection, segmentation-assisted classification, and query-based vision-language models. The results demonstrate that the direct object detection and vision language model–based approach outperformed other approaches, achieving macro F1-scores exceeding 0.84. The larger vision-language model, in particular, exhibited better robustness and performance for identifying small, safety-critical objects such as traffic lights, especially in nighttime conditions. Conversely, the segmentation-assisted paradigm suffers from a part-versus-whole semantic gap, which led to low recall. The results reveal a trade-off between the real-time efficiency of traditional detectors and the contextual understanding and robustness offered by large vision-language models. These findings provide guidance for the design of future human-aware driver monitoring systems.
| [1] | National Center for Statistics and Analysis, Early estimate of motor vehicle traffic fatalities in 2024, National Highway Traffic Safety Administration, 2025. Available from: https://crashstats.nhtsa.dot.gov/Api/Public/ViewPublication/813710. |
| [2] | WHO, Global Status Report on Road Safety 2023, World Health Organization, 2023. Available from: https://www.who.int/teams/social-determinants-of-health/safety-and-mobility/global-status-report-on-road-safety-2023. |
| [3] |
S. García-Herrero, J. Febres, W. Boulagouas, J. Gutiérrez, M. Saldaña, Assessment of the influence of technology-based distracted driving on drivers' infractions and their subsequent impact on traffic accidents severity, Int. J. Environ. Res. Public Health, 18 (2021), 7155. https://doi.org/10.3390/ijerph18137155 doi: 10.3390/ijerph18137155
|
| [4] |
C. Ahlström, K. Kircher, M. Nyström, B. Wolfe, Eye tracking in driver attention research—how gaze data interpretations influence what we learn, Front. Neuroergonomics, 2 (2021), 778043. https://doi.org/10.3389/fnrgo.2021.778043 doi: 10.3389/fnrgo.2021.778043
|
| [5] | National Center for Statistics and Analysis, Distracted driving in 2023, National Highway Traffic Safety Administration, 2025. Available from: https://crashstats.nhtsa.dot.gov/Api/Public/ViewPublication/813703. |
| [6] |
S. Klauer, J. Ehsani, D. McGehee, M. Manser, The effect of secondary task engagement on adolescents' driving performance and crash risk, J. Adolescent Health, 57 (2015), S36–S43. https://doi.org/10.1016/j.jadohealth.2015.03.014 doi: 10.1016/j.jadohealth.2015.03.014
|
| [7] |
B. Simons-Morton, F. Guo, S. Klauer, J. Ehsani, A. Pradhan, Keep your eyes on the road: young driver crash risk increases according to duration of distraction, J. Adolescent Health, 54 (2014), S61–S67. https://doi.org/10.1016/j.jadohealth.2013.11.021 doi: 10.1016/j.jadohealth.2013.11.021
|
| [8] |
F. Guo, S. Klauer, Y. Fang, J. Hankey, J. Antin, M. Perez, et al., The effects of age on crash risk associated with driver distraction, Int. J. Epidemiol., 46 (2017), 258–265. https://doi.org/10.1093/ije/dyw234 doi: 10.1093/ije/dyw234
|
| [9] |
Y. Dong, Z. Hu, K. Uchimura, N. Murayama, Driver inattention monitoring system for intelligent vehicles: a review, IEEE Trans. Intell. Transp., 12 (2011), 596–614. https://doi.org/10.1109/TITS.2010.2092770 doi: 10.1109/TITS.2010.2092770
|
| [10] |
P. Deng, C. Cui, Z. Cheng, Q. Zhang, Y. Bu, Fatigue damage prognosis of orthotropic steel deck based on data-driven LSTM, J. Constr. Steel Res., 202 (2023), 107777. https://doi.org/10.1016/j.jcsr.2023.107777 doi: 10.1016/j.jcsr.2023.107777
|
| [11] |
P. Deng, J. Yang, T. Yee, Deep learning-based flood detection for bridge monitoring using accelerometer data, Infrastructures, 9 (2024), 140. https://doi.org/10.3390/infrastructures9090140 doi: 10.3390/infrastructures9090140
|
| [12] |
M. Khan, S. Lee, Gaze and eye tracking: techniques and applications in ADAS, Sensors, 19 (2019), 5540. https://doi.org/10.3390/s19245540 doi: 10.3390/s19245540
|
| [13] |
A. Ledezma, V. Zamora, O. Sipele, M. Sesmero, A. Sanchis, Implementing a gaze tracking algorithm for improving advanced driver assistance systems, Electronics, 10 (2021), 1480. https://doi.org/10.3390/electronics10121480 doi: 10.3390/electronics10121480
|
| [14] |
Y. Feng, Y. Chen, J. Zhang, C. Tian, R. Ren, T. Han, et al., Human-centred design of next generation transportation infrastructure with connected and automated vehicles: a system-of-systems perspective, Theor. Iss. Ergon. Sci., 25 (2024), 287–315. https://doi.org/10.1080/1463922X.2023.2182003 doi: 10.1080/1463922X.2023.2182003
|
| [15] |
J. Kim, Sustainable real-time driver gaze monitoring for enhancing autonomous vehicle safety, Sustainability, 17 (2025), 4114. https://doi.org/10.3390/su17094114 doi: 10.3390/su17094114
|
| [16] |
S. Haghzare, J. Campos, A. Mihailidis, Classifying older drivers' gaze behaviour during automated versus non-automated driving: a preliminary step towards detecting mode confusion, Int. J. Hum. -Comput. Int., 40 (2024), 241–254. https://doi.org/10.1080/10447318.2022.2112933 doi: 10.1080/10447318.2022.2112933
|
| [17] |
F. Walker, J. Wang, M. Martens, W. Verwey, Gaze behaviour and electrodermal activity: objective measures of drivers' trust in automated vehicles, Transport. Res. F-Traf., 64 (2019), 401–412. https://doi.org/10.1016/j.trf.2019.05.021 doi: 10.1016/j.trf.2019.05.021
|
| [18] |
Y. Zhu, Y. Geng, R. Huang, X. Zhang, L. Wang, W. Liu, Driving towards the future: exploring human-centered design and experiment of glazing projection display systems for autonomous vehicles, Int. J. Hum. -Comput. Int., 40 (2024), 4087–4102. https://doi.org/10.1080/10447318.2023.2209836 doi: 10.1080/10447318.2023.2209836
|
| [19] |
M. Winlaw, S. Steiner, R. MacKay, A. Hilal, Using telematics data to find risky driver behaviour, Accident Anal. Prev., 131 (2019), 131–136. https://doi.org/10.1016/j.aap.2019.06.003 doi: 10.1016/j.aap.2019.06.003
|
| [20] |
D. Kumar, N. Muhammad, Object detection in adverse weather for autonomous driving through data merging and YOLOv8, Sensors, 23 (2023), 8471. https://doi.org/10.3390/s23208471 doi: 10.3390/s23208471
|
| [21] |
G. Liu, K. Wu, W. Lan, Y. Wu, YOLO-FDCL: improved YOLOv8 for driver fatigue detection in complex lighting conditions, Sensors, 25 (2025), 4832. https://doi.org/10.3390/s25154832 doi: 10.3390/s25154832
|
| [22] |
T. Tammi, J. Pekkanen, B. Cowley, O. Lappi, Quantifying tracking quality during occlusion with an integrated gaze metric anchored to task performance, Sci. Rep., 15 (2025), 31858. https://doi.org/10.1038/s41598-025-17519-8 doi: 10.1038/s41598-025-17519-8
|
| [23] | D. Mardanbegi, T. Langlotz, H. Gellersen, Resolving target ambiguity in 3d gaze interaction through vor depth estimation, Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 2019, 1–12. https://doi.org/10.1145/3290605.3300842 |
| [24] |
J. Wolfe, A. Kosovicheva, B. Wolfe, Normal blindness: when we look but fail to see, Trends Cogn. Sci., 26 (2022), 809–819. https://doi.org/10.1016/j.tics.2022.06.006 doi: 10.1016/j.tics.2022.06.006
|
| [25] |
M. Herslund, N. Jørgensen, Looked-but-failed-to-see-errors in traffic, Accident Anal. Prev., 35 (2003), 885–891. https://doi.org/10.1016/S0001-4575(02)00095-7 doi: 10.1016/S0001-4575(02)00095-7
|
| [26] |
K. Kennedy, J. Bliss, Inattentional blindness in a simulated driving task, Proceedings of the Human Factors and Ergonomics Society Annual Meeting, 57 (2013), 1899–1903. https://doi.org/10.1177/1541931213571423 doi: 10.1177/1541931213571423
|
| [27] |
J. Hoffmann, H. Tosso, M. Santos, J. Justo, A. Malik, A. Rahman, Real-time adaptive object detection and tracking for autonomous vehicles, IEEE Transactions on Intelligent Vehicles, 6 (2021), 450–459. https://doi.org/10.1109/TIV.2020.3037928 doi: 10.1109/TIV.2020.3037928
|
| [28] |
Z. Wang, K. Yao, F. Guo, Driver attention detection based on improved YOLOv5, Appl. Sci., 13 (2023), 6645. https://doi.org/10.3390/app13116645 doi: 10.3390/app13116645
|
| [29] |
N. Youssouf, Traffic sign classification using CNN and detection using faster-RCNN and YOLOV4, Heliyon, 8 (2022), e11792. https://doi.org/10.1016/j.heliyon.2022.e11792 doi: 10.1016/j.heliyon.2022.e11792
|
| [30] | M. Maity, S. Banerjee, S. Chaudhuri, Faster R-CNN and yolo based vehicle detection: a survey, Proceedings of the 5th International Conference on Computing Methodologies and Communication, 2021, 1442–1447. https://doi.org/10.1109/ICCMC51019.2021.9418274 |
| [31] | F. Tonini, N. Dall'Asen, C. Beyan, E. Ricci, Object-aware gaze target detection, Proceedings of IEEE/CVF International Conference on Computer Vision (ICCV), 2023, 21803–21812. https://doi.org/10.1109/ICCV51070.2023.01998 |
| [32] |
B. Mirzaei, H. Nezamabadi-Pour, A. Raoof, R. Derakhshani, Small object detection and tracking: a comprehensive review, Sensors, 23 (2023), 6887. https://doi.org/10.3390/s23156887 doi: 10.3390/s23156887
|
| [33] |
H. Zhao, S. Wang, X. Peng, J. Pan, R. Wang, X. Liu, Road surface semantic segmentation for autonomous driving, PeerJ Comput. Sci., 10 (2024), e2250. https://doi.org/10.7717/peerj-cs.2250 doi: 10.7717/peerj-cs.2250
|
| [34] |
F. Hazzaa, I. Udoidiong, A. Qashou, S. Yousef, Segment anything: a review, Mesopotamian Journal of Computer Science, 2024 (2024), 150–161. https://doi.org/10.58496/MJCSC/2024/012 doi: 10.58496/MJCSC/2024/012
|
| [35] |
Z. Yuan, Principles, applications, and advancements of the segment anything model, Appl. Comput. Eng., 53 (2024), 73–78. https://doi.org/10.54254/2755-2721/53/20241270 doi: 10.54254/2755-2721/53/20241270
|
| [36] | N. Ravi, V. Gabeur, Y. Hu, R. Hu, C. Ryali, T. Ma, et al., Sam 2: segment anything in images and videos, Proceedings of International Conference on Learning Representations (ICLR), 2025, 1–44. |
| [37] |
M. Elhenawy, H. Ashqar, A. Rakotonirainy, T. Alhadidi, A. Jaber, M. Tami, Vision-language models for autonomous driving: clip-based dynamic scene understanding, Electronics, 14 (2025), 1282. https://doi.org/10.3390/electronics14071282 doi: 10.3390/electronics14071282
|
| [38] | A. Keskar, S. Perisetla, R. Greer, Evaluating multimodal vision-language model prompting strategies for visual question answering in road scene understanding, Proceedings of IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW), 2025, 937–946. https://doi.org/10.1109/WACVW65960.2025.00115 |
| [39] | J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, et al., Qwen-vl: a versatile vision-language model for understanding, localization, text reading, and beyond, arXiv: 2308.12966. https://doi.org/10.48550/arXiv.2308.12966 |
| [40] | S. Wang, D. Kim, A. Taalimi, C. Sun, W. Kuo, Learning visual grounding from generative vision and language model, Proceedings of IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2025, 8057–8067. https://doi.org/10.1109/WACV61041.2025.00782 |
| [41] | F. Yu, H. Chen, X. Wang, W. Xian, Y. Chen, F. Liu, et al., Bdd100k: a diverse driving dataset for heterogeneous multitask learning, Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, 2636–2645. https://doi.org/10.1109/CVPR42600.2020.00271 |
| [42] | T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, et al., Microsoft coco: common objects in context, In: Computer Vision-ECCV 2014, Cham: Springer, 2014,740–755. https://doi.org/10.1007/978-3-319-10602-1_48 |
| [43] | M. Lei, S. Li, Y. Wu, H. Hu, Y. Zhou, X. Zheng, et al., Yolov13: real-time object detection with hypergraph-enhanced adaptive visual perception, arXiv: 2506.17733. https://doi.org/10.48550/arXiv.2506.17733 |
| [44] | M. Tan, Q. Le, Efficientnetv2: smaller models and faster training, Proceedings of the 38th International Conference on Machine Learning, 2021, 10096–10106. |