Abstract
Identifying a fixation is a crucial step of an eye-tracking pipeline, especially for scenarios where a user might need to operate a device hands-free, either due to mobility impairments or because they are in a situation where manual operation is not feasible. However, distinguishing between a prolonged fixation and a random eye movement is often done arbitrarily, with fixed thresholds applied to the gaze points. Our work proposes a user-tailored fixation identification method, calculating the gaze's dispersion and velocity to effectively discriminate between fixations and saccades. To do so, a solid gaze estimation is needed: we fine tuned a multi-Mobile Vision Transformer model on the GazeCapture dataset, obtaining lower errors than the reference iTracker model. Finally, we adapted the standard 13-points calibration procedure to incorporate the computation of the dispersion and velocity of the gaze points based on the I-VT and I-DT algorithms, obtaining calibration errors as low as 0.51±0.12 cm.
Paper Type
Poster
DOI
10.62036/ISD.2026.69
Handless Interaction with Technology: Improving Fixation Point Estimate with Multi-MobileViT and Gaze Dynamics
Identifying a fixation is a crucial step of an eye-tracking pipeline, especially for scenarios where a user might need to operate a device hands-free, either due to mobility impairments or because they are in a situation where manual operation is not feasible. However, distinguishing between a prolonged fixation and a random eye movement is often done arbitrarily, with fixed thresholds applied to the gaze points. Our work proposes a user-tailored fixation identification method, calculating the gaze's dispersion and velocity to effectively discriminate between fixations and saccades. To do so, a solid gaze estimation is needed: we fine tuned a multi-Mobile Vision Transformer model on the GazeCapture dataset, obtaining lower errors than the reference iTracker model. Finally, we adapted the standard 13-points calibration procedure to incorporate the computation of the dispersion and velocity of the gaze points based on the I-VT and I-DT algorithms, obtaining calibration errors as low as 0.51±0.12 cm.
Recommended Citation
Fiani, F., Carnebella, M., Di Lorenzo, F., Scherer, M. & Napoli, C.(2026). Handless Interaction with Technology: Improving Fixation Point Estimate with Multi-MobileViT and Gaze Dynamics. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.69