Abstract

Large multimodal models recently achieved impressive performance on a wide range of vision-language tasks, but training such models from scratch is computationally expensive and hardware demanding, which motivates fine-tuning and model‑reuse strategies. In this work, we present a practical example of reusing a pretrained image-text-to-text model (LLaVA) for predicting human gaze fixation heatmaps. Instead of instruction tuning with prompt and response pairs, we attach a gaze prediction model. In the presented method, we operate directly on the hidden representations of a frozen multimodal backbone that runs in a quantized configuration to reduce GPU memory and enable training on consumer-grade hardware. We also demonstrate that multimodal features can be effectively repurposed to map internal states to human gaze fixation from the LLaVA model, which was originally developed for text generation, after a relatively brief training period.

Recommended Citation

Najgebauer, P., Scherer, R. & Skrzyński, M.(2026). Reusing a Multimodal LLaVA Model for Rapid Learning of Human Gaze Prediction. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.93

Paper Type

Short Paper

DOI

10.62036/ISD.2026.93

Share

COinS
 

Reusing a Multimodal LLaVA Model for Rapid Learning of Human Gaze Prediction

Large multimodal models recently achieved impressive performance on a wide range of vision-language tasks, but training such models from scratch is computationally expensive and hardware demanding, which motivates fine-tuning and model‑reuse strategies. In this work, we present a practical example of reusing a pretrained image-text-to-text model (LLaVA) for predicting human gaze fixation heatmaps. Instead of instruction tuning with prompt and response pairs, we attach a gaze prediction model. In the presented method, we operate directly on the hidden representations of a frozen multimodal backbone that runs in a quantized configuration to reduce GPU memory and enable training on consumer-grade hardware. We also demonstrate that multimodal features can be effectively repurposed to map internal states to human gaze fixation from the LLaVA model, which was originally developed for text generation, after a relatively brief training period.