Paper Type

Complete

Abstract

In today’s visual-first digital environment, brands increasingly rely on video to communicate experiential product qualities. However, while video can deliver rich sight and sound, it cannot directly transmit touch, taste, or smell that often shape product experience. This study examines whether visuals can nonetheless evoke these “missing” senses through cross-sensory imagery. Drawing on cross-modal correspondence theory, we propose that three classes of visual cues, namely semantic objects, vicarious consumption depictions, and sensory amplification, elicit smell-, taste-, and touch-related imagery among viewers. We test this framework using 63,389 product-relevant Instagram brand videos and 5,305,001 viewer comments. We quantify visual cues with computer vision and vision-language models and capture viewer-expressed cross-sensory imagery from comments. Results show that each cue class predicts higher cross-sensory imagery and greater engagement.

Paper Number

1521

Comments

VCC

Share

COinS
Best Paper Nominee badge
 
Aug 15th, 12:00 AM

Tactile, Taste, and Aroma in Pixels: Eliciting Multi-Sensory Imagery Through Videos

In today’s visual-first digital environment, brands increasingly rely on video to communicate experiential product qualities. However, while video can deliver rich sight and sound, it cannot directly transmit touch, taste, or smell that often shape product experience. This study examines whether visuals can nonetheless evoke these “missing” senses through cross-sensory imagery. Drawing on cross-modal correspondence theory, we propose that three classes of visual cues, namely semantic objects, vicarious consumption depictions, and sensory amplification, elicit smell-, taste-, and touch-related imagery among viewers. We test this framework using 63,389 product-relevant Instagram brand videos and 5,305,001 viewer comments. We quantify visual cues with computer vision and vision-language models and capture viewer-expressed cross-sensory imagery from comments. Results show that each cue class predicts higher cross-sensory imagery and greater engagement.

When commenting on articles, please be friendly, welcoming, respectful and abide by the AIS eLibrary Discussion Thread Code of Conduct posted here.