Paper Type
Short
Paper Number
PACIS2026-1827
Description
Generative AI is shifting from text-based chatbots to real-time voice assistants, expanding the modalities through which users co-create with AI in rule-bound tasks. Whether voice’s expressive ease still serves users when the task demands deliberate specification and evaluation, and through what channel, remains unclear. This research-in-progress study examines how interaction modality shapes user engagement, output quality, and trust calibration in rule-bound human-AI co-creation tasks. A pilot study shows that voice significantly reduces output quality and directionally widens a Trust Calibration Gap, accompanied by fragmented prompting and lower deliberate engagement; domain expertise directionally buffers these effects. We term this the Fluency-Control Paradox: voice’s perceptual fluency dissolves the friction that ordinarily forces careful thought while supplying a counterfeit signal that careful thought has already occurred. The study contributes a dual-engagement account of modality effects and identifies expertise as a first-stage boundary condition for modality-aware GenAI interface design.
Recommended Citation
Shi, Zhuzhi and Xue, Mei, "Voice Modality, Engagement, and Trust Calibration in Human-AI Co-Creation" (2026). PACIS 2026 Proceedings. 10.
https://aisel.aisnet.org/pacis2026/ai_fow/ai_fow/10
Voice Modality, Engagement, and Trust Calibration in Human-AI Co-Creation
Generative AI is shifting from text-based chatbots to real-time voice assistants, expanding the modalities through which users co-create with AI in rule-bound tasks. Whether voice’s expressive ease still serves users when the task demands deliberate specification and evaluation, and through what channel, remains unclear. This research-in-progress study examines how interaction modality shapes user engagement, output quality, and trust calibration in rule-bound human-AI co-creation tasks. A pilot study shows that voice significantly reduces output quality and directionally widens a Trust Calibration Gap, accompanied by fragmented prompting and lower deliberate engagement; domain expertise directionally buffers these effects. We term this the Fluency-Control Paradox: voice’s perceptual fluency dissolves the friction that ordinarily forces careful thought while supplying a counterfeit signal that careful thought has already occurred. The study contributes a dual-engagement account of modality effects and identifies expertise as a first-stage boundary condition for modality-aware GenAI interface design.
Comments
02-FutureofWork