Abstract

Large language models (LLMs) are increasingly deployed in interactive settings where users provide feedback in natural language, often with an emotional component. This paper examines whether such feedback is associated with systematic changes in LLM outputs. Using 42 Polish prompts across six everyday domains and 20 additional feedback prompts, we compare responses from GPT-4o mini, Mistral Small 24B Instruct, and Phi-4 before and after emotionally valenced contextual prompting. We analyze response length, type-token ratio, syllables per word, and Jaccard similarity. A survey of 50 participants evaluates perceived changes in tone, structure, and usefulness. Results provide preliminary evidence that emotionally valenced feedback modulates verbosity and lexical diversity, and that human raters prefer output produced after feedback. We interpret these findings as behavioral modulation driven by context rather than adaptation at the parameter level and discuss implications for information systems design.

Recommended Citation

Skierkowska, M., Kurowski, A., Zaporowski, S. & Kostek, B.(2026). Contextual Effects of Emotionally Valenced Feedback Prompts on LLM Response Style: A Pilot Study. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.54

Paper Type

Poster

DOI

10.62036/ISD.2026.54

Share

COinS
 

Contextual Effects of Emotionally Valenced Feedback Prompts on LLM Response Style: A Pilot Study

Large language models (LLMs) are increasingly deployed in interactive settings where users provide feedback in natural language, often with an emotional component. This paper examines whether such feedback is associated with systematic changes in LLM outputs. Using 42 Polish prompts across six everyday domains and 20 additional feedback prompts, we compare responses from GPT-4o mini, Mistral Small 24B Instruct, and Phi-4 before and after emotionally valenced contextual prompting. We analyze response length, type-token ratio, syllables per word, and Jaccard similarity. A survey of 50 participants evaluates perceived changes in tone, structure, and usefulness. Results provide preliminary evidence that emotionally valenced feedback modulates verbosity and lexical diversity, and that human raters prefer output produced after feedback. We interpret these findings as behavioral modulation driven by context rather than adaptation at the parameter level and discuss implications for information systems design.