Abstract

This work analyzes the vulnerability of Large Language Models (LLMs) deployed in voice-based conversational systems. While most prior research focuses on text-based interfaces, less attention has been given to voice-based assistants, where user input is transcribed from speech and responses must conform to structured output constraints. We compare a simple, intuitively written prompt with a prompt engineered according to best practices identified in prior research. We also evaluate the impact of incorporating additional behavioral policies directly into the prompt. Our results show that increasing prompt template complexity significantly improves structured-output compliance and reduces injection attack success rates. For language-forcing attacks, the attack's success rate decreased from 56.7% for a simple prompt without policies to 14.2% for a complex prompt with embedded policies. However, some format-override attacks remain highly successful, highlighting the limitations of prompt-only defense mechanisms.

Recommended Citation

Kręt, C. & Wieczorkowska, A.(2026). Evaluating Prompt-Level Defenses Against Injection in Voice-Based LLM Systems. In M. Valenta, B. Mannová, R. Pergl, A. Przybylek, M. Lang, H. Linger, C. Schneider, N. Iivari, & E. Insfran (Eds.), Making ISD Sustainable: Reloaded with AI and Automation (ISD2026 Proceedings). Prague, Czech Republic: Czech Technical University in Prague. ISBN: 978-80-01-07585-2. https://doi.org/10.62036/ISD.2026.64

Paper Type

Poster

DOI

10.62036/ISD.2026.64

Share

COinS
 

Evaluating Prompt-Level Defenses Against Injection in Voice-Based LLM Systems

This work analyzes the vulnerability of Large Language Models (LLMs) deployed in voice-based conversational systems. While most prior research focuses on text-based interfaces, less attention has been given to voice-based assistants, where user input is transcribed from speech and responses must conform to structured output constraints. We compare a simple, intuitively written prompt with a prompt engineered according to best practices identified in prior research. We also evaluate the impact of incorporating additional behavioral policies directly into the prompt. Our results show that increasing prompt template complexity significantly improves structured-output compliance and reduces injection attack success rates. For language-forcing attacks, the attack's success rate decreased from 56.7% for a simple prompt without policies to 14.2% for a complex prompt with embedded policies. However, some format-override attacks remain highly successful, highlighting the limitations of prompt-only defense mechanisms.