Location
Hilton Waikoloa Village, Hawaii
Event Website
https://hicss.hawaii.edu/
Start Date
7-1-2025 12:00 AM
End Date
10-1-2025 12:00 AM
Description
Fine-tuning efforts have led to progress in reducing overt, obvious gender and racial biases in the latest generation of large language models (LLMs). Here we study covert, non-obvious bias in LLM-based chat systems. We run a two-stage experiment in the hiring context consisting of resume creation and selection. We use ChatGPT-4o to create resumes for minority, ethnic candidates and majority, baseline candidates. After removal of all identifying markers, we run pair-wise selection tests and find that resumes of majority candidates are stronger, winning contests in 80% of the time. This suggests that racial markers lead to encoding of biases in resume generation in imperceptible ways. Covert biases are difficult to spot and hard to address, but they deserve urgent attention, as the latest models are becoming increasingly capable of inferring user characteristics from conversations, potentially biasing content in unwanted and unexpected ways. We discuss implications and avenues for future research.
Recommended Citation
Riemer, Kai and Peter, Sandra, "Exploring Covert Bias in Large Language Models - Experimental Evidence of Racial Discrimination in Resume Creation and Selection" (2025). Hawaii International Conference on System Sciences 2025 (HICSS-58). 7.
https://aisel.aisnet.org/hicss-58/sj/discrimination/7
Exploring Covert Bias in Large Language Models - Experimental Evidence of Racial Discrimination in Resume Creation and Selection
Hilton Waikoloa Village, Hawaii
Fine-tuning efforts have led to progress in reducing overt, obvious gender and racial biases in the latest generation of large language models (LLMs). Here we study covert, non-obvious bias in LLM-based chat systems. We run a two-stage experiment in the hiring context consisting of resume creation and selection. We use ChatGPT-4o to create resumes for minority, ethnic candidates and majority, baseline candidates. After removal of all identifying markers, we run pair-wise selection tests and find that resumes of majority candidates are stronger, winning contests in 80% of the time. This suggests that racial markers lead to encoding of biases in resume generation in imperceptible ways. Covert biases are difficult to spot and hard to address, but they deserve urgent attention, as the latest models are becoming increasingly capable of inferring user characteristics from conversations, potentially biasing content in unwanted and unexpected ways. We discuss implications and avenues for future research.
https://aisel.aisnet.org/hicss-58/sj/discrimination/7