Paper Type
Complete
Abstract
The objective of this exploratory pilot study was to investigate the preliminary effectiveness of the wisdom of crowds approach in the context of AI agents based on Large Language Models (LLMs) to improve fraud detection in emails. To this end, a customized digital platform was developed to collect assessments through which responses from 122 human agents and 505 autonomous agents regarding the likelihood that emails were fraudulent were gathered. Diversity among the autonomous agents was introduced by varying the LLMs’ temperature parameter, which influences the level of randomness and propensity for hallucinations. Individual assessments were aggregated to compare collective versus individual performance. The initial results suggest that, within this proof-of-concept, aggregating the evaluation of a collective of LLM agents, even when individual agents exhibited a tendency toward hallucination, has the potential to lead to enhanced and more robust fraud detection, mitigating individual errors and leveraging distributed knowledge. Although many AI agents displayed a somewhat “paranoid” bias, suggesting the inheritance of human behavioral patterns, the AI “collective” achieved a notably strong performance, offering better predictions than individual agents in almost all cases.
Paper Number
1719
Recommended Citation
dos Santos Salles, Daniel and Graeml, Alexandre R., "Echoes of Humanity: The Collective Mind of LLM Agents" (2026). AMCIS 2026 Proceedings. 8.
https://aisel.aisnet.org/amcis2026/ai_aiaa/ai_aiaa/8
Echoes of Humanity: The Collective Mind of LLM Agents
The objective of this exploratory pilot study was to investigate the preliminary effectiveness of the wisdom of crowds approach in the context of AI agents based on Large Language Models (LLMs) to improve fraud detection in emails. To this end, a customized digital platform was developed to collect assessments through which responses from 122 human agents and 505 autonomous agents regarding the likelihood that emails were fraudulent were gathered. Diversity among the autonomous agents was introduced by varying the LLMs’ temperature parameter, which influences the level of randomness and propensity for hallucinations. Individual assessments were aggregated to compare collective versus individual performance. The initial results suggest that, within this proof-of-concept, aggregating the evaluation of a collective of LLM agents, even when individual agents exhibited a tendency toward hallucination, has the potential to lead to enhanced and more robust fraud detection, mitigating individual errors and leveraging distributed knowledge. Although many AI agents displayed a somewhat “paranoid” bias, suggesting the inheritance of human behavioral patterns, the AI “collective” achieved a notably strong performance, offering better predictions than individual agents in almost all cases.
When commenting on articles, please be friendly, welcoming, respectful and abide by the AIS eLibrary Discussion Thread Code of Conduct posted here.
Comments
SIG AIAA