Abstract

Crisis mapping platforms such as Ushahidi depend on the rapid categorization of crowdsourced Requests for Help (RFH) to coordinate disaster response. As large language models (LLMs) grow in capability, a critical question emerges: under what conditions can these models augment or substitute human categorizers, specifically registered nurses (RNs) and crowd volunteers, in real-world crisis settings? Using crisis messages from the 2010 Haiti and Chile earthquakes, we evaluate multiple LLMs’ categorization decisions against those of RN and crowd volunteers across crisis categories, analyzing agreement patterns and category-level decision variability to establish conditions under which LLMs can reliably support substitution versus augmentation. Across two studies, we formulate a set of design principles that show LLM integration should be conditionally configured based on crisis category and annotator role. LLMs require human supervision in categories involving community needs and context-dependent interpretation, such as vital lines, public health, and infrastructure damage, where agreement is lower. In contrast, LLMs can substitute human annotators in categories where agreement is higher, including certain individual need contexts for expert RNs and selected community-level categories such as security and services available for crowd volunteers. At the same time, LLMs require human oversight in cases involving specific or locally grounded details, particularly where volunteers possess contextual knowledge. Our work furthers collective sensemaking in crisis environments to human-AI collaborative contexts and reframes automation versus augmentation as a sensemaking problem in which AI supports human sensemaking rather than replacing it.

DOI

10.17705/1jais.01022

Share

COinS