Location
Hilton Waikoloa Village, Hawaii
Event Website
https://hicss.hawaii.edu/
Start Date
7-1-2025 12:00 AM
End Date
10-1-2025 12:00 AM
Description
In recent years, large language models (LLMs) have made substantial strides in mimicking human language and coherently presenting information. However, researchers continue to debate the accuracy and robustness of LLMs’ reasoning abilities. The reasoning abilities of thirteen LLMs were tested on two long-text analogy datasets, named Rattermann and Wharton, which required them to rank a series of stories from most analogous to least analogous compared to a source story. On the Rattermann dataset, GPT-4 obtained the highest accuracy of 70%. As a whole, LLMs seem to struggle with over-emphasizing similar story entities (characters and settings) and a lack of awareness of higher-order relationship(s) between stories. LLMs struggled more with the Wharton dataset, with the highest accuracy achieved being 46.4% by GPT-4o, and all but nine LLMs performing below random chance accuracy. Although LLMs are improving, they still struggle with higher-cognitive tasks such as analogical reasoning.
Recommended Citation
Combs, Kara; Bihl, Trevor; Howlett, Spencer; and Adams, Yuki, "Zero-shot Comparison of Large Language Models (LLMs) Reasoning Abilities on Long-text Analogies" (2025). Hawaii International Conference on System Sciences 2025 (HICSS-58). 8.
https://aisel.aisnet.org/hicss-58/da/nlp_and_llms/8
Zero-shot Comparison of Large Language Models (LLMs) Reasoning Abilities on Long-text Analogies
Hilton Waikoloa Village, Hawaii
In recent years, large language models (LLMs) have made substantial strides in mimicking human language and coherently presenting information. However, researchers continue to debate the accuracy and robustness of LLMs’ reasoning abilities. The reasoning abilities of thirteen LLMs were tested on two long-text analogy datasets, named Rattermann and Wharton, which required them to rank a series of stories from most analogous to least analogous compared to a source story. On the Rattermann dataset, GPT-4 obtained the highest accuracy of 70%. As a whole, LLMs seem to struggle with over-emphasizing similar story entities (characters and settings) and a lack of awareness of higher-order relationship(s) between stories. LLMs struggled more with the Wharton dataset, with the highest accuracy achieved being 46.4% by GPT-4o, and all but nine LLMs performing below random chance accuracy. Although LLMs are improving, they still struggle with higher-cognitive tasks such as analogical reasoning.
https://aisel.aisnet.org/hicss-58/da/nlp_and_llms/8