Abstract

Background: Large language models (LLMs) are increasingly used to simulate human responses in social science. However, empirical research remains limited on whether expert-guided prompting adds value beyond model-native reasoning in newer reasoning-oriented models, how well LLMs reproduce psychological–behavioral associations, how robust simulation performance is across repeated runs and models, and how locally deployable models compare with cloud-based models.

Method: Using non-public Japanese survey data from 600 respondents to reduce the risk of pretraining exposure, we compared GPT-5, GPT-4o, DeepSeek-R1-Distill-Qwen, and GPT-OSS under Standard-CoT (unguided model-native reasoning) and holistic-to-analytic Expert-CoT. Across repeated runs, we evaluated behavioral fidelity and the reproduction of associations between psychological constructs and platform use. DeepSeek-R1-Distill-Qwen reasoning texts were assessed through LLM-as-a-judge evaluation and small-scale human validation.

Results: DeepSeek-R1-Distill-Qwen showed the most balanced distributional fidelity. GPT-OSS achieved the highest binary recall and F1-score but tended to overpredict platform use, whereas GPT-4o achieved the highest binary precision but tended to underpredict users. All models largely preserved the positive direction of the two psychological–behavioral associations but underestimated anticipated regret and overestimated face orientation relative to the human benchmark; GPT-5 showed the largest deviations. Expert-CoT generally improved aggregate distributional agreement and reasoning-text quality but did not uniformly improve psychological association reproduction. In the human validation sample, Expert-CoT elicited stronger activating positive affect without significantly increasing trust.

Conclusion: These findings support a multidimensional view of simulation fidelity encompassing aggregate behavioral patterns, psychological–behavioral associations, and repeated-run robustness. Accordingly, effective LLM-based behavioral simulation requires aligning model choice, prompting strategy, and evaluation approach with the simulation objective. Overall, the study provides empirical evidence and practical guidance for the design and evaluation of LLM-based behavioral simulations in research and decision-support contexts.

[ Online Supplemental Materials ]

Share

COinS