Location
Hilton Waikoloa Village, Hawaii
Event Website
https://hicss.hawaii.edu/
Start Date
7-1-2025 12:00 AM
End Date
10-1-2025 12:00 AM
Description
Many visual generative artificial intelligence (AI) models use textual “prompts” as input(s) to guide the development of the resulting image(s). Converting text to images utilizes pragmatics and semantics, which can make an impact on the output. To facilitate more precise prompting, we propose the three-dimensional vector space of textual similarity which uses textual representation, auditory representation, and meaning similarity as its axes. Next, we show that meaning similarity between two words does not necessarily yield visual similarity between corresponding AI-generated images of those words. We quantitively justify this by leveraging eight image generators to generate images for abstract and concrete synonyms, antonyms, and hypernyms-hyponym pairs and compare their image-image CLIPScores to their corresponding text-text CLIPScores. Across all models and relationship types the average similarity comparing text-text and image-image similarity decreased from 92.8% to 70.1% for synonyms, 89% to 58.9% for antonyms, and 85.6% to 68.1% for hypernym-hyponym pairs.
Recommended Citation
Combs, Kara and Bihl, Trevor, "The Visual Analogs of Linguistic Concepts and Their Implications on Generative AI" (2025). Hawaii International Conference on System Sciences 2025 (HICSS-58). 4.
https://aisel.aisnet.org/hicss-58/cl/technological_advancements/4
The Visual Analogs of Linguistic Concepts and Their Implications on Generative AI
Hilton Waikoloa Village, Hawaii
Many visual generative artificial intelligence (AI) models use textual “prompts” as input(s) to guide the development of the resulting image(s). Converting text to images utilizes pragmatics and semantics, which can make an impact on the output. To facilitate more precise prompting, we propose the three-dimensional vector space of textual similarity which uses textual representation, auditory representation, and meaning similarity as its axes. Next, we show that meaning similarity between two words does not necessarily yield visual similarity between corresponding AI-generated images of those words. We quantitively justify this by leveraging eight image generators to generate images for abstract and concrete synonyms, antonyms, and hypernyms-hyponym pairs and compare their image-image CLIPScores to their corresponding text-text CLIPScores. Across all models and relationship types the average similarity comparing text-text and image-image similarity decreased from 92.8% to 70.1% for synonyms, 89% to 58.9% for antonyms, and 85.6% to 68.1% for hypernym-hyponym pairs.
https://aisel.aisnet.org/hicss-58/cl/technological_advancements/4