Location

Hilton Waikoloa Village, Hawaii

Event Website

https://hicss.hawaii.edu/

Start Date

7-1-2025 12:00 AM

End Date

10-1-2025 12:00 AM

Description

Many visual generative artificial intelligence (AI) models use textual “prompts” as input(s) to guide the development of the resulting image(s). Converting text to images utilizes pragmatics and semantics, which can make an impact on the output. To facilitate more precise prompting, we propose the three-dimensional vector space of textual similarity which uses textual representation, auditory representation, and meaning similarity as its axes. Next, we show that meaning similarity between two words does not necessarily yield visual similarity between corresponding AI-generated images of those words. We quantitively justify this by leveraging eight image generators to generate images for abstract and concrete synonyms, antonyms, and hypernyms-hyponym pairs and compare their image-image CLIPScores to their corresponding text-text CLIPScores. Across all models and relationship types the average similarity comparing text-text and image-image similarity decreased from 92.8% to 70.1% for synonyms, 89% to 58.9% for antonyms, and 85.6% to 68.1% for hypernym-hyponym pairs.

Share

COinS
 
Jan 7th, 12:00 AM Jan 10th, 12:00 AM

The Visual Analogs of Linguistic Concepts and Their Implications on Generative AI

Hilton Waikoloa Village, Hawaii

Many visual generative artificial intelligence (AI) models use textual “prompts” as input(s) to guide the development of the resulting image(s). Converting text to images utilizes pragmatics and semantics, which can make an impact on the output. To facilitate more precise prompting, we propose the three-dimensional vector space of textual similarity which uses textual representation, auditory representation, and meaning similarity as its axes. Next, we show that meaning similarity between two words does not necessarily yield visual similarity between corresponding AI-generated images of those words. We quantitively justify this by leveraging eight image generators to generate images for abstract and concrete synonyms, antonyms, and hypernyms-hyponym pairs and compare their image-image CLIPScores to their corresponding text-text CLIPScores. Across all models and relationship types the average similarity comparing text-text and image-image similarity decreased from 92.8% to 70.1% for synonyms, 89% to 58.9% for antonyms, and 85.6% to 68.1% for hypernym-hyponym pairs.

https://aisel.aisnet.org/hicss-58/cl/technological_advancements/4