Language does more than describe: on the lack of figurative speech in text-to-image models

Kleinlein, Ricardo; Luna-Jiménez, Cristina; Fernández-Martínez, Fernando

doi:10.48550/arXiv.2210.10578

search hit 18 of 559

Back to Result List

Language does more than describe: on the lack of figurative speech in text-to-image models

Ricardo Kleinlein, Cristina Luna-Jiménez, Fernando Fernández-Martínez

The impressive capacity shown by recent text-to-image diffusion models to generate high-quality pictures from textual input prompts has leveraged the debate about the very definition of art. Nonetheless, these models have been trained using text data collected from content-based labelling protocols that focus on describing the items and actions in an image but neglect any subjective appraisal. Consequently, these automatic systems need rigorous descriptions of the elements and the pictorial style of the image to be generated, otherwise failing to deliver. As potential indicators of the actual artistic capabilities of current generative models, we characterise the sentimentality, objectiveness and degree of abstraction of publicly available text data used to train current text-to-image diffusion models. Considering the sharp difference observed between their language style and that typically employed in artistic contexts, we suggest generative models should incorporate additionalThe impressive capacity shown by recent text-to-image diffusion models to generate high-quality pictures from textual input prompts has leveraged the debate about the very definition of art. Nonetheless, these models have been trained using text data collected from content-based labelling protocols that focus on describing the items and actions in an image but neglect any subjective appraisal. Consequently, these automatic systems need rigorous descriptions of the elements and the pictorial style of the image to be generated, otherwise failing to deliver. As potential indicators of the actual artistic capabilities of current generative models, we characterise the sentimentality, objectiveness and degree of abstraction of publicly available text data used to train current text-to-image diffusion models. Considering the sharp difference observed between their language style and that typically employed in artistic contexts, we suggest generative models should incorporate additional sources of subjective information in their training in order to overcome (or at least to alleviate) some of their current limitations, thus effectively unleashing a truly artistic and creative generation.…

Metadaten
Author:	Ricardo Kleinlein, Cristina Luna-Jiménez ORCiD GND, Fernando Fernández-Martínez
Frontdoor URL	https://opus.bibliothek.uni-augsburg.de/opus4/122681
Parent Title (English):	arXiv
Publisher:	arXiv
Type:	Preprint
Language:	English
Date of Publication (online):	2025/06/04
Year of first Publication:	2022
Publishing Institution:	Universität Augsburg
Release Date:	2025/06/05
Issue:	arXiv:2210.10578
DOI:	https://doi.org/10.48550/arXiv.2210.10578
Institutes:	Fakultät für Angewandte Informatik
	Fakultät für Angewandte Informatik / Institut für Informatik
	Fakultät für Angewandte Informatik / Institut für Informatik / Lehrstuhl für Menschzentrierte Künstliche Intelligenz
Dewey Decimal Classification:	0 Informatik, Informationswissenschaft, allgemeine Werke / 00 Informatik, Wissen, Systeme / 004 Datenverarbeitung; Informatik
Latest Publications (not yet published in print):	Aktuelle Publikationen (noch nicht gedruckt erschienen)

Open Access

Language does more than describe: on the lack of figurative speech in text-to-image models

Export metadata

Statistics

Additional Services