Improve Image Text Descriptions using Large Language Models

Бесплатный доступ

In this paper, a multi-stage approach to improve text queries (prompts) for image generation is proposed and comprehensively investigated. First, the GPT-2 model, pre-trained on 18,000 raw query-quality query pairs from the Lexica.art platform, automatically expands and refines the original prompts to produce more detailed and semantically accurate images when generated by the diffusion network Stable Diffusion 1.5. Next, the Image Captioning task (BLIP2) and four large language models (DeepSeek, Grok, ChatGPT, YandexGPT) are used to compare the quality of signature expansion, demonstrating different stylistic strategies for augmenting initial descriptions. The proposed "Captioning → Prompt Enhancer (Mistral) → Stable Diffusion+LoRA" Pipeline additionally includes a pre-training of Mistral's own model on BLEU, METEOR and CIDEr metrics, providing a steady increase in the quality of textual descriptions (BLEU from 0.12 to 0.435 after 500 epochs) and a significant reduction in the FID metric for image generation (from 0.3482 to 0.1873). In the final stage, the LoRA modules embedded in UNet and the Stable Diffusion text encoder allow efficient learning of the generation of previously "unknown" objects (rare or fictional), reducing FID to 0.172 at rank = 64. An expert survey (134 respondents) confirmed the visual preference of images generated by optimized queries, demonstrating the potential of the proposed technique to improve the quality of multimodal systems.

image generation \ large language models \ prompt \ text enhancement \ multimodal models \ LoRA

Короткий адрес: https://sciup.org/140316471

IDS: 140316471   |   DOI: 10.18287/COJ1860