🤖 AI 资讯

· ·
← 返回列表

CompArt: Operationalizing Aesthetic Alignment in Text-to-Image Generation via Principles of Art

arXiv cs.AI2026-09-17 04:00:00大模型,多模态,办公效率,扩散模型,强化学习,预训练,提示工程,模型安全对齐,论文,开发者生态原文 ↗

arXiv:2503.12018v2 Announce Type: replace-cross

Abstract: Text-to-Image (T2I) diffusion models have made rapid progress on semantic alignment (generating what is described in the prompt), yet users still lack reliable control over aesthetic composition (how visual elements are put together). Prior work often treats aesthetics as a single, preference-driven notion (e.g., "high quality", "detailed", "breathtaking"), which does not map cleanly to compositional intent.

We propose Aesthetic Alignment: aligning generated images to explicit, user-specified compositional constraints. We operationalize these constraints using the Principles of Art (PoA)-e.g., Balance, Rhythm, and Emphasis-commonly used in art education to describe composition. To support this task, we introduce CompArt, a dataset of 80,032 WikiArt images augmented with captions and PoA analyses produced by a multimodal LLM under structured prompting. We further propose ArtDapter, a lightweight and disentangled adapter that enables steering a pretrained T2I model along 10 PoA dimensions while retaining the base model's semantic capability. Experiments on CompArt show improved adherence to PoA controls over strong baselines under a dual evaluation protocol.