IMAG-EVAL

Interpretable Text-to-Image Instruction Following Evaluation

🏆 EMNLP 2026 📊 1,140 Prompts 🔬 8,842 Rules
💻 GitHub 🤗 Dataset 📄 Paper

⚙️ Evaluation Settings

Benchmark All Skills
Skills Counting + Spatial + Size + Color + Emotion + Text + Cohesiveness
Resolution 1024 × 1024
Annotation Type Manual
Seed 42
Inference Settings Default values reported in the corresponding model documentation

🏆 Leaderboard

Results reported on the IMAG-EVAL benchmark using the All Skills setting (Counting, Spatial, Size, Color, Emotion, Text, and Cohesiveness).

Rank Model Counting Spatial Size Emotion Color Cohesiveness Text (WER ↓)

📊 Evaluated Skills

🔢 Counting
📍 Spatial Relations
📏 Size Relations
🎨 Color Attribution
😊 Emotion Attribution
🔤 Text Rendering
🧩 Cohesiveness

📖 About IMAG-EVAL

IMAG-EVAL is a controlled benchmark designed to evaluate instruction-following capabilities in Text-to-Image generation models. Unlike existing benchmarks that primarily vary prompt length, IMAG-EVAL independently manipulates the number of grounded instances and compositional constraints, enabling fine-grained analysis of multimodal instruction following.