Interpretable Text-to-Image Instruction Following Evaluation
Results reported on the IMAG-EVAL benchmark using the All Skills setting (Counting, Spatial, Size, Color, Emotion, Text, and Cohesiveness).
| Rank | Model | Counting | Spatial | Size | Emotion | Color | Cohesiveness | Text (WER ↓) |
|---|
IMAG-EVAL is a controlled benchmark designed to evaluate instruction-following capabilities in Text-to-Image generation models. Unlike existing benchmarks that primarily vary prompt length, IMAG-EVAL independently manipulates the number of grounded instances and compositional constraints, enabling fine-grained analysis of multimodal instruction following.