AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation

arXiv · AI, language, vision and robotics · article · Sep 11, 2026 · UTC

Multi-modal image generation, particularly subject-driven customization, has garnered growing attention in recent years. Despite the rapid advancement of generative models, their evaluation remains largely lagging. Existing methods, whether embedding-based or Multi-modal Large Language Model (MLLM)-based, evaluate alignment with each modal condition in isolation, which contradicts the simultaneous condition alignment objective of multi-modal image generation, leading to poor consistency with human judgments. To address this challenge, we propose UFO, the first unified framework for omni-condit

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T18:42:18.733Z. This is not the publication date.