AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability

📄 arXiv: 2608.26713v1 📥 PDF

作者: Xuanwei Hu, Haoyu Dong, Kejun Wu, Tianyi Liu, Jianjun Gao

分类: cs.CV, cs.AI

发布日期: 2026-08-27

备注: 10 pages, 4 figures, 6 tables. Supplementary material included


💡 一句话要点

提出AesCanvas以解决图像美学评估中的上下文适用性问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 图像美学评估 多模态大语言模型 上下文适用性 批评生成 文化背景

📋 核心要点

  1. 现有的美学评估方法主要集中于内在视觉质量,缺乏对图像在特定上下文中的适用性评估。
  2. 本文提出AesCanvas,通过CritiqueCanvas和ContextCanvas两个组件,提供多维度的美学批评和上下文适用性评估。
  3. 实验结果表明,传统的评估指标无法全面捕捉批评质量,且美学专门模型在上下文适用性上表现不如强大的通用模型。

📝 摘要(中文)

近年来,多模态大语言模型(MLLMs)的进展使得图像美学评估(IAA)超越了简单的分数评估,向可解释的批评和指导发展。然而,现有基准主要评估内在视觉质量或固定领域标准,未能探讨图像在特定目的、受众、文化背景或领域惯例下的适用性。本文提出了AesCanvas,一个统一的套件,包含两个互补组件:CritiqueCanvas和ContextCanvas,前者支持跨摄影、绘画和虚拟图像的多维度长形式批评,后者评估在现实使用场景中的上下文美学适用性。通过统一协议,我们评估了多种MLLMs,结果显示批评生成与上下文敏感判断之间存在明显差异。

🔬 方法详解

问题定义:本文旨在解决现有图像美学评估方法未能考虑图像在特定上下文中的适用性这一问题。现有方法主要关注内在视觉质量,缺乏对文化和使用场景的考量。

核心思路:AesCanvas通过引入CritiqueCanvas和ContextCanvas两个组件,提供了一个全面的评估框架,支持多维度的批评和上下文适用性分析。这样的设计使得评估不仅限于视觉质量,还包括文化和情境因素。

技术框架:AesCanvas的整体架构包括两个主要模块:CritiqueCanvas,包含519,136个指令-响应对,支持长形式批评;ContextCanvas,包含301个专家审核的使用场景,评估上下文美学适用性。

关键创新:AesCanvas的最大创新在于将批评生成与上下文适用性评估分开,强调了文化背景和使用场景对美学评估的重要性。这与现有方法的单一视觉质量评估形成了鲜明对比。

关键设计:在设计中,CritiqueCanvas使用了多维度的指令-响应对,而ContextCanvas则依赖于专家审核的真实场景,确保评估的准确性和可靠性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,传统的基于参考的词汇和语义指标仅部分捕捉批评质量,而美学专门模型在ContextCanvas上的表现显著落后于强大的通用模型。这表明美学专门化并不一定能有效转移到上下文适用性评估上。

🎯 应用场景

AesCanvas的研究成果在多个领域具有广泛的应用潜力,包括广告设计、艺术创作、社交媒体内容生成等。通过提供更为细致的美学评估,能够帮助创作者和设计师更好地理解其作品在特定文化和情境下的适用性,从而提升作品的影响力和接受度。

📄 摘要(原文)

Recent advances in Multimodal Large Language Models (MLLMs) have extended Image Aesthetic Assessment (IAA) beyond scalar scores toward interpretable critique and guidance. Yet existing benchmarks mainly assess intrinsic visual quality or fixed domain criteria, leaving open whether an appealing image is appropriate for a specific purpose, audience, cultural setting, or domain convention. We introduce AesCanvas, a unified suite with two complementary components: CritiqueCanvas with 519,136 instruction-response pairs from 54,300 images supports long-form, multi-dimensional critique across photography, painting, and virtual imagery, whereas ContextCanvas with 301 expert-reviewed use scenarios evaluates contextual aesthetic suitability in realistic use scenarios. Under a unified protocol, we evaluate closed-source frontier, open-weight general, and aesthetic-specific MLLMs. Results reveal a clear separation between critique generation and context-sensitive judgment: reference-based lexical and semantic metrics only partially capture critique quality, while aesthetic specialists remain competitive on selected critique metrics yet substantially lag strong general-purpose MLLMs on ContextCanvas. Further analyses show that aesthetic specialization does not reliably transfer to contextual suitability and that model decisions may fail to track or ground themselves in decisive contextual visual cues. These findings establish culturally situated, evidence-grounded suitability as a distinct objective for aesthetic modeling.