How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment
作者: Guang Yang, Fengchen Liu, Alex Wang, Homa Hosseinmardi, Amir Ghasemian
分类: cs.CR, cs.AI, cs.CL
发布日期: 2026-08-12
备注: 41 pages, 31 figures, 9 tables. Preprint
💡 一句话要点
构建多模态基准以揭示中国源模型的国家对齐偏差
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 多模态系统 国家对齐 视觉语言模型 信息重构 政治敏感话题
📋 核心要点
- 现有的多模态系统在国家对齐扭曲方面缺乏系统性研究,尤其是如何在视觉语言模型中体现这一现象。
- 论文通过构建一个包含200个条目的基准,系统评估了九个视觉语言模型在不同提示下的表现,揭示了国家对齐的隐性重构现象。
- 实验结果显示,中文提示显著提高了国家对齐框架的概率,中国源模型在信息重构方面的表现优于非中国模型,且拒绝率下降。
📝 摘要(中文)
本研究系统性地探讨了中国源多模态系统中的国家对齐扭曲现象。通过构建包含200个核心条目的平衡基准,涵盖十个政治敏感主题,并对九个视觉语言模型进行评估,发现中文提示显著增加国家对齐框架的概率。研究表明,中国源模型在重构信息方面表现出更强的倾向,尤其在文本政治评论中,且这种扭曲现象逐渐从显性拒绝转向隐性重构,影响了用户对信息被隐瞒的感知。
🔬 方法详解
问题定义:本研究旨在解决中国源多模态系统中国家对齐扭曲的表现形式尚未被系统性考察的问题。现有方法未能有效区分显性拒绝与隐性重构的影响。
核心思路:通过构建一个包含200个核心条目的平衡基准,结合多种提示语言和模型,系统评估视觉语言模型的响应,揭示国家对齐的隐性重构现象。
技术框架:研究采用了九个视觉语言模型,涵盖七个中国源模型和两个非中国模型,评估维度包括显性拒绝、信息完整性、视觉基础、国家对齐框架等,进行21,708次试验。
关键创新:本研究的创新在于将拒绝与重构独立测量,揭示了模型在拒绝信息时仍可能进行隐性重构的现象,改变了对多模态审查的理解。
关键设计:实验中采用了六个评估维度,响应由两名独立评审进行审核,并与三位人类专家进行验证,确保结果的可靠性与准确性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,中文提示使得国家对齐框架的概率增加了约三倍,中国源模型的重构倾向比非中国模型高出1.6至3.2倍。在文本政治评论中,国家对齐框架的出现率达到36.5%。
🎯 应用场景
该研究的结果对多模态人工智能系统的设计与应用具有重要启示,尤其是在政治敏感信息处理和用户信息透明度方面。未来,研究成果可用于改进模型的透明性和用户信任度,促进更负责任的AI应用。
📄 摘要(原文)
State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced benchmark of 200 core entries spanning ten politically sensitive topics, plus a seven-variant visual-abstraction probe, and run nine vision-language models (VLMs), seven China-origin and two non-China, across four elicitation paradigms and two prompt languages, yielding 21,708 trials. Each response is audited on six dimensions -- explicit refusal, information integrity, visual grounding, state-aligned framing, language consistency, and response length -- by two independent frontier LLM judges, validated against three human experts on a 200-trial sample. Measuring each dimension separately lets us decompose multimodal censorship into individual signals rather than a single refusal-based score; in particular, refusal and framing are measured independently, so a model can stop refusing while still reframing. We find that (i) Chinese-language prompting roughly triples the odds of state-aligned framing, within every model; (ii) China-origin models reframe more than non-China models (direction robust across judges and human raters; magnitude 1.6--3.2x); (iii) the effect is strongest in text-only political commentary (36.5%) and is gated by recognition of the depicted subject rather than pixel detail, persisting even at silhouette for iconic images; and (iv) across four Qwen generations, state-aligned framing rises while explicit refusal falls: censorship migrates from a visible act (refusal) to an invisible one (fluent reframing). We argue this shift to invisible reframing is fundamentally a problem of human-AI interaction: it removes the very signal users rely on to recognize that information has been withheld.