Language Chain in Alignment: Cross-Lingual Ranking Preference Optimization

📄 arXiv: 2608.23149v1 📥 PDF

作者: Seungyoon Lee, Minhyuk Kim, Jungseob Lee, Heuiseok Lim

分类: cs.CL, cs.AI

发布日期: 2026-08-24

备注: EMNLP 2026 Main


💡 一句话要点

提出跨语言排名偏好优化框架以提升多语言模型性能

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 跨语言对齐 偏好优化 多语言模型 LambdaLoss 机器学习 自然语言处理 模型适应性

📋 核心要点

  1. 现有方法过于依赖英语偏好数据,导致其他语言的模型性能不佳,缺乏有效的跨语言对齐机制。
  2. CRPO框架通过设计层次结构,利用英语的偏好知识来优化目标语言的偏好,从而实现跨语言的偏好对齐。
  3. 实验结果显示,CRPO在五种语言的指令遵循和知识利用能力上均优于传统方法,且在不同加权方案下表现出显著的性能提升。

📝 摘要(中文)

大型语言模型的对齐通常依赖于以英语为中心的高质量偏好数据,这导致其他语言的性能不佳。本文提出了跨语言排名偏好优化(CRPO)框架,利用英语的偏好知识来促进目标语言的偏好对齐。通过设计并行偏好对的层次结构,CRPO能够联合优化语言内部和语言间的偏好,从而提升语言适应性和输出质量。实验结果表明,CRPO在五种不同语言的指令遵循和知识利用能力上均优于标准方法,验证了其在多语言环境中的有效性。

🔬 方法详解

问题定义:本研究旨在解决大型语言模型在多语言环境中对齐性能不足的问题,现有方法主要依赖英语偏好数据,导致其他语言的模型效果不理想。

核心思路:CRPO框架通过引入层次结构,利用英语的偏好知识来优化目标语言的偏好,旨在实现更有效的跨语言对齐。

技术框架:CRPO的整体架构包括并行偏好对的设计,分为语言内部和语言间的偏好优化模块,结合LambdaLoss框架进行多候选响应的相对排名信号优化。

关键创新:CRPO的主要创新在于超越了传统的二元比较优化,提供了跨多个候选响应的相对排名信号,从而增强了模型的适应性和输出质量。

关键设计:在损失函数设计上,CRPO采用LambdaLoss,结合层次结构的偏好对,优化了模型的学习过程,确保了在多语言设置下的有效性。具体的参数设置和网络结构细节在实验中进行了验证。

🖼️ 关键图片

fig_0
img_1
img_2

📊 实验亮点

实验结果表明,CRPO在五种语言的指令遵循和知识利用能力上均优于标准方法,尤其在不同加权方案下,性能提升幅度显著,验证了其在多语言环境中的有效性。

🎯 应用场景

该研究的潜在应用领域包括多语言对话系统、跨语言信息检索和多语言内容生成等。通过提升不同语言模型的对齐性能,CRPO框架能够为全球用户提供更高质量的语言服务,具有重要的实际价值和未来影响。

📄 摘要(原文)

The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal performance in other languages. In this paper, we propose Cross-Lingual Ranking Preference Optimization (CRPO), a novel framework that leverages robust preference knowledge from English to facilitate preference alignment in the target language. We design a hierarchical structure within parallel preference pairs across the target language and English to jointly optimize intra- and inter-lingual preferences, thereby enhancing language adaptation and output quality. Building on the LambdaLoss framework, CRPO goes beyond the binary comparison based optimization by providing a relative ranking signal across multiple candidate responses. Our experiments across five languages with varying resource scales demonstrate that CRPO consistently outperforms standard approaches in both instruction-following and knowledge utilization capability. Notably, the robust performance gains observed across various weighting schemes further validate the empirical effectiveness of our hierarchical design in a multilingual setup. Furthermore, our findings highlight that CRPO significantly improves both reward margins and the log-probability of desirable responses, contributing to a more stable preference manifold for cross-lingual alignment.