Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation

📄 arXiv: 2608.10385v1 📥 PDF

作者: Samaneh Mohtadi, Pietro Bernardelle, Joel Mackenzie, Gianluca Demartini

分类: cs.IR, cs.AI

发布日期: 2026-08-11

备注: Accepted at CIKM 2026


💡 一句话要点

提出人格条件化作为LLM评估敏感性探测器

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 大型语言模型 信息检索 评估者敏感性 人格条件化 系统排名一致性 评估框架 神经排名系统

📋 核心要点

  1. 现有方法在使用大型语言模型进行信息检索评估时,缺乏对评估者框架影响的深入理解,导致判断可靠性问题。
  2. 论文提出通过人格条件化来探测LLM评估者的敏感性,利用不同的评估者角色来分析其对判断的影响。
  3. 实验结果表明,高容量模型在系统排名一致性上表现良好,而小模型则显示出更大的敏感性和不稳定性。

📝 摘要(中文)

大型语言模型(LLMs)在信息检索(IR)评估中越来越多地被用作相关性评估者,这引发了关于评估者框架如何影响判断可靠性和下游系统比较的问题。本文研究了人格条件化作为一种诊断机制,以揭示LLM评估者的敏感性。通过使用来自两个互补来源(PersonaHub和NVIDIA Nemotron-Personas-USA)的任务导向人格,我们实例化了五种评估者角色,强调意图解释、领域专业知识、对比判断、证据验证和全球搜索质量评估,并与标准的UMBRELA基线进行比较。分析结果显示,评估者的敏感性是结构化的而非统一的,判断通常保持接近基线,同时在评估严格性、证据阈值或解释重点上有所变化,而不是产生广泛的相关性逆转。高容量模型保持系统排名一致性,而较小模型则放大了人格引起的不稳定性。

🔬 方法详解

问题定义:本文旨在解决大型语言模型作为信息检索评估者时,评估者框架对判断可靠性的影响。现有方法未能充分探讨这一问题,导致评估结果的不一致性和不可靠性。

核心思路:通过人格条件化,论文提出了一种新的诊断机制,利用不同的评估者角色来揭示LLM的敏感性,从而更好地理解评估者的判断过程。

技术框架:研究采用了来自PersonaHub和NVIDIA Nemotron-Personas-USA的任务导向人格,构建了五种评估者角色,分别关注意图解释、领域专业知识、对比判断、证据验证和全球搜索质量评估,并与标准UMBRELA基线进行比较。

关键创新:最重要的创新在于将人格条件化作为一种控制敏感性探测工具,能够有效识别出对评估者框架敏感的系统,进而优化LLM在IR评估中的应用。

关键设计:在实验中,采用了六种LLM骨干网络,分析了在TREC DL20和RAG24数据集上的表现,重点考察了评估者角色、模型容量和系统类型对评估结果的影响。

🖼️ 关键图片

fig_0
fig_1

📊 实验亮点

实验结果显示,使用人格条件化的评估者角色能够有效揭示LLM的敏感性。高容量模型在系统排名一致性上表现良好,而较小模型则在评估中显示出更大的不稳定性,特别是在神经排名系统上,强调了人格条件化的重要性。

🎯 应用场景

该研究的潜在应用领域包括信息检索系统的评估与优化,尤其是在需要高可靠性和一致性的场景中。通过理解评估者的敏感性,研究可以帮助开发更为稳健的评估框架,提升信息检索系统的整体性能和用户体验。

📄 摘要(原文)

Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment reliability and downstream system comparison. We study persona conditioning as a diagnostic mechanism for exposing LLM assessor sensitivity. Using task-oriented personas drawn from two complementary sources (PersonaHub and NVIDIA Nemotron-Personas-USA), we instantiate five assessor roles emphasizing intent interpretation, domain expertise, contrastive judgment, evidence verification, and global search-quality assessment, compared with a standard UMBRELA baseline. Across six LLM backbones on TREC DL20 and RAG24, our analyses reveal structured rather than uniform assessor sensitivity. Judgments usually remain close to the baseline while shifting assessment strictness, evidential threshold, or interpretation emphasis rather than producing widespread relevance reversals. At the system level, high-capacity models preserve system-ranking agreement, while smaller models amplify persona-induced instability. Local rank-displacement analysis shows sensitivity concentrates on particular retrieval systems and system types, especially neural ranking/reranking systems on DL20 and RAG-oriented pipelines on RAG24. Persona source matters less than assessor role and model capacity. These findings position persona-conditioned judging as a controlled sensitivity probe for stress-testing LLM-based IR evaluation pipelines and identifying systems whose evaluation outcomes are sensitive to assessor framing.