Challenges in annotations by humans and LLMs: A case study of evaluative language
作者: Mirela Imamovic, Aenne Cecilia Kristine Knierim, Khushi Pitroda, Ekaterina Lapshinova-Koltunski
分类: cs.CL, cs.SI
发布日期: 2026-07-30
💡 一句话要点
比较人类与LLMs在复杂语言注释中的表现
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 评估性语言 大型语言模型 注释一致性 数字人文学科 TED演讲 主观性注释 Appraisal理论
📋 核心要点
- 现有方法在处理复杂语言现象时,尤其是主观性强的注释任务上存在一致性不足的问题。
- 论文通过分析TED演讲稿中的评估性语言,提出了三种提示语以优化LLMs的自动分类性能。
- 实验结果显示,LLMs在注释一致性上优于训练中的语言学家,F1-score达0.77,表明其在复杂注释任务中的有效性。
📝 摘要(中文)
本文比较了在培训中的语言学家、训练有素的语言学家与大型语言模型(LLMs)在复杂语言现象注释中的表现。研究以TED演讲稿为例,分析了评估性语言,重点关注评估理论及其情感、判断和欣赏三个子类。通过对人类注释的评估和LLMs的自动分类性能比较,发现LLMs在复杂注释任务中表现优于培训中的语言学家,表明LLMs在数字人文学科的复杂理论注释和分析中具有潜在的应用价值。
🔬 方法详解
问题定义:本文旨在解决人类和LLMs在复杂语言注释中的一致性问题,尤其是在评估性语言的注释上,现有方法在主观性强的任务中表现不佳。
核心思路:通过分析TED演讲稿中的评估性语言,比较不同注释者的表现,探索LLMs在复杂注释任务中的潜力。设计三种提示语以优化模型性能,旨在提高注释的一致性和准确性。
技术框架:研究分为几个主要阶段:首先,进行人类注释的句子级评估;其次,开发三种提示语并进行模型性能比较;最后,使用最佳提示语对LLMs进行微调。
关键创新:最重要的创新在于通过对比不同注释者的表现,发现LLMs在复杂注释任务中能够超越训练中的语言学家,提供了一种新的注释方法。
关键设计:在实验中,使用了特定的损失函数和参数设置,以确保模型在评估性语言的分类上达到最佳性能,最终实现了F1-score为0.77的结果。
🖼️ 关键图片
📊 实验亮点
实验结果显示,LLMs在评估性语言的注释一致性上表现优于训练中的语言学家,F1-score达0.77,表明其在复杂注释任务中的有效性和潜力,开启了数字人文学科的新路径。
🎯 应用场景
该研究的潜在应用领域包括数字人文学科、语言学研究和教育技术等。通过利用LLMs进行复杂语言注释,可以提高研究效率和准确性,推动相关领域的发展,尤其是在需要处理大量文本数据的场景中。
📄 摘要(原文)
In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.