The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era, 2021-2025
作者: Przemysław Czuma
分类: cs.DL, cs.AI, cs.CL, cs.CY
发布日期: 2026-08-06
备注: Preregistered study (OSF: 10.17605/OSF.IO/U5NEY); deviations from the registered plan, including a formal validation-gate breach, are disclosed in Section 4.6. Companion study: arXiv:2606.29540. 3 figures, 4 tables
💡 一句话要点
分析美国国会新闻稿中无间隔破折号频率的变化
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 大型语言模型 文本风格 国会新闻稿 破折号频率 数据分析 文本生成 风格痕迹
📋 核心要点
- 现有的新闻稿写作风格与大型语言模型生成的文本存在显著差异,尤其是在破折号的使用上。
- 本研究通过分析大量国会新闻稿,探讨无间隔破折号的频率变化,以评估LLMs对文本风格的影响。
- 分析结果显示,2025年无间隔破折号的使用频率显著上升,表明LLMs辅助写作的影响逐渐显现。
📝 摘要(中文)
大型语言模型(LLMs)在文本中留下小的风格痕迹,尤其是无间隔破折号(U+2014)。本研究探讨这一痕迹在美国国会新闻稿中的可测量性。通过分析2021至2025年间的146,239份新闻稿,发现无间隔破折号的密度在2025年显著上升,超过了四年基线的两倍。研究结果表明,LLMs辅助写作的广泛传播可能与模型的成熟有关,但未能满足所有预注册决策规则,因而不支持因果关系的声明。
🔬 方法详解
问题定义:本研究旨在探讨大型语言模型在美国国会新闻稿中留下的风格痕迹,特别是无间隔破折号的使用频率。现有方法未能有效量化这种风格变化的影响。
核心思路:通过分析146,239份国会新闻稿,研究无间隔破折号的密度变化,以评估LLMs对文本风格的影响。设计了预注册的研究框架以确保结果的可靠性。
技术框架:研究采用了Poisson/负二项模型,结合文本长度偏移进行分析,按办公室进行聚类。数据来源于开放的国会新闻稿数据集,涵盖480个众议院和参议院办公室。
关键创新:本研究首次系统性地量化了LLMs对国会新闻稿风格的影响,发现无间隔破折号的使用频率在2025年显著上升,且这一变化在不同办公室和党派间均表现出一致性。
关键设计:研究中使用的主要参数包括每千字符的无间隔破折号密度,采用了预注册的设计以确保结果的有效性,并进行了多项虚假检验以验证结果的稳健性。实验结果显示,75.6%的持续办公室在无间隔破折号的使用上有所增加。
🖼️ 关键图片
📊 实验亮点
研究发现,2025年无间隔破折号的使用密度达到0.217,较2021-2024年的基线水平翻倍,且相关新闻稿中使用无间隔破折号的比例从约13%上升至24.8%。这一变化在不同办公室和党派间表现出一致性,表明LLMs的影响逐渐显现。
🎯 应用场景
该研究为理解大型语言模型在文本生成中的影响提供了重要视角,尤其是在新闻写作领域。未来,相关发现可用于指导新闻机构在使用LLMs时的风格调整,提升文本质量与一致性。
📄 摘要(原文)
Large language models (LLMs) can leave small stylistic traces in text written with their help. The most discussed is the em-dash (U+2014), especially the unspaced form word---word, which is normal in typeset English prose but unusual in U.S. press writing, where AP style calls for spaced dashes. This study asks whether that trace is measurable in congressional press releases. In a preregistered design (OSF: 10.17605/OSF.IO/U5NEY), 146,239 scraper-sourced releases from 480 House and Senate offices (2021-2025, the open congress-press dataset) were analyzed: density of unspaced prose-form em-dashes per 1,000 characters of cleaned text, Poisson/negative-binomial models with a length offset, clustering by office. Density stayed within 0.10-0.12 per 1,000 characters through 2021-2024, then rose to 0.217 in 2025, more than twice the four-year baseline; the share of releases with such an em-dash rose from ~13% to 24.8%. The primary frequency ratio (2023-2025 vs 2021-2022) was 1.55 (95% CI 1.28-1.93; exact registered cut-off: 1.528), just above the prespecified 1.5x threshold. The rise was net-new (hyphen density stable), held within authors (75.6% of 262 continuous offices increased; p ~ 1e-16) and in a closed panel of 224 offices, and survived falsification tests: three placebo cut-offs were null, the pipeline showed no step at the 2024/2025 boundary, and continuing offices carried the rise. A segmented regression finds no step at the ChatGPT cut-off but a clear post-period acceleration; the 2025 rise is symmetric across parties and chambers. Because the registered validation gate was formally breached, the full preregistered decision rule was not met; the interpretation (broad diffusion of LLM-assisted writing as the models matured) is offered as exploratory. The em-dash remains a population-level marker, not a per-release authorship detector, and the design supports no causal claim.