Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators

📄 arXiv: 2608.30110v1 📥 PDF

作者: Xinyue Zhao, Ruiyi Zhang, Liqin Ye, Rui Cao, Pengtao Xie, Sudheer Chava

分类: cs.CL, cs.AI

发布日期: 2026-08-31


💡 一句话要点

提出LiveMacroEval以实时评估大型语言模型的宏观经济指标预测能力

🎯 匹配领域: 支柱八:物理动画 (Physics-based Animation) 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 宏观经济预测 大型语言模型 实时评估 数据污染 经济指标 金融市场 货币政策

📋 核心要点

  1. 现有的宏观经济指标预测方法面临数据污染的挑战,尤其是历史数据的评估容易受到影响。
  2. 本文提出LiveMacroEval基准,允许LLM在官方发布前实时生成宏观经济指标的预测,避免了数据污染问题。
  3. 实验结果表明,经过六个月的测试,LLM的预测准确性与专业机构的基准相当,显示出其作为实时估计工具的潜力。

📝 摘要(中文)

本研究聚焦于宏观经济指标的实时预测,即在官方发布前估计指标的当前值,这对货币政策和金融市场至关重要。大型语言模型(LLM)因其广泛的知识和实时网络搜索能力,成为这一任务的有力候选者。为评估LLM的预测能力,本文提出了LiveMacroEval,一个实时且抗污染的基准,LLM在每次官方发布前的窗口内每小时生成16个主要美国宏观经济指标的预测。通过与联邦储备银行的预测、Bloomberg ECOS专业共识和auto-ARIMA基线进行比较,结果显示,LLM的整体预测准确性与这些机构和专业基准相当,尽管不同指标的表现差异较大,突显了LLM作为实时宏观经济条件估计工具的潜力。

🔬 方法详解

问题定义:本研究旨在解决大型语言模型在宏观经济指标预测中的评估问题,现有方法在历史数据评估中容易受到污染,影响结果的可靠性。

核心思路:提出LiveMacroEval基准,通过实时生成预测并在官方发布前进行评估,避免了数据污染的影响,从而准确评估LLM的预测能力。

技术框架:整体架构包括数据收集、LLM预测生成、实时评估和结果比较四个主要模块。LLM利用网络搜索获取最新信息,生成每小时的宏观经济指标预测。

关键创新:LiveMacroEval基准的提出是本研究的核心创新,它通过实时生成预测和抗污染的评估机制,显著提高了LLM在宏观经济预测中的应用潜力。

关键设计:在实验中,使用了四种最先进的LLM,结合网络搜索功能,评估指标包括LiveMacro Score和LiveBetting Score,确保了预测的准确性和可靠性。实验持续六个月,涵盖了多个宏观经济指标。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,经过六个月的测试,LLM的整体预测准确性与联邦储备银行的预测、Bloomberg ECOS专业共识和auto-ARIMA基线相当,表明LLM在宏观经济条件实时估计中的潜力。不同指标的表现差异显著,强调了LLM在特定经济指标预测中的适用性。

🎯 应用场景

该研究的潜在应用领域包括金融市场分析、货币政策制定和经济研究等。通过实时预测宏观经济指标,决策者和投资者能够更快地响应经济变化,从而提高决策的有效性和市场的稳定性。未来,LLM在经济预测中的应用可能会进一步拓展到其他领域,如政策分析和风险管理。

📄 摘要(原文)

Nowcasting headline macroeconomic indicators, i.e., estimating an indicator's value for the current reference period before its official release, is critical for monetary policy and financial markets, and central banks devote dedicated teams of expert economists to producing such estimates. Large language model (LLM) agents are a promising candidate for this task, combining broad world knowledge with real-time web search and supporting queries at higher frequency than institutional nowcasts. Evaluating their nowcasting capability is, however, challenging: headline indicators such as GDP and CPI are widely reported and likely memorized during pretraining, so any evaluation on historical releases is vulnerable to data contamination. To address this, we introduce LiveMacroEval, a live, contamination-resistant benchmark in which LLM agents produce hourly nowcasts for sixteen major U.S. macroeconomic indicators over a pre-release window closing at each official release. Nowcast quality is assessed through a LiveMacro Score against announcement-window equity returns and a LiveBetting Score from simulated Polymarket-style trading, with Federal Reserve regional-bank nowcasts, the Bloomberg ECOS professional consensus, and an auto-ARIMA baseline as comparators. Over six months with four state-of-the-art LLM agents configured with web search, aggregate nowcast accuracy is broadly comparable to the institutional and professional benchmarks, with performance varying widely across individual indicators. This highlights LLM agents' potential as real-time estimators of macroeconomic conditions.