Eliciting ESG Preferences for Reinforcement Learning-Based Portfolio Optimization

📄 arXiv: 2609.02677v1 📥 PDF

作者: Giovanni Dispoto, Marcello Restelli, Carmine Ventre

分类: q-fin.PM, cs.CE, cs.LG

发布日期: 2026-09-02


💡 一句话要点

提出多目标强化学习框架以优化ESG投资组合

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture) 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 投资组合优化 多目标强化学习 环境社会治理 高斯过程 偏好引导 金融科技 可持续投资

📋 核心要点

  1. 现有的强化学习方法在ESG投资组合优化中存在单一评级提供者的局限,无法反映行业内的多样性。
  2. 本文提出将ESG投资组合优化视为多目标强化学习问题,整合多个ESG机构的评级,并引入偏好引导框架。
  3. 通过模拟不同地区的投资组合经理,实证结果显示地区背景对偏好权重有显著影响,提供了灵活的适应性框架。

📝 摘要(中文)

现代投资组合管理日益要求在传统风险调整收益与严格的环境、社会和治理(ESG)标准之间取得平衡。现有的强化学习方法通常只优化单一ESG提供者,忽视了行业内评级方法的显著差异以及手动权衡冲突目标的复杂性。本文将ESG意识的投资组合优化形式化为多目标强化学习问题,同时整合来自三个不同ESG机构的评级。为弥合高维算法权衡与人类决策之间的差距,我们采用高斯过程的偏好引导框架,使从业者能够通过直观的成对比较推断其潜在效用函数。通过使用大型语言模型(LLM)角色模拟不同地区背景下的投资组合经理,我们系统评估了该框架。实证结果表明,地区背景显著影响偏好权重,例如,基于欧洲的角色更倾向于优先考虑ESG一致性,而德克萨斯州的角色则更关注风险调整后的表现。

🔬 方法详解

问题定义:本文旨在解决现有强化学习方法在ESG投资组合优化中的局限性,特别是单一评级提供者的使用导致的偏差和复杂的目标权衡问题。

核心思路:论文提出将ESG投资组合优化视为多目标强化学习问题,整合来自多个ESG评级机构的信息,并通过偏好引导框架帮助决策者更好地理解和选择投资组合。

技术框架:整体架构包括三个主要模块:1) 多目标强化学习算法,2) 高斯过程偏好引导框架,3) 投资组合评估与选择模块。该框架通过成对比较来推断用户的潜在效用函数。

关键创新:最重要的技术创新在于将多目标强化学习与高斯过程结合,允许用户通过直观的方式表达其对不同投资组合的偏好,从而克服了传统方法的局限性。

关键设计:在技术细节上,采用高斯过程回归来建模用户的偏好,并设计了适应性损失函数以优化多目标的权衡,确保算法能够有效处理高维数据。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,基于不同地区背景的投资组合经理在偏好权重上存在显著差异。例如,欧洲背景的角色更倾向于优先考虑ESG一致性,而德克萨斯州背景的角色则更关注风险调整后的表现。这一发现强调了地区文化对投资决策的影响。

🎯 应用场景

该研究的潜在应用领域包括金融投资、资产管理及可持续投资决策。通过提供一个灵活的框架,能够帮助投资者在遵循ESG标准的同时实现风险调整后的收益,具有重要的实际价值和未来影响。

📄 摘要(原文)

Modern portfolio management increasingly demands a balance between traditional risk-adjusted returns and strict Environmental, Social, and Governance (ESG) mandates. Current Reinforcement Learning (RL) approaches typically optimize for a single ESG provider, neglecting the significant divergence in rating methodologies across the industry and the unintuitive nature of manually weighting conflicting objectives. This paper addresses these limitations by formulating ESG-aware portfolio optimization as a Multi-Objective Reinforcement Learning (MORL) problem that simultaneously incorporates ratings from three distinct ESG agencies. To bridge the gap between high-dimensional algorithmic trade-offs and human decision-making, we integrate a Preference Elicitation framework using Gaussian Processes. This system enables practitioners to infer their latent utility functions through intuitive pairwise comparisons of candidate portfolios based on their Sharpe ratios and aggregate ESG scores. We systematically evaluate our framework by employing Large Language Model (LLM) personas to simulate Portfolio Managers operating under varied regional contexts. Empirical results using historical market data reveal that regional backgrounds fundamentally shift the derived preference weights. For instance, European-based personas tend to prioritize ESG alignment over financial returns, while Texas-based personas favor risk-adjusted performance. This work offers a highly adaptable framework that successfully aligns multi-objective algorithmic trading with diverse, real-world human sustainability preferences.