FAR-DPO: Feasibility-Aware and Robust Direct Preference Optimization for Cyclic Peptide Design

📄 arXiv: 2608.19808v1 📥 PDF

作者: Guofeng Zhang, Rong Han, Xiaoyu Wang, Zhiyun Li, Zongbo Han, Xiaohong Liu, Guangyu Wang

分类: cs.LG

发布日期: 2026-08-20

备注: 16 pages, 6 figures, and 7 tables. Includes supplementary materials


💡 一句话要点

提出FAR-DPO以解决循环肽设计中的可行性与鲁棒性问题

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)

关键词: 循环肽设计 生成模型 药物发现 多目标优化 生物物理可行性 鲁棒性优化 机器学习 生物医药工程

📋 核心要点

  1. 现有循环肽设计方法在生成模型扩展方面存在挑战,尤其是由于环化带来的几何和生物物理限制,导致可行设计空间受限。
  2. FAR-DPO通过可行性意识的偏好构建和难度意识的群体鲁棒优化,引导生成模型生成结构和生物物理上可行的循环肽设计。
  3. 在CPSea LNR基准上,FAR-DPO在固定生成预算下,成功率从46.89%提升至57.79%,显示出显著的性能提升。

📝 摘要(中文)

循环肽因其高结合亲和力和结构稳定性而在药物发现中备受关注。然而,将生成模型从线性肽扩展到循环肽设计面临挑战,主要由于环化限制了可行设计空间。现有方法依赖于零-shot生成或后处理过滤,导致可行设计的产出率低,且对多目标权衡的控制有限。为此,本文提出了FAR-DPO(可行性意识和鲁棒性直接偏好优化),该框架能够引导生成模型朝向结构和生物物理上可行的循环肽设计,特别是针对困难目标。FAR-DPO通过可行性意识的偏好构建和难度意识的群体鲁棒优化来实现这一目标。实验结果表明,FAR-DPO在CPSea LNR基准上显著提高了成功率,证明了其在可行性和目标鲁棒性方面的有效性。

🔬 方法详解

问题定义:本文旨在解决循环肽设计中的可行性和鲁棒性问题。现有方法多依赖于零-shot生成或后处理,导致可行设计的产出率低,且对多目标权衡的控制能力不足。

核心思路:FAR-DPO的核心思路在于通过构建可行性意识的偏好和难度意识的群体鲁棒优化,引导生成模型生成更具生物物理可行性的循环肽设计。这样的设计能够有效应对复杂的设计目标。

技术框架:FAR-DPO的整体架构包括两个主要模块:可行性意识的偏好构建和难度意识的群体鲁棒优化。前者通过可行性门控的多目标优势构建目标偏好对,后者则根据当前的偏好损失自适应地重新加权预定义的难度组。

关键创新:FAR-DPO的最重要创新在于其可行性意识的偏好构建和难度意识的群体鲁棒优化,这与现有方法的单一偏好优化或简单后处理形成鲜明对比。

关键设计:在具体实现中,FAR-DPO采用了多目标损失函数,并通过动态调整难度组的权重来优化生成过程,从而提高了设计的成功率和目标鲁棒性。实验中使用的基准数据集CPSea LNR为验证提供了良好的基础。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

FAR-DPO在CPSea LNR基准上表现出色,成功率在PepGLAD上从46.89%提升至57.79%,在PepFlow上从47.96%提升至49.57%。这些提升不仅体现在整体成功率上,还扩展到最困难目标的四分之一,且伴随更有利的最佳结合评分。

🎯 应用场景

该研究的潜在应用领域包括药物发现和生物医药工程,尤其是在需要设计高亲和力和高稳定性的循环肽时。FAR-DPO的框架能够为新药研发提供更为有效的设计工具,提升药物的开发效率和成功率,具有重要的实际价值和未来影响。

📄 摘要(原文)

Cyclic peptides are emerging as promising molecular scaffolds in drug discovery due to their high binding affinity and structural stability. However, extending generative models from linear to cyclic peptide design remains challenging, as cyclization sharply restricts the feasible design space through coupled geometric and biophysical constraints. Moreover, limited training data has led existing approaches to rely largely on zero-shot generation or post hoc filtering, resulting in low yields of feasible designs and limited control over multi-objective trade-offs. To address these limitations, we propose FAR-DPO (Feasibility-Aware and Robust Direct Preference Optimization), an architecture-agnostic framework that steers generative models toward structurally and biophysically feasible cyclic peptide designs, particularly for challenging targets. FAR-DPO integrates feasibility-aware preference construction with difficulty-aware group-robust optimization. Specifically, it constructs within-target preference pairs through feasibility-gated multi-objective dominance and adaptively reweights predefined difficulty groups according to their current preference losses. On the CPSea LNR benchmark, under a fixed generation budget, FAR-DPO increases overall success rate from 46.89% to 57.79% on PepGLAD and from 47.96% to 49.57% on PepFlow. These gains also extend to the hardest target quartile and are accompanied by more favorable best-per-target binding scores. Together, these results demonstrate FAR-DPO's effectiveness in improving feasibility and target-wise robustness.