| 1 |
Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry |
系统评估离线强化学习在卒中抗血栓治疗中的应用 |
reinforcement learning offline RL offline reinforcement learning |
|
|
| 2 |
Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry |
提出信用可寻址推理以解决多模态几何推理问题 |
reinforcement learning multimodal |
|
|
| 3 |
Language-Informed Flow Matching for Trend-Guided Structure-Based 3D Molecular Generation |
提出LiFT框架以解决结构基础的3D分子生成问题 |
flow matching foundation model |
✅ |
|
| 4 |
PRACTICE: From Experience to Expertise in Self-Evolving Embodied Agents |
提出PRACTICE以解决自我进化体代理的技能更新问题 |
distillation large language model multimodal |
✅ |
|
| 5 |
PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization |
提出PLC-DPO以解决偏好优化中的标签噪声问题 |
preference learning DPO direct preference optimization |
|
|
| 6 |
Tail-Replay: Escaping the Curse of Linear Attention in Prefix Caching for Hybrid LLMs |
提出Tail-Replay以解决混合大语言模型中的前缀缓存问题 |
linear attention large language model |
|
|
| 7 |
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement |
提出On-Policy自适应方法以解决教师监督噪声问题 |
reinforcement learning distillation |
|
|
| 8 |
Event-Driven Language Models with Sparse Neural Activity for Neuromorphic Hardware |
提出稀疏神经活动模型以优化神经形态硬件上的语言模型推理 |
SSM linear attention large language model |
|
|
| 9 |
Online Estimation of Dynamic Origin-Destination Matrices Using Reinforcement Learning with Link-Flow Propagation Guidance |
提出LFPG-RL以解决动态OD矩阵在线估计问题 |
reinforcement learning PPO |
|
|
| 10 |
BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning |
提出BCPPO以解决高成本事件的安全强化学习问题 |
reinforcement learning PPO |
|
|
| 11 |
Balancing Privacy, Utility, and Safety in LLM Alignment through Preference Optimization |
提出P3M方法以平衡LLM对隐私、效用与安全的优化 |
DPO direct preference optimization large language model |
|
|
| 12 |
A Human-in-the-Loop Autonomous Agent for Industry Time Series Forecasting |
提出CastClaw以解决工业时间序列预测中的人机协作问题 |
MAE large language model |
|
|
| 13 |
One Policy Is Enough: Single-Agent Reinforcement Learning Outperforms Tree Search for Chemistry Tool Learning |
提出单一策略以超越树搜索解决化学工具学习问题 |
reinforcement learning |
|
|
| 14 |
PAC: Progress-Augmented Advantage Curriculum for Multi-Task Reinforcement Learning of LLMs |
提出PAC以优化多任务强化学习中的任务分配问题 |
reinforcement learning |
|
|
| 15 |
DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving |
提出DASC以解决混合线性注意力模型的状态压缩问题 |
linear attention |
|
|
| 16 |
Reinforcement Learning for Symbolic Equation Solving |
提出强化学习方法以逐步解决符号方程问题 |
reinforcement learning |
|
|
| 17 |
Locally-Guided Actor-Critic: Training a Goal-conditioned Actor with a Subgoal-aware Critic |
提出局部引导演员-评论家以解决目标条件强化学习中的稀疏奖励问题 |
reinforcement learning reward shaping |
|
|