| 1 |
Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs |
提出多教师自蒸馏策略优化以解决多领域LLM集成问题 |
reinforcement learning distillation large language model |
✅ |
|
| 2 |
DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation |
提出文档介导的强化学习以优化广告推荐中的技能 |
reinforcement learning representation learning large language model |
|
|
| 3 |
Cliff: Learning Process Rewards from the First Mistake |
提出Cliff以解决强化学习中奖励指导不足的问题 |
reinforcement learning distillation reward shaping |
|
|
| 4 |
Eliciting ESG Preferences for Reinforcement Learning-Based Portfolio Optimization |
提出多目标强化学习框架以优化ESG投资组合 |
reinforcement learning large language model |
|
|
| 5 |
Online Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Coupling |
通过分布式模型-智能体耦合提出在线强化学习以改进天气预报 |
reinforcement learning MAE |
|
|
| 6 |
Post-Training Language Models for Gold-Medal Performance in Coding Competitions |
提出后训练语言模型以在编程竞赛中实现金牌表现 |
reinforcement learning large language model |
|
|
| 7 |
Rethinking the Teacher-Student Framework for Test-Time Adaptation |
提出不更新教师模型以解决测试时适应中的错误累积问题 |
teacher-student |
✅ |
|
| 8 |
AGI Maze Prediction Datasets: A Compact Benchmark for Learning World Dynamics with Transformers |
提出AGI迷宫预测数据集以研究Transformer的世界动态学习 |
world model world models predictive model |
|
|
| 9 |
Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents |
提出SPACE以解决长时间交互任务中的动作选择效率问题 |
reinforcement learning large language model |
|
|
| 10 |
Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation |
提出低秩克隆蒸馏以解决MLP可达性差距问题 |
distillation |
|
|
| 11 |
A Comparative Study of Graph Representations for GNN-Based Power Grid Control in L2RPN |
比较不同图表示以优化电网控制的深度强化学习 |
reinforcement learning deep reinforcement learning |
|
|