| 19 |
Faster-WAM: Do World Action Models Need Deep Action Modules? |
提出Faster-WAM以解决现有世界动作模型的计算开销问题 |
world model world models world action model |
|
|
| 20 |
Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation |
提出指令条件探索以解决大语言模型的探索问题 |
reinforcement learning distillation large language model |
|
|
| 21 |
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning |
提出PCSD以解决强化学习中的稀疏奖励问题 |
reinforcement learning distillation large language model |
|
|
| 22 |
Antares: Foundation Models for Agentic Vulnerability Localization |
提出Antares以解决软件安全中的漏洞定位问题 |
reinforcement learning foundation model |
|
|
| 23 |
ProWorld: Progress-Aware Hyperbolic World Models for Long-Horizon Visual Goal Reaching |
提出ProWorld以解决长时间视觉目标规划中的进展问题 |
world model world models JEPA |
|
|
| 24 |
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models |
提出SpeechAgent-R以解决复杂音频推理问题 |
reinforcement learning multimodal |
|
|
| 25 |
Chess on Ice: Curling Tactical Decision-Making via Backward Induction and Deep Reinforcement Learning |
提出深度强化学习框架以解决冰壶战术决策问题 |
reinforcement learning deep reinforcement learning |
|
|
| 26 |
CoNav-UAV: Cooperative Dual-Altitude Aerial Navigation via Stackelberg Learning |
提出CoNav-UAV以解决无人机协同导航问题 |
distillation privileged information VLN |
|
|
| 27 |
Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling |
提出层次记忆Mamba以解决长序列建模中的表示瓶颈问题 |
Mamba linear attention |
|
|
| 28 |
DAPD: Dual-Anchored Policy Distillation |
提出双锚政策蒸馏以解决信息不对称问题 |
distillation privileged information |
|
|
| 29 |
Is More Privileged Information Better? From Solution Traces to Problem-Solving Structure in Self-Distilled Reasoning |
提出PS-OPSD以提升自蒸馏推理的准确性 |
distillation privileged information |
|
|
| 30 |
RL-Lock: Reinforcement Learning for Generating Interlocking Assemblies |
提出RL-Lock以解决生成互锁装配问题 |
reinforcement learning |
|
|
| 31 |
Agentic Incident Response through Digital Twin-Enhanced Multiscale Planning |
提出基于数字双胞胎的多尺度规划以优化事件响应 |
reinforcement learning large language model |
|
|
| 32 |
Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories |
提出Harness-R1以实现可执行运行时的智能编辑 |
reinforcement learning large language model |
|
|
| 33 |
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning |
提出多时刻策略优化方法以提升大语言模型推理能力 |
reinforcement learning large language model |
|
|
| 34 |
CoEvoKG: Co-Evolving Knowledge Graphs with Self-Evolving Search Agents |
提出CoEvoKG框架以解决自我演化搜索代理知识积累不足问题 |
reinforcement learning large language model |
✅ |
|
| 35 |
Syntax Meets Semantics: Understanding Scientific Formulae |
提出跨模态对齐方法以提升科学公式检索性能 |
representation learning contrastive learning |
|
|