| 22 |
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning |
提出EnvACE以解决长时间工具使用的环境交互问题 |
reinforcement learning policy learning world model |
✅ |
|
| 23 |
DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model |
提出DreamGuard以解决LLM代理的风险管理问题 |
world model world models large language model |
|
|
| 24 |
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning |
提出AgentOPSD以解决长时间多回合任务中的信用分配问题 |
reinforcement learning teacher-student distillation |
|
|
| 25 |
AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents |
提出AppDeltaWorld以解决移动GUI代理的环境建模问题 |
reinforcement learning world model world models |
|
|
| 26 |
ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion |
提出ViSR-KGC以解决多模态知识图谱补全问题 |
representation learning multimodal |
|
|
| 27 |
When Agentic AI Meets Integrated Sensing and Communication |
提出AISAC框架以整合智能感知与通信技术 |
reinforcement learning world model world models |
|
|
| 28 |
GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models |
提出GAUGE基准以评估模拟引擎和视频世界模型的物理真实度 |
world model world models |
|
|
| 29 |
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models |
提出DASH以解决标准OPSD在时间结构利用上的不足 |
reinforcement learning distillation large language model |
✅ |
|
| 30 |
iARCS: Iterative Agentic RL for Controllable 3D Scene Generation |
提出iARCS框架以解决3D场景生成中的功能约束问题 |
reinforcement learning traversability embodied AI |
|
|
| 31 |
StepReflect: Structured UI Transition Reflection for Mobile GUI Agents |
提出StepReflect以解决移动GUI代理的准确动作反映问题 |
teacher-student distillation multimodal |
|
|
| 32 |
From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over Networks |
提出HDT-Nets框架以解决物理AI协调问题 |
world model world models spatiotemporal |
|
|
| 33 |
Training a Conditioned Video Game Agent on a VLM Annotated Dataset |
提出基于视觉语言模型的强化学习视频游戏代理训练方法 |
reinforcement learning policy learning offline RL |
|
|
| 34 |
Contextual Information Policy Optimization for Search Agents |
提出上下文信息策略优化以解决搜索代理的推理问题 |
reinforcement learning large language model |
|
|
| 35 |
Subliminal Learning is Non-Semantic Distillation |
提出隐性学习以解决AI系统可预测性问题 |
distillation |
|
|
| 36 |
VLMs for Videogame Data Annotation |
利用视觉语言模型进行视频游戏数据标注以提升训练效果 |
reinforcement learning offline reinforcement learning |
|
|
| 37 |
TaskSense: Focusing on What Matters in World Models |
提出TaskSense以解决视觉控制中的任务相关性问题 |
world model world models dreamer |
|
|
| 38 |
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models |
提出DASH以解决标准OPSD在时间结构利用上的不足 |
reinforcement learning distillation large language model |
✅ |
|
| 39 |
Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation |
提出基于LLM的SYF系统以解决推荐系统用户偏好表达不足问题 |
DPO direct preference optimization multimodal |
|
|
| 40 |
WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader |
提出WebGrader以解决大语言模型在网页开发中的功能缺口问题 |
reinforcement learning reward design large language model |
|
|
| 41 |
Contextual Information Policy Optimization for Search Agents |
提出上下文信息策略优化以解决搜索代理的推理问题 |
reinforcement learning large language model |
|
|