| 1 |
No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models |
提出对比逆动态方法以解决JEPA世界模型中的崩溃问题 |
world model worldmodel world models |
✅ |
|
| 2 |
Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics |
提出基于相对迁移的动态课程采样方法以优化大语言模型训练 |
curriculum learning large language model |
|
|
| 3 |
MoRAX: Mobility-based Representation Augmentation for Geospatial Foundation Models |
提出MoRAX以增强地理基础模型的城市表示能力 |
representation learning foundation model |
|
|
| 4 |
Repetition as Reinforcement: Enhancing Sample Efficiency via Instant Episode Repetition in Reinforcement Learning |
提出即时情节重复机制以提升强化学习样本效率 |
reinforcement learning SAC TD3 |
|
|
| 5 |
Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents |
提出基于LLM反馈的政策不变奖励塑形框架以优化混合RL代理 |
reinforcement learning reward shaping large language model |
|
|
| 6 |
An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models |
提出采样验证危险法则以解决连续控制中的模式遗漏问题 |
world model world models |
|
|
| 7 |
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL |
提出Co-RL框架以解决无监督推理中的反馈循环问题 |
reinforcement learning multimodal |
✅ |
|
| 8 |
Composing Flow-Matching Energies with Known Physics: Generation, OOD Detection, and Inversion on PDE Fields |
提出流匹配能量模型以解决PDE场的生成与OOD检测问题 |
flow matching |
|
|
| 9 |
Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation |
提出图结构在线难度估计以优化RLVR调度 |
reinforcement learning large language model |
|
|
| 10 |
rl-triton: High-Performance Triton GPU Kernels for Reinforcement Learning Credit Assignment |
提出rl-triton以解决强化学习信用分配问题 |
reinforcement learning |
✅ |
|
| 11 |
Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents |
提出EvalXRL基准以评估可解释强化学习方法的有效性 |
reinforcement learning large language model |
|
|
| 12 |
Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning |
提出NSPER以提升图像基础强化学习的样本效率 |
reinforcement learning |
|
|