cs.LG(2026-08-03)

📊 共 28 篇论文 | 🔗 4 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (15 🔗1) 支柱九:具身大模型 (Embodied Foundation Models) (9 🔗2) 支柱一:机器人控制 (Robot Control) (2) 支柱八:物理动画 (Physics-based Animation) (1 🔗1) 支柱四:生成式动作 (Generative Motion) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (15 篇)

#题目一句话要点标签🔗
1 WorldDynCache: Risk-Controlled Latent Dynamics Approximation for Diffusion World Model 提出WorldDynCache以解决扩散世界模型推理速度慢的问题 world model world models latent dynamics
2 Diffusion Policy with Behavioral Advantage Correction for Offline Reinforcement Learning 提出行为优势校正的扩散策略以解决离线强化学习中的Q值偏差问题 reinforcement learning offline reinforcement learning diffusion policy
3 Understanding and Correcting Low-Frequency Bias in EEG Foundation Model 提出FAME框架以解决EEG基础模型中的低频偏差问题 masked autoencoder foundation model
4 Start Classifying: Categorical Critics for LLM Reinforcement Learning 提出HL-Gauss PPO以优化大语言模型的强化学习评估 reinforcement learning PPO large language model
5 Learning-Based Collaborative MEC for LLM Inference with Soft-Deadline Awareness via Transformer-Enhanced PPO 提出基于Transformer增强的PPO以解决MEC服务器的LLM推理问题 PPO large language model
6 Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning 提出CoKL以解决LLM强化学习中的能力保留问题 reinforcement learning large language model
7 Beyond On-Policy Exploration: Integrating External Policy Rollouts for Reinforcement Learning in Diffusion Language Models 提出ERILS以解决扩展策略回放在强化学习中的挑战 reinforcement learning large language model
8 DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling 提出DART以解决长序列建模中的注意力效率问题 Mamba SSM state space model
9 HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning 提出HindSearch以解决搜索增强强化学习中的失败轨迹信息损失问题 reinforcement learning distillation
10 Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection 提出PRECOG以解决检索增强生成模型的预填充成本问题 SSM
11 Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning 提出有限时间分析以优化折扣指数效用强化学习 reinforcement learning
12 LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation 提出LEAP以解决低级代码生成中的反馈稀疏问题 reinforcement learning large language model
13 Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints 提出HeLyMARL以解决无线网络中的资源管理问题 reinforcement learning
14 Progressive Agent Skill Generation via Reinforcement Learning 提出Skill-α以解决技能生成中的监督信号缺失问题 reinforcement learning
15 Analytic Planning under Uncertainty with Moment Closure 提出基于时刻闭合的分析规划方法以应对不确定性问题 reinforcement learning deep reinforcement learning

🔬 支柱九:具身大模型 (Embodied Foundation Models) (9 篇)

#题目一句话要点标签🔗
16 Why Large Language Models Fail at Tabular Prediction 探讨大型语言模型在表格预测中的失败原因 large language model foundation model
17 CENTILE: A Telemetry Foundation Model Evaluated by the Decisions It Drives 提出CENTILE以优化网络和系统遥测决策 foundation model TAMP
18 RamanPFN: learning from Raman spectral structure with a tabular foundation model 提出RamanPFN以解决拉曼光谱数据分析中的结构依赖问题 foundation model
19 How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models 提出可观察性阶梯以评估大型语言模型的推理透明度 large language model
20 GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning 提出GradCuit以解决测试时潜在推理的鲁棒性与可解释性问题 large language model chain-of-thought
21 HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses 提出HarnessCompass以解决自动化工具演化中的过拟合问题 large language model
22 ReFP-AD: Rectified Flow Preconditioning for Energy-Based Anomaly Detection 提出ReFP-AD以解决高维异常检测中的不稳定性问题 foundation model
23 Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal standard 提出F-ICL基准以评估语言模型的算法推理能力 large language model
24 Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning 提出FedGD以解决大语言模型个性化奖励建模中的偏好异质性问题 large language model

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
25 Foundations of Reinforcement Learning and Control:Connections and New Perspectives 提出自适应控制与演员-评论家算法结合以优化动态系统控制 locomotion reinforcement learning
26 Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning 提出期望值n步Q学习以解决离线强化学习中的偏差问题 manipulation reinforcement learning

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
27 LieStoNet: Learning Lie Symmetries from Spatiotemporal Data for Stochastic Dynamical Systems 提出LieStoNet以从时空数据中学习随机动力系统的李对称性 spatiotemporal

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
28 Probabilistic Deep Learning for Drought Forecasting: Role of Internal Climate Variability 提出基于深度学习的干旱预测框架以应对气候内部变率问题 physically plausible

⬅️ 返回 cs.LG 首页 · 🏠 返回主页