cs.LG(2026-07-30)

📊 共 38 篇论文 | 🔗 2 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (20 🔗1) 支柱九:具身大模型 (Embodied Foundation Models) (14 🔗1) 支柱一:机器人控制 (Robot Control) (2) 支柱四:生成式动作 (Generative Motion) (1) 支柱八:物理动画 (Physics-based Animation) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (20 篇)

#题目一句话要点标签🔗
1 Contrastive Reinforced Policy Optimization via Privileged Self-Distillation 提出对比强化策略优化以解决自蒸馏中的偏见问题 reinforcement learning contrastive learning distillation
2 QQWorld: Quantile-Quantile Matching for World Model Regularization 提出QQWorld以解决潜在世界模型的尾部样本问题 world model worldmodel world models
3 FedOGL: Combating Catastrophic Forgetting in Federated Open-World Multimodal Graph Learning 提出FedOGL以解决联邦开放世界多模态图学习中的灾难性遗忘问题 distillation multimodal
4 Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective 提出统一理论框架以理解子次模信息度量在表示学习中的应用 representation learning contrastive learning multimodal
5 Cybersecurity Detection Classification with Reasoning-enabled Language Models 提出基于推理的语言模型以解决网络安全检测分类问题 reinforcement learning large language model chain-of-thought
6 Flux-OPD: On-Policy Distillation with Evolving Contexts 提出Flux-OPD以解决开放领域任务偏好蒸馏问题 distillation large language model
7 TAPO: Transition-Aware Policy Optimization for LLM Agents 提出TAPO以优化大型语言模型代理的策略学习 reinforcement learning large language model foundation model
8 S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring 提出S-CEReBrO以解决连续EEG监测中的内存瓶颈问题 linear attention spatiotemporal foundation model
9 $β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation 提出$β$-OPSD以提升推理语言模型的稳定性与性能 reinforcement learning distillation
10 LM-GRASP: Instance-Specific Language Models for Combinatorial Construction via Online Imitation Learning 提出LM-GRASP以解决组合优化中的实例特定学习问题 reinforcement learning imitation learning
11 Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning 提出奖励设计框架以提升强化学习去学习效率 reinforcement learning reward design
12 Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold 提出扩展与压缩框架以提升推理解决方案的有效性 reinforcement learning distillation instruction following
13 ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow 提出ODEWorld以解决离散时间预测的低效问题 policy learning world model world models
14 Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular Histories 提出双重稳健功能表示学习以解决不规则历史的因果推断问题 representation learning
15 On-Policy and Off-Policy Learning for Large Action Spaces 提出结构化贝叶斯方法以解决大动作空间中的策略学习问题 policy learning
16 Multi-channel Uplift Policy Learning 提出ReAlloc框架以解决多渠道营销预算分配问题 policy learning
17 Class-Aware Reinforcement Learning for Counterfactual Explanation Generation 提出类感知强化学习以生成反事实解释 reinforcement learning
18 Improving the Robustness/Accuracy Tradeoff Against Adversarial Attacks Using Information Bottleneck Distillation Through Dual Teachers 通过双教师信息瓶颈蒸馏提升对抗攻击的鲁棒性与准确性 distillation
19 Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning 提出KGPS以解决RL微调中的动态提示选择问题 reinforcement learning large language model
20 Policy Gradient Steering: Interventions from Behavioral Objectives 提出政策梯度引导以解决现有行为干预方法不足问题 reinforcement learning large language model

🔬 支柱九:具身大模型 (Embodied Foundation Models) (14 篇)

#题目一句话要点标签🔗
21 Revisiting Predictive Process Monitoring in the Age of Foundation Models: A Comparative Study of Sequence, Tabular, and LLM Approaches 比较序列、表格和大语言模型在预测过程监控中的应用 large language model foundation model
22 LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger 提出LedgerMind以解决多模态推理中的证据追溯问题 multimodal
23 Building a User Foundation Model for the Open Web 提出用户基础模型以解决开放网络中的身份碎片化问题 foundation model
24 What Makes Graph Unified? Principles and Generative Sliding-Window Transformer for Graph Foundation Models 提出SliGFM以解决跨域图特征统一问题 foundation model
25 LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference 提出LightRot以解决低比特大语言模型推理的能效与准确性问题 large language model
26 Memory Efficient Tabular Foundation Models 提出内存高效的表格基础模型以解决部署问题 foundation model
27 Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees 提出自适应预期策略树以解决GUI代理决策延迟问题 multimodal
28 Uncertainty quantification for trustworthy deep learning: Methods and measures 提出深度学习不确定性量化方法以提升预测可信度 large language model
29 TriShield: Zero-Utility-Loss Defense Against Privacy Backdoors in Federated Language Model Fine-Tuning via Orthogonal Gradient Projection and Optimizer State Entanglement 提出TriShield以解决联邦语言模型微调中的隐私后门问题 large language model
30 Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting 提出贝叶斯领域加权方法以优化多领域数据混合 large language model
31 Train Small, Deploy Large: Zero-Shot GNN Transfer Through Geometric Renormalization 提出零-shot转移协议以解决GNN在大规模图上的训练问题 zero-shot transfer
32 Towards joint scaling laws with optimal batch size schedules 提出动态批量大小调度以优化深度学习训练效果 large language model
33 Back from the Future: Key-Value Cache Management by Counter-Causal Surprise 提出基于反因果惊讶的KV缓存管理方案以优化LLM性能 large language model
34 Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs 提出Prox框架以解决LLM中FFN激活稀疏性问题 large language model

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
35 Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control 提出基于TD3的强化学习微调方法以优化多区域VAV控制 model predictive control reinforcement learning TD3
36 Real-Time Hard Peak Age-of-Information Safety with No-Regret Learning 提出OCO-PAoI-Hard以解决IoT系统的实时安全问题 teleoperation reinforcement learning deep reinforcement learning

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
37 MUGEN: A Unified Framework for Efficient Motion Understanding and Generation 提出MUGEN框架以高效理解与生成运动 text-to-motion human motion motion representation

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
38 QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error Correction 提出QAdapt以解决量子错误纠正中的噪声适应性问题 spatiotemporal

⬅️ 返回 cs.LG 首页 · 🏠 返回主页