cs.LG(2026-08-31)

📊 共 33 篇论文 | 🔗 7 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (17 🔗2) 支柱九:具身大模型 (Embodied Foundation Models) (14 🔗5) 支柱八:物理动画 (Physics-based Animation) (1) 支柱一:机器人控制 (Robot Control) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (17 篇)

#题目一句话要点标签🔗
1 Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry 系统评估离线强化学习在卒中抗血栓治疗中的应用 reinforcement learning offline RL offline reinforcement learning
2 Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry 提出信用可寻址推理以解决多模态几何推理问题 reinforcement learning multimodal
3 Language-Informed Flow Matching for Trend-Guided Structure-Based 3D Molecular Generation 提出LiFT框架以解决结构基础的3D分子生成问题 flow matching foundation model
4 PRACTICE: From Experience to Expertise in Self-Evolving Embodied Agents 提出PRACTICE以解决自我进化体代理的技能更新问题 distillation large language model multimodal
5 PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization 提出PLC-DPO以解决偏好优化中的标签噪声问题 preference learning DPO direct preference optimization
6 Tail-Replay: Escaping the Curse of Linear Attention in Prefix Caching for Hybrid LLMs 提出Tail-Replay以解决混合大语言模型中的前缀缓存问题 linear attention large language model
7 Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement 提出On-Policy自适应方法以解决教师监督噪声问题 reinforcement learning distillation
8 Event-Driven Language Models with Sparse Neural Activity for Neuromorphic Hardware 提出稀疏神经活动模型以优化神经形态硬件上的语言模型推理 SSM linear attention large language model
9 Online Estimation of Dynamic Origin-Destination Matrices Using Reinforcement Learning with Link-Flow Propagation Guidance 提出LFPG-RL以解决动态OD矩阵在线估计问题 reinforcement learning PPO
10 BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning 提出BCPPO以解决高成本事件的安全强化学习问题 reinforcement learning PPO
11 Balancing Privacy, Utility, and Safety in LLM Alignment through Preference Optimization 提出P3M方法以平衡LLM对隐私、效用与安全的优化 DPO direct preference optimization large language model
12 A Human-in-the-Loop Autonomous Agent for Industry Time Series Forecasting 提出CastClaw以解决工业时间序列预测中的人机协作问题 MAE large language model
13 One Policy Is Enough: Single-Agent Reinforcement Learning Outperforms Tree Search for Chemistry Tool Learning 提出单一策略以超越树搜索解决化学工具学习问题 reinforcement learning
14 PAC: Progress-Augmented Advantage Curriculum for Multi-Task Reinforcement Learning of LLMs 提出PAC以优化多任务强化学习中的任务分配问题 reinforcement learning
15 DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving 提出DASC以解决混合线性注意力模型的状态压缩问题 linear attention
16 Reinforcement Learning for Symbolic Equation Solving 提出强化学习方法以逐步解决符号方程问题 reinforcement learning
17 Locally-Guided Actor-Critic: Training a Goal-conditioned Actor with a Subgoal-aware Critic 提出局部引导演员-评论家以解决目标条件强化学习中的稀疏奖励问题 reinforcement learning reward shaping

🔬 支柱九:具身大模型 (Embodied Foundation Models) (14 篇)

#题目一句话要点标签🔗
18 TSPFN: A Temporal Tabular Foundation Model for Physiological Time Series Classification 提出TSPFN以解决生理时间序列分类中的数据不足问题 foundation model
19 TrainSDC: Characterizing and Mitigating Silent Data Corruption in Large Language Model Training 提出TrainSDC以解决大语言模型训练中的静默数据损坏问题 large language model
20 Foundation Models Meet Agriculture: Challenges Beyond Pretraining 提出针对农业的基础模型以解决预训练后的应用挑战 foundation model
21 Uncertainty of Vision Medical Foundation Models 提出基于领域特定模型的医疗视觉不确定性估计方法 foundation model
22 A Model with No Head and Many Thoughts 提出Soft Latent Thinking以提升语言模型推理效率 large language model chain-of-thought
23 A Universal Context-Reuse Layer for Cross-Model KV Sharing 提出跨模型KV共享层以解决冗余计算问题 large language model
24 Fine-Tuning Low-Bit Models with Gradient in Quantized Code Space 提出代码替代梯度以优化低比特模型的微调 instruction following
25 E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation 提出E-Commerce Bench以评估长时间自主商业操作中的LLM代理 large language model
26 The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal 提出机制可解释性分析以解决角色扮演越狱中的安全识别问题 large language model
27 Q-Strata: Hierarchical Bit Allocation for Mixed-Precision Quantization of Mixture-of-Experts LLMs 提出Q-Strata以优化混合精度量化中的位宽分配问题 large language model
28 Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability 提出张量方法以优化大语言模型的训练与推理 large language model
29 When the Martingale Never Stops Firing: Anytime-Valid Gating on Real Forecast Streams 提出实时监控的任何时刻有效推断方法以解决数据依赖性问题 foundation model
30 CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration 提出CateKV以解决长上下文LLM推理加速问题 large language model
31 Graph4BiLO: Graph Neural Network Approximation for Bilevel Mixed-Integer Linear Optimization 提出Graph4BiLO以解决双层混合整数线性优化问题 zero-shot transfer

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
32 Coarse composition suffices: tabular in-context learning for multi-activity antimicrobial peptide profiling 提出基于表格的上下文学习方法以解决多标签抗菌肽活性预测问题 AMP foundation model multimodal

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
33 T3S: Improving Multi-Task Reinforcement Learning with Task-Specific Feature Selector and Scheduler 提出T3S框架以解决多任务强化学习中的干扰问题 manipulation reinforcement learning

⬅️ 返回 cs.LG 首页 · 🏠 返回主页