cs.LG(2026-08-04)

📊 共 19 篇论文 | 🔗 3 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (11 🔗3) 支柱九:具身大模型 (Embodied Foundation Models) (7) 支柱八:物理动画 (Physics-based Animation) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (11 篇)

#题目一句话要点标签🔗
1 Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging 提出Any-OPD以解决异构模型间的在线蒸馏问题 flow matching distillation
2 Muon Meets Mamba: Spectral Optimization for State Space Models 提出Muon优化器以提升状态空间模型的训练效率 Mamba state space model
3 Agentic Reinforcement Learning with Self-Distilled Reward Shaping 提出自蒸馏奖励塑形方法以提升代理强化学习性能 reinforcement learning reward shaping
4 SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation 提出SMOPD以解决多奖励强化学习中的信号平衡问题 reinforcement learning distillation
5 PLAN: Parallel Liquid-Inspired Approximation Network for Efficient Representation Learning in Flexible Job Shop Scheduling 提出PLAN以解决灵活作业车间调度中的效率问题 reinforcement learning deep reinforcement learning DRL
6 Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL 提出CSDG以控制离线强化学习中的局部修正传播问题 reinforcement learning offline RL offline reinforcement learning
7 Bi-semantic Chemical Embedder for Joint Representation Learning of SMILES and Natural Language 提出CheMatE以解决化学领域语义表示不足问题 representation learning contrastive learning
8 CausalOPD: First-Wrong-Step Supervision for Distilling Causal Chain Reasoning 提出CausalOPD以解决因果链推理中的错误传播问题 reinforcement learning distillation large language model
9 Robust General Utility for Reinforcement Learning 提出鲁棒一般效用强化学习以解决效用误指定问题 reinforcement learning
10 AS-FedBridge: Pseudo-Spike Bridge Distillation for Heterogeneous ANN-SNN Federated Learning 提出AS-FedBridge以解决混合ANN-SNN联邦学习中的表示不一致问题 distillation
11 Latent Reward Registers for Diffusion Preference Alignment 提出潜在奖励寄存器以解决扩散模型偏好对齐问题 reinforcement learning distillation

🔬 支柱九:具身大模型 (Embodied Foundation Models) (7 篇)

#题目一句话要点标签🔗
12 The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics 利用链式思维动态检测大型语言模型推理失败 large language model chain-of-thought
13 Noise-Aware Shrinkage for Differentially Private Zeroth-Order Fine-Tuning of Large Language Models 提出SAGE方法以解决差分隐私下大语言模型微调中的噪声问题 large language model
14 Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility 提出测试时间缩放方法以提升推理LLMs的性能 large language model
15 Omega-S: A Functional Resilience Index for LLM Fine-Tuning 提出Omega-S以解决大语言模型微调中的能力退化问题 large language model
16 Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces 评估推理接口以缩短推理时间和提高回答准确性 large language model
17 LLM-Derived Priors for Thompson Sampling in Cold-Start Comment Recommendation 提出基于LLM的先验知识以解决冷启动评论推荐问题 large language model
18 A Graph Signal Processing Perspective on Numerical Sequence Representations in LLM In-Context Learning 提出图信号处理视角以解析LLM中的数字序列表示 large language model

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
19 LAEF: A Lead-Agnostic ECG Foundation Model Towards Point-of-Care Diagnostics 提出LAEF以解决心电图基础模型在点-of-care诊断中的局限性 spatiotemporal foundation model

⬅️ 返回 cs.LG 首页 · 🏠 返回主页