cs.LG(2026-08-10)

📊 共 21 篇论文 | 🔗 1 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (12 🔗1) 支柱九:具身大模型 (Embodied Foundation Models) (7) 支柱一:机器人控制 (Robot Control) (1) 支柱八:物理动画 (Physics-based Animation) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (12 篇)

#题目一句话要点标签🔗
1 Twin Rollouts: Noise-Coupled Counterfactual Branching in Interactive Video World Models 提出噪声耦合双重回放以解决交互视频世界模型的反事实生成问题 world model world models spatiotemporal
2 DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models 提出DreOPD以解决流匹配模型的优化冲突问题 reinforcement learning flow matching distillation
3 Multimodal Federated Learning under Dual-Axis Modality Missingness 提出Flux框架以解决双轴模态缺失问题的多模态联邦学习 representation learning multimodal
4 SR-OPSD: Self-Referenced On-Policy Self-Distillation 提出SR-OPSD以解决现有自蒸馏方法的不稳定性问题 reinforcement learning distillation large language model
5 Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation 提出SKALD框架以解决强化学习中的奖励信号不足问题 reinforcement learning distillation
6 Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation 提出OP²SD以探讨教师行为对自蒸馏的影响 distillation privileged information
7 MARA: Flow-Matching-Guided Multi-Agent Resource Allocation for Computational Resource Efficient Learning 提出MARA以解决多任务计算资源分配问题 reinforcement learning flow matching
8 Bayesian Symbolic Regression with Entropic Reinforcement Learning 提出ERRLESS以解决符号回归中的不确定性量化问题 reinforcement learning
9 Learning to Modulate, Not to Cycle: Soft Actor---Critic Recovers Inverter-Style Heat-Pump Control 提出基于SAC的控制策略以解决热泵压缩机循环问题 reinforcement learning PPO SAC
10 WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training 提出WDL-OPD以解决在线蒸馏的不稳定性问题 distillation
11 Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training 提出TrajVal以解决任务学习性评估不足问题 reinforcement learning large language model
12 Dynamic Distribution-Aware Uncertainty Tracking in Vision-Language Representation Learning 提出动态分布感知的不确定性跟踪以解决视觉语言模型的可靠性问题 representation learning

🔬 支柱九:具身大模型 (Embodied Foundation Models) (7 篇)

#题目一句话要点标签🔗
13 Deep Multimodal Wearable Sensor Fusion for Detection of Body-Focused Repetitive Behaviors 提出深度多模态传感器融合方法以检测身体专注重复行为 multimodal
14 Hyperbolic Multimodal Continual Learning 提出超曲面多模态持续学习框架以解决遗忘问题 multimodal
15 Rethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive Approach 提出FedAS-LoRA以优化联邦学习中的低秩适应问题 large language model
16 Test-Time Scaling for CAD Generation via Verifier-Free Consensus Selection 提出无验证者共识选择以优化CAD生成 large language model
17 Activation Probes Surface Code-Security Signals that the Model's Output Misses 提出激活探针以揭示模型输出遗漏的安全信号 chain-of-thought
18 SwiftQK: Fast and Communication-Efficient Tensor Parallelism for Query-Key Normalization 提出SwiftQK以解决QK-Norm在多GPU下的通信效率问题 large language model
19 Multitask Scanning Probe Microscopy 提出多任务扫描探针显微镜以优化材料表征效率 multimodal

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
20 Satellite Trajectory Optimization via Proximal Policy Optimization for Space Debris Avoidance 提出基于近端策略优化的卫星轨迹优化方法以避免太空碎片碰撞 trajectory optimization reinforcement learning PPO

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
21 F2STNet: Fair and Federated Spectral-Temporal Modeling for Graph Forecasting 提出F2STNet以解决图结构数据的时空预测问题 spatiotemporal

⬅️ 返回 cs.LG 首页 · 🏠 返回主页