cs.LG(2026-08-11)

📊 共 25 篇论文 | 🔗 3 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (11 🔗1) 支柱九:具身大模型 (Embodied Foundation Models) (8 🔗2) 支柱八:物理动画 (Physics-based Animation) (2) 支柱一:机器人控制 (Robot Control) (2) 支柱六:视频提取与匹配 (Video Extraction) (1) 支柱四:生成式动作 (Generative Motion) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (11 篇)

#题目一句话要点标签🔗
1 Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving 提出Dreamer-SAC框架以解决自主驾驶中的样本效率问题 reinforcement learning policy learning PPO
2 Efficient Hypergradient Descent for Inverse Reinforcement Learning 提出高效超梯度下降法以解决逆强化学习中的计算挑战 reinforcement learning inverse reinforcement learning
3 Scheduling Mixed RL Rollouts Beyond Prefix Locality 提出MISA-T以解决异构RL回滚调度问题 reinforcement learning RLHF large language model
4 ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation 提出ReOrder-OPD以解决在政策蒸馏中的教师监督不可靠问题 teacher-student distillation
5 IADD-TR: Intervention-Aware Dynamics Decoupling with Targeted Regularization for Model-Based Reinforcement Learning 提出IADD-TR以解决模型基强化学习中的数据偏差问题 reinforcement learning policy learning
6 Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation 提出探索驱动的个性化联邦强化学习框架以解决隐私和探索问题 reinforcement learning distillation
7 SQuaT: Self-Supervised Knowledge Distillation via Student-Aware Quantized Teacher Features 提出SQuaT以解决量化模型蒸馏中的标签缺失问题 distillation
8 Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning 提出无评论预训练以解决在线强化学习微调效率问题 reinforcement learning
9 TideRL: Boosting Agentic RL Goodput with Readiness-Aware Scheduling 提出TideRL以提升多轮强化学习的训练效率 reinforcement learning large language model
10 Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks 提出SINKFLEX-RL以解决长时间工具使用代理任务的挑战 reinforcement learning
11 Partially Observable Learning for Multi-Platform Dispatch Optimization 提出POLO框架以解决多平台调度优化中的部分可观测性问题 reinforcement learning reward shaping

🔬 支柱九:具身大模型 (Embodied Foundation Models) (8 篇)

#题目一句话要点标签🔗
12 ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes 提出ProTAGAD以解决TAG异常检测中的模糊边界问题 large language model foundation model
13 Mapping and Measuring the Behavioral Evolution of Large Language Models 提出行为映射与测量方法以分析大语言模型的演变 large language model
14 Can Bayesian Optimization Efficiently Find a Strong Single Expert in Neural Thickets? 提出贝叶斯优化以高效寻找强单一专家模型 large language model
15 Diffract: Spectral View of LLM Domain Adaptation 提出Diffract工具以优化大语言模型领域适应性 large language model
16 TACTICL: Task-Aware Compression of Tabular ICL Models 提出TACTICL以解决表格任务模型压缩问题 foundation model
17 Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control 提出行为模式轴以控制大型语言模型的行为风格 large language model
18 ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions 提出ProbGuard以解决LLM输出安全风险评估问题 large language model
19 Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy 提出LLM引导的控制平面策略以解决HAProxy后端故障隔离问题 large language model

🔬 支柱八:物理动画 (Physics-based Animation) (2 篇)

#题目一句话要点标签🔗
20 DEFT: Data-Efficient Frequency-domain Top-k Sampling via Inverse Discrete Fourier Transform for Spatiotemporal Dynamical Systems Modeling 提出DEFT以解决时空动力系统建模中的数据效率问题 spatiotemporal
21 Link-adaptive digital twin for robust physical-layer modeling in hybrid-amplified ultra-wideband optical networks 提出链接自适应数字双胞胎以解决超宽带光网络建模问题 ASE

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
22 Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique 提出Latent Critic以解决大语言模型的幻觉检测问题 manipulation large language model
23 Automatic Field-of-View Adjustment for a View-Expansive Microscope via LSTM-Based Gaze and Pipette Motion Interpretation 提出基于LSTM的自动视野调整方法以提升显微操作效率 manipulation

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
24 Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives 提出统一框架以解决跨视角特征匹配问题 feature matching foundation model

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
25 Physics-informed Diffusion Generative Model for Time-Series Data Synthesis in Dynamic Systems 提出PhysDGM以解决动态系统时间序列数据合成问题 physics-informed diffusion

⬅️ 返回 cs.LG 首页 · 🏠 返回主页