cs.LG(2026-09-02)

📊 共 20 篇论文 | 🔗 4 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (11 🔗2) 支柱九:具身大模型 (Embodied Foundation Models) (6 🔗1) 支柱三:空间感知与语义 (Perception & Semantics) (2) 支柱五:交互与反应 (Interaction & Reaction) (1 🔗1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (11 篇)

#题目一句话要点标签🔗
1 Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs 提出多教师自蒸馏策略优化以解决多领域LLM集成问题 reinforcement learning distillation large language model
2 DMRL: Document-Mediated Reinforcement Learning for Skill Optimization in Advertising Recommendation 提出文档介导的强化学习以优化广告推荐中的技能 reinforcement learning representation learning large language model
3 Cliff: Learning Process Rewards from the First Mistake 提出Cliff以解决强化学习中奖励指导不足的问题 reinforcement learning distillation reward shaping
4 Eliciting ESG Preferences for Reinforcement Learning-Based Portfolio Optimization 提出多目标强化学习框架以优化ESG投资组合 reinforcement learning large language model
5 Online Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Coupling 通过分布式模型-智能体耦合提出在线强化学习以改进天气预报 reinforcement learning MAE
6 Post-Training Language Models for Gold-Medal Performance in Coding Competitions 提出后训练语言模型以在编程竞赛中实现金牌表现 reinforcement learning large language model
7 Rethinking the Teacher-Student Framework for Test-Time Adaptation 提出不更新教师模型以解决测试时适应中的错误累积问题 teacher-student
8 AGI Maze Prediction Datasets: A Compact Benchmark for Learning World Dynamics with Transformers 提出AGI迷宫预测数据集以研究Transformer的世界动态学习 world model world models predictive model
9 Act More, Decide Less: Skill-Guided Adaptive Action Chunking for Long-Horizon LLM Agents 提出SPACE以解决长时间交互任务中的动作选择效率问题 reinforcement learning large language model
10 Train What You Deploy: Closing the MLP Reachability Gap in Low-Rank Clone Distillation 提出低秩克隆蒸馏以解决MLP可达性差距问题 distillation
11 A Comparative Study of Graph Representations for GNN-Based Power Grid Control in L2RPN 比较不同图表示以优化电网控制的深度强化学习 reinforcement learning deep reinforcement learning

🔬 支柱九:具身大模型 (Embodied Foundation Models) (6 篇)

#题目一句话要点标签🔗
12 TC-Next: Zero-Shot Multimodal Cyclone Forecasting 提出TC-Next以实现零样本多模态热带气旋预报 foundation model multimodal
13 Do Tabular Foundation Models Know Physics? Contamination, Units, and the Deterministic Limit 探讨表格基础模型在物理知识学习中的局限性 foundation model
14 The Implications of Linguistic Illegibility for LLM Security 提出语言不清晰性以解决LLM安全性问题 chain-of-thought
15 WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading 提出WeaveMark以解决大语言模型水印提取准确性与文本质量的矛盾问题 large language model
16 GenCAR: Generative Counterfactual Alignment with Risk-Controlled Selection for Out-of-Distribution Recommendation 提出GenCAR以解决分布转移下的推荐风险控制问题 large language model
17 Network-Aware Forecasting on Wireless Access Points 提出网络感知预测方法以优化无线接入点的机器学习部署 foundation model

🔬 支柱三:空间感知与语义 (Perception & Semantics) (2 篇)

#题目一句话要点标签🔗
18 A Common Measure of Communication for Speech Brain-Computer Interfaces 提出开放词汇互信息以解决语音脑机接口的测量问题 open-vocabulary open vocabulary
19 What Is Worth Representing? Representational Empowerment for Continual Model Construction 提出代表性赋能方法以解决持续模型构建问题 open-vocabulary open vocabulary

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
20 Compositional Spectral Prompts for LLM-based Online Time Series Forecasting 提出CoSPOT以解决在线时间序列预测中的长期适应问题 CHOIS

⬅️ 返回 cs.LG 首页 · 🏠 返回主页