cs.CL(2026-09-01)
📊 共 32 篇论文 | 🔗 3 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (19 🔗3)
支柱二:RL算法与架构 (RL & Architecture) (12)
支柱三:空间感知与语义 (Perception & Semantics) (1)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (19 篇)
🔬 支柱二:RL算法与架构 (RL & Architecture) (12 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 20 | Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict | 研究多模态大语言模型在知识冲突下的鲁棒性问题 | direct preference optimization large language model multimodal | ||
| 21 | CaRL-EM: Cost-Aware Reinforcement Learning for Entity Matching with LLMs | 提出CaRL-EM以解决实体匹配中的成本意识问题 | reinforcement learning large language model zero-shot transfer | ||
| 22 | Instella-MoE Technical Report | 提出Instella-MoE以提升大规模语言模型训练效率 | reinforcement learning direct preference optimization distillation | ||
| 23 | PersuaRL: Reinforcement Learning-Driven Multi-Expert Selection for Persuasive Dialogue Generation in Insurance | 提出PersuaRL以解决保险领域对话生成的说服力不足问题 | reinforcement learning large language model | ||
| 24 | VerTox: Verifiable Reward-Guided Corpus Poisoning Against Neural Ranking Models | 提出VerTox框架以解决神经排名模型的语料中毒问题 | reinforcement learning reward shaping large language model | ||
| 25 | SFAD: Speculative Factuality-Aware Decoding | 提出SFAD框架以解决大语言模型的上下文可信性问题 | reinforcement learning direct preference optimization large language model | ||
| 26 | Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO | 提出审计GRPO、SFT和DPO以增强语言模型的上下文理解能力 | DPO | ||
| 27 | The Rise of Verbal Reinforcement Learning | 提出语言强化学习以提升语言代理的反馈能力 | reinforcement learning | ||
| 28 | Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall | 提出Switch Distillation以解决中期训练中的知识蒸馏问题 | distillation | ||
| 29 | From Rollouts to Recipes: Self-Contained Post-Training for LLMs | 提出自适应后训练框架以优化大语言模型的样本处理 | distillation large language model | ||
| 30 | OUTLETS: Output-Length Prediction from Speculative Decoding Backbones | 提出OUTLETS以解决大语言模型输出长度预测问题 | MAE large language model | ||
| 31 | A Dataset for Modeling Iterative Problem-Solving | 构建CodeInsight数据集以建模迭代问题解决过程 | state space model TAMP |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 32 | Inspicio: Open-Vocabulary, LLM-Based Sense Retrieval for Historical Languages | 提出Inspicio以解决历史语言的词义消歧问题 | open-vocabulary open vocabulary |