cs.AI(2026-08-19)

📊 共 19 篇论文

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (9) 支柱二:RL算法与架构 (RL & Architecture) (7) 支柱一:机器人控制 (Robot Control) (2) 支柱八:物理动画 (Physics-based Animation) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (9 篇)

#题目一句话要点标签🔗
1 Preference Reasoning under Indeterminacy in Large Language Models 提出应对不确定性偏好的推理方法以提升语言模型决策能力 large language model
2 Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models 利用自提示与跨模型共识实现科学文献数据提取的可重复性 large language model
3 rEDMRec: Distilling Large Language Model Reasoning into an Editable Experience Memory for Recommendation 提出rEDMRec以解决推荐系统中的推理效率问题 large language model
4 DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning 提出DentAgent以解决多模态牙科推理中的证据整合问题 multimodal
5 Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference 提出BudgetDoc以优化文档推理中的推理预算分配问题 multimodal
6 What is Missing from AI Post-Training AI: An Empirical Analysis 提出后训练AI的策略重评机制以提升执行能力 large language model
7 Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication 提出可验证潜在对齐框架以监测隐秘多智能体通信 VLA
8 Science Done on a Machine by a Machine: AI Agents in Computational Chemistry 提出自主AI系统以推动计算化学研究的自动化 generalist agent
9 LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents 提出LEDGER以解决LLM代理输出审计问题 large language model

🔬 支柱二:RL算法与架构 (RL & Architecture) (7 篇)

#题目一句话要点标签🔗
10 UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval 提出UMER框架以解决多模态检索中的语义推理问题 distillation multimodal chain-of-thought
11 Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models 提出EvoResearcher以解决大语言模型推理效率问题 reinforcement learning large language model
12 Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson 提出基于成对排序的离线解释选择方法以降低LLM成本 PPO DPO teacher-student
13 AlphaClifford: Efficient Clifford Synthesis and Transpilation with Model-based RL 提出AlphaClifford以解决高门数Clifford电路合成问题 reinforcement learning model-based RL
14 Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services 提出Flama框架以简化API和机器学习服务的开发与部署 predictive model large language model
15 Eureka: Task-Conditioned Meta-Agent Orchestration for Scientific Discovery 提出Eureka以解决科学发现中的任务调度问题 Eureka
16 SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution 提出SkillForge框架以主动解决项目特定问题 distillation large language model

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
17 RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training 提出RTPO以解决多回合强化学习训练不稳定问题 trajectory optimization reinforcement learning large language model
18 Breaking the weakest link to evade vision language models 提出针对视觉语言模型的对抗攻击方法以提升安全性 manipulation multimodal

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
19 ORBITER: Conflict-Aware Decision-Making for Agentic Last-Mile Delivery 提出ORBITER以解决最后一公里配送中的决策冲突问题 spatiotemporal

⬅️ 返回 cs.AI 首页 · 🏠 返回主页