cs.RO(2026-08-17)

📊 共 25 篇论文 | 🔗 4 篇有代码

🎯 兴趣领域导航

支柱一:机器人控制 (Robot Control) (17 🔗3) 支柱二:RL算法与架构 (RL & Architecture) (3 🔗1) 支柱九:具身大模型 (Embodied Foundation Models) (2) 支柱六:视频提取与匹配 (Video Extraction) (2) 支柱三:空间感知与语义 (Perception & Semantics) (1)

🔬 支柱一:机器人控制 (Robot Control) (17 篇)

#题目一句话要点标签🔗
1 HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL 提出HAF框架以解决人形机器人全身运动协调问题 humanoid humanoid robot locomotion
2 $τ_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation 提出$τ_0$-VLA以解决长时间机器人操作中的决策问题 manipulation world model world models
3 When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents 提出状态语义注入方法以解决LLM驱动的实体代理安全问题 manipulation affordance vision-language-action
4 US-VLA: An Ultrasound Vision-Language-Action Model for Embodied Abdomina 提出US-VLA模型以解决超声扫描自动化问题 manipulation reinforcement learning vision-language-action
5 NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation 提出NebulaVLA以解决机器人操作中的效率与性能权衡问题 manipulation cross-embodiment vision-language-action
6 Throwing a Tight Spiral American Football by a Humanoid Robot 提出人形机器人控制美式足球投掷以解决投掷精度问题 humanoid humanoid robot whole-body control
7 HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction 提出HiPHI数据集以解决人类运动与物体交互的高精度学习问题 humanoid policy learning human motion
8 ViHaTeleop: A Low-Cost, Lightweight Visual-Haptic Teleoperation System for Dexterous Manipulation Learning 提出ViHaTeleop以解决低成本遥操作中的高质量示范收集问题 manipulation dexterous manipulation teleoperation
9 SparkVLA: Stop-Aware Hierarchical VLA with Adaptive Action Chunking for Long-Horizon Manipulation 提出SparkVLA以解决长时间操作中的停止与执行决策问题 manipulation vision-language-action VLA
10 Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory 提出BATON以解决长时间机器人操作中的任务链问题 manipulation vision-language-action VLA
11 MatchingPolicy: Correspondence-Aware Policy Enables Cross-Object In-Context Learning 提出MatchingPolicy以解决跨对象场景学习中的性能下降问题 manipulation policy learning imitation learning
12 RoboStriker: Latent-Space Strategic Games for Autonomous Humanoid Boxing 提出RoboStriker以解决人形机器人拳击中的策略与执行问题 humanoid humanoid robot reinforcement learning
13 FlexWorm: Primitive-augmented Hybrid Contact-motion Planning for Suction-based Multi-segment Deformable Robots 提出FlexWorm以解决多段软体机器人运动规划问题 motion planning
14 Robot-Body-Aware Traversal Risk Graph Planning for Wheeled-Legged Robots in Complex Terrain 提出机器人身体感知的行驶风险图规划以解决复杂地形导航问题 legged robot
15 Semantic- and Density-Aware Planning for Accessibility-Preserving Multi-Object Placement 提出语义与密度感知规划以解决多物体放置问题 manipulation motion planning
16 Unified Condition-Action Modeling for Accurate One-Step Action Generation 提出UCA-Flow以解决机器人动作生成中的条件与动作联合建模问题 manipulation representation learning
17 Observation-Constrained Joint-Space Viewpoint Optimization for Robotic Inspection of Cylindrical Cavities 提出观察约束的关节空间视角优化方法以解决机器人圆柱腔体检查问题 Unitree

🔬 支柱二:RL算法与架构 (RL & Architecture) (3 篇)

#题目一句话要点标签🔗
18 Orbit-Planner: Towards Latent World Models for On-Orbit Obstacle Avoidance of Satellite Agents 提出Orbit-Planner以解决卫星在轨障碍规避问题 world model world models
19 SurgVIL: Scaling Surgical Robot Imitation Learning with Open-source Surgical Videos 提出SurgVIL框架以解决外科机器人模仿学习数据稀缺问题 policy learning imitation learning
20 Planner-Conditioned Diffusion for Coordinated Multi-Agent Exploration 提出规划条件扩散策略以解决多智能体协调探索问题 diffusion policy multimodal

🔬 支柱九:具身大模型 (Embodied Foundation Models) (2 篇)

#题目一句话要点标签🔗
21 Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation 提出信任边界中心的安全框架以应对基础模型驱动的实体代理安全问题 foundation model multimodal
22 Co-design of Neural and Muscle Network based on Embodied Perceptron Representation 提出一种新框架以优化机器人控制与身体设计 embodied AI

🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)

#题目一句话要点标签🔗
23 Neurosymbolic Embodied Agents 提出神经符号体代理以解决可执行计划生成问题 egocentric visual grounding
24 Exposing the Long-tail in Embodied Urban Navigation via Scalable Learning from In-the-Wild Videos 提出可扩展框架以解决城市导航中的长尾问题 egocentric vision-language-action

🔬 支柱三:空间感知与语义 (Perception & Semantics) (1 篇)

#题目一句话要点标签🔗
25 OccamView: Object-Conditioned View Selection for Frame-Budgeted Active 3D Gaussian Reconstruction 提出OccamView以解决有限帧预算下的3D高斯重建问题 3DGS open-vocabulary open vocabulary

⬅️ 返回 cs.RO 首页 · 🏠 返回主页