cs.CV(2026-08-05)
📊 共 50 篇论文 | 🔗 10 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (22 🔗4)
支柱三:空间感知与语义 (Perception & Semantics) (12 🔗2)
支柱二:RL算法与架构 (RL & Architecture) (9 🔗2)
支柱一:机器人控制 (Robot Control) (5 🔗2)
支柱六:视频提取与匹配 (Video Extraction) (2)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (22 篇)
🔬 支柱三:空间感知与语义 (Perception & Semantics) (12 篇)
🔬 支柱二:RL算法与架构 (RL & Architecture) (9 篇)
🔬 支柱一:机器人控制 (Robot Control) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 44 | MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight | 提出MobileWAM以解决移动操控中的动态协调问题 | whole-body control locomotion manipulation | ||
| 45 | CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention | 提出CofactVLA以解决视觉-语言-动作模型中的因果混淆问题 | manipulation flow matching vision-language-action | ||
| 46 | Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models | 提出Faster-WAM以解决高效推理时间未来条件化问题 | manipulation world action model world action models | ||
| 47 | Differential 6-DOF Pose Estimation with Provable First-Order Immunity to Camera Calibration Errors | 提出一种差分6自由度姿态估计方法以解决相机标定误差问题 | manipulation motion estimation | ✅ | |
| 48 | Towards a satellite image manipulation and deepfake localization benchmark dataset | 构建卫星图像操控与深伪检测基准数据集以解决验证真实性问题 | manipulation | ✅ |
🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 49 | The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering | 提出EgoCross基准以解决跨领域第一人称视频问答问题 | egocentric large language model multimodal | ||
| 50 | Promptable Animal Pose Tracking Across Species | 提出基于视觉基础模型的动物姿态跟踪方法以解决标注数据不足问题 | feature matching foundation model |