cs.CV(2026-08-14)
📊 共 19 篇论文 | 🔗 5 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (8 🔗1)
支柱二:RL算法与架构 (RL & Architecture) (5 🔗2)
支柱三:空间感知与语义 (Perception & Semantics) (3 🔗1)
支柱一:机器人控制 (Robot Control) (2 🔗1)
支柱七:动作重定向 (Motion Retargeting) (1)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (8 篇)
🔬 支柱二:RL算法与架构 (RL & Architecture) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 9 | ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models | 提出ForgeWM以解决交互式视频生成中的因果训练问题 | world model world models distillation | ||
| 10 | Self-Supervised Visual On-Policy Distillation | 提出自监督视觉在线蒸馏方法以解决信息不对称问题 | policy learning teacher-student distillation | ✅ | |
| 11 | CSG-Mamba: A Convolutional Scoring Gating Vision State Space Network for Endoscopic Polyp Segmentation | 提出CSG-Mamba以解决内窥镜息肉分割中的低对比度问题 | Mamba SSM state space model | ||
| 12 | ProFocus: Interpreting Affective Experience in Artistic Images with Progressive Visual Focusing | 提出ProFocus以解决艺术图像情感解读问题 | representation learning large language model multimodal | ✅ | |
| 13 | Marionette: Predicting World States, Rendering Geometry, Painting Appearance | 提出Marionette以解决长时间序列生成中的一致性问题 | world model world models penetration |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 14 | HiCo-GS: Hierarchical Context Aggregation and Geometric Consistency for Octree Gaussian Splatting | 提出HiCo-GS以解决现有Octree Gaussian Splatting的特征隔离问题 | gaussian splatting splatting geometric consistency | ✅ | |
| 15 | PISA: A Pseudo-Individual Source-Domain Feature Adaptation Framework for Test-Time Open-Vocabulary Object Detection | 提出PISA框架以解决开放词汇目标检测中的性能下降问题 | open-vocabulary open vocabulary | ||
| 16 | SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation | 提出SPARGen以统一空间感知与推理问题 | 3D reconstruction multimodal |
🔬 支柱一:机器人控制 (Robot Control) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 17 | OpenBelief-Nav: Evidence-Preserving Object Memory for Open-Vocabulary Language-Guided Navigation | 提出OpenBelief-Nav以解决开放词汇语言引导导航中的对象记忆问题 | Unitree open-vocabulary open vocabulary | ✅ | |
| 18 | Zero-Shot Skeleton-Based Action Anticipation | 提出零样本骨架动作预测方法以解决新动作识别问题 | humanoid |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 19 | Beyond Control Points: Arcsecond Relative-Motion Estimation of Vision Measurement Platforms With Incomplete or Absent Control Fields | 提出控制自适应差分框架以解决视觉测量平台运动估计问题 | motion estimation |