cs.CV(2026-08-20)
📊 共 31 篇论文 | 🔗 5 篇有代码
🎯 兴趣领域导航
支柱二:RL算法与架构 (RL & Architecture) (9 🔗2)
支柱九:具身大模型 (Embodied Foundation Models) (7 🔗1)
支柱一:机器人控制 (Robot Control) (5 🔗1)
支柱三:空间感知与语义 (Perception & Semantics) (5 🔗1)
支柱四:生成式动作 (Generative Motion) (3)
支柱六:视频提取与匹配 (Video Extraction) (1)
支柱七:动作重定向 (Motion Retargeting) (1)
🔬 支柱二:RL算法与架构 (RL & Architecture) (9 篇)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (7 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 10 | Robust Cross-Modal Foundation Model Perception for Underwater Robots under Degraded Visual Conditions | 提出跨模态基础模型以提升水下机器人在恶劣视觉条件下的感知能力 | foundation model multimodal | ||
| 11 | Question-Guided Evidence Acquisition for Multimodal Visual Question Answering | 提出Q-Guide以解决文档视觉问答中的证据获取问题 | multimodal | ||
| 12 | MUST-PET: MUltimodal Self-supervised learning across Tracers for whole-body PET/CT-based lesion segmentation | 提出MUST-PET以解决全身PET/CT肿瘤分割的标注稀缺问题 | multimodal | ||
| 13 | ID-VTG: Image-Disambiguated Video Temporal Grounding | 提出ID-VTG以解决视频时间定位中的歧义问题 | multimodal | ✅ | |
| 14 | V-REX: Efficient Specialist VLM Training for Veterinary X-Rays | 提出V-REX以高效训练兽医X光影像专用VLM | foundation model | ||
| 15 | StreamSoccer: Event-Driven Memory for Streaming Soccer Commentary | 提出StreamSoccer以解决实时足球解说中的状态更新问题 | TAMP | ||
| 16 | PL-NBA: A Possession-level Universal Basketball Video Dataset Supporting Multiple Visual Understanding Tasks | 提出PL-NBA数据集以解决篮球视频理解中的时序连续性问题 | TAMP |
🔬 支柱一:机器人控制 (Robot Control) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 17 | DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery | 提出DreamHand以解决视频中手部运动恢复的遮挡问题 | manipulation bi-manual egocentric | ||
| 18 | RoMAN-Flow: Taming Autoregressive Normalizing Flows for Offline Reinforcement Learning in Robotic Manipulation | 提出RoMAN-Flow以解决离线强化学习中的采样瓶颈问题 | manipulation reinforcement learning offline RL | ✅ | |
| 19 | Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis | 提出Inter-X++以解决多模态人际互动分析的瓶颈问题 | dexterous hand multimodal | ||
| 20 | Towards Surgical World-Action Modeling: A Preliminary Joint Visual-Trajectory Forecasting for Surgical Motion Planning | 提出联合视觉-轨迹预测模型以解决外科手术规划问题 | motion planning | ||
| 21 | Zero-Shot Color Image Manipulation Localization via Noise Residual Artifact Pattern Analysis | 提出零-shot方法以解决图像篡改定位问题 | manipulation |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (5 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 22 | Point-Based 3D Reconstruction from Sparse Views under Known Illumination | 提出基于透明度的点渲染方法以解决稀疏视图下的3D重建问题 | 3D reconstruction gaussian splatting splatting | ||
| 23 | Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models | 提出Stream4D以解决视频生成中的几何漂移问题 | 3D reconstruction gaussian splatting splatting | ✅ | |
| 24 | 4DAnyone: Create Anyone in 4D from a Casual Monocular Video | 提出4DAnyone以解决单目视频生成4D人类重建问题 | gaussian splatting splatting | ||
| 25 | Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation | 提出Grounded-Exo2Ego以解决外部视角到内部视角视频生成问题 | 3D reconstruction egocentric | ||
| 26 | Sparse Light Field Sampling Improves Casual 3D and 4D Reconstruction | 提出稀疏光场采样以改善休闲3D和4D重建 | 3DGS |
🔬 支柱四:生成式动作 (Generative Motion) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 27 | Learning to Beat: Phenotype-Guided Latent Flow with Regional Motion Priors for Biventricular Motion Synthesis | 提出基于表型引导的区域运动先验以合成双心室运动 | motion synthesis motion latent | ||
| 28 | STEP: Score-Based Temporal Energy for Human Pose Video Anomaly Detection | 提出STEP框架以解决骨架视频异常检测中的噪声注入问题 | physically plausible | ||
| 29 | When Guidance Goes Off-Scale: Recalibrating Diffusion Transformers under Analog Compute-in-Memory Nonidealities | 提出重校准方法以解决扩散变换器在模拟计算中的非理想性问题 | classifier-free guidance |
🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 30 | G3Ego: Gaze-Guided Graphs for Egocentric Action Understanding | 提出G3Ego以解决第一人称动作理解中的实体识别问题 | egocentric |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 31 | Keep Your Friends Close, and the Right Neighbours Closer: Disaster-Conditioned Kernel-Regularized Graph Attention for Building Damage Classification | 提出灾害条件下的图注意力机制以改进建筑损坏分类 | spatial relationship |