cs.CV(2026-08-25)
📊 共 25 篇论文 | 🔗 7 篇有代码
🎯 兴趣领域导航
支柱九:具身大模型 (Embodied Foundation Models) (9 🔗2)
支柱二:RL算法与架构 (RL & Architecture) (6 🔗2)
支柱一:机器人控制 (Robot Control) (3 🔗1)
支柱四:生成式动作 (Generative Motion) (2 🔗1)
支柱六:视频提取与匹配 (Video Extraction) (2 🔗1)
支柱七:动作重定向 (Motion Retargeting) (1)
支柱八:物理动画 (Physics-based Animation) (1)
支柱三:空间感知与语义 (Perception & Semantics) (1)
🔬 支柱九:具身大模型 (Embodied Foundation Models) (9 篇)
🔬 支柱二:RL算法与架构 (RL & Architecture) (6 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 10 | Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training | 提出Game2World引擎以解决游戏视频训练数据问题 | world model world models multimodal | ✅ | |
| 11 | TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation | 提出TurboT2VA以解决大规模文本到视频音频生成的效率问题 | distillation multimodal | ✅ | |
| 12 | On-Policy Self-Distillation in Diffusion Models | 提出DiffusionOPSD以解决扩散模型与人类偏好对齐问题 | reinforcement learning distillation | ||
| 13 | Joint-Embedding Prediction of Masked Point Tubes for Self-Supervised Learning on 4D Point Cloud Videos | 提出JEPA框架以解决4D点云视频自监督学习问题 | JEPA representation learning spatiotemporal | ||
| 14 | IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves | 提出IDeaL方法以解决无数据多教师蒸馏问题 | distillation | ||
| 15 | Representation Learning in Diffusion and Flow-based Model: An Application Aspect | 提出三层进阶框架以提升生成模型的表示学习能力 | representation learning |
🔬 支柱一:机器人控制 (Robot Control) (3 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 16 | LeFlow: Generative Latent Flow Planning for World Models | 提出LeFlow以解决世界模型中的规划效率问题 | trajectory optimization world model world models | ✅ | |
| 17 | NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation | 提出NeoWorld-Pro以解决单目图像到交互场景构建问题 | manipulation scene reconstruction physically plausible | ||
| 18 | VizAnchor: Decoding Manipulation Intent from Tampering Visualizations via Dual-Anchor Reasoning | 提出VizAnchor以解决数据可视化操控理解问题 | manipulation |
🔬 支柱四:生成式动作 (Generative Motion) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 19 | SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling | 提出SeMoCo以解决运动语言建模中的语义编码问题 | text-to-motion language-conditioned motion motion generation | ||
| 20 | Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis | 提出基于元数据的生成模型以改善心脏磁共振图像合成 | classifier-free guidance foundation model | ✅ |
🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 21 | From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms | 提出统一框架以提升智能眼镜的第一人称智能平台能力 | egocentric egocentric vision multimodal | ||
| 22 | EgoErrorVQA: Assess Egocentric Comprehension Capabilities through Procedural Errors for Ego-Agentic AI | 提出EgoErrorVQA以解决程序性理解能力评估问题 | egocentric | ✅ |
🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 23 | Event-Based Motion Estimation via Oriented Distance Fields | 提出定向距离场运动估计以解决事件驱动运动估计的低延迟问题 | motion estimation |
🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 24 | Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning | 提出多模态框架以解决在线学习中学生参与度预测问题 | spatiotemporal multimodal |
🔬 支柱三:空间感知与语义 (Perception & Semantics) (1 篇)
| # | 题目 | 一句话要点 | 标签 | 🔗 | ⭐ |
|---|---|---|---|---|---|
| 25 | KLTNet: Learning Sparse Feature Tracking for Robust and Accurate Monocular Visual-Inertial Odometry | 提出KLTNet以解决稀疏特征跟踪的鲁棒性与准确性问题 | VIO optical flow |