cs.CV(2026-08-20)

📊 共 31 篇论文 | 🔗 5 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (9 🔗2) 支柱九:具身大模型 (Embodied Foundation Models) (7 🔗1) 支柱一:机器人控制 (Robot Control) (5 🔗1) 支柱三:空间感知与语义 (Perception & Semantics) (5 🔗1) 支柱四:生成式动作 (Generative Motion) (3) 支柱六:视频提取与匹配 (Video Extraction) (1) 支柱七:动作重定向 (Motion Retargeting) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (9 篇)

#题目一句话要点标签🔗
1 Open-Vocabulary 3D Object Detection with Co-Distillation Discovery and Dual Guidance Robust Training 提出共蒸馏发现与双重引导训练以解决开放词汇3D目标检测问题 distillation open-vocabulary open vocabulary
2 PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment 提出PEA-DPO以解决多模态偏好优化中的视觉不敏感问题 DPO direct preference optimization large language model
3 Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning 提出Scaffolding Minds以优化多模态推理中的潜在视觉目标表示 reinforcement learning multimodal chain-of-thought
4 RIPE++: Reinforced Keypoint Learning from Positive Pairs Only 提出RIPE++以解决稀疏关键点学习中的监督不足问题 reinforcement learning representation learning visual SLAM
5 CVSD-Reg: Cross-Modal Visual Semantic Prior Distillation for Robust LiDAR Registration 提出CVSD-Reg以解决LiDAR注册的鲁棒性问题 distillation foundation model
6 ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation 提出ArmorOCR以解决对抗性视觉文本识别问题 distillation multimodal
7 Flow Matching-Based PET Image Reconstruction 提出基于流匹配的PET图像重建方法以提升重建质量 flow matching
8 Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs 提出OraRL以解决视频多模态大语言模型的样本效率问题 reinforcement learning large language model multimodal
9 RISE: Adaptive Imagination for World Action Models 提出RISE框架以解决固定想象预算问题 world model world models world action model

🔬 支柱九:具身大模型 (Embodied Foundation Models) (7 篇)

#题目一句话要点标签🔗
10 Robust Cross-Modal Foundation Model Perception for Underwater Robots under Degraded Visual Conditions 提出跨模态基础模型以提升水下机器人在恶劣视觉条件下的感知能力 foundation model multimodal
11 Question-Guided Evidence Acquisition for Multimodal Visual Question Answering 提出Q-Guide以解决文档视觉问答中的证据获取问题 multimodal
12 MUST-PET: MUltimodal Self-supervised learning across Tracers for whole-body PET/CT-based lesion segmentation 提出MUST-PET以解决全身PET/CT肿瘤分割的标注稀缺问题 multimodal
13 ID-VTG: Image-Disambiguated Video Temporal Grounding 提出ID-VTG以解决视频时间定位中的歧义问题 multimodal
14 V-REX: Efficient Specialist VLM Training for Veterinary X-Rays 提出V-REX以高效训练兽医X光影像专用VLM foundation model
15 StreamSoccer: Event-Driven Memory for Streaming Soccer Commentary 提出StreamSoccer以解决实时足球解说中的状态更新问题 TAMP
16 PL-NBA: A Possession-level Universal Basketball Video Dataset Supporting Multiple Visual Understanding Tasks 提出PL-NBA数据集以解决篮球视频理解中的时序连续性问题 TAMP

🔬 支柱一:机器人控制 (Robot Control) (5 篇)

#题目一句话要点标签🔗
17 DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery 提出DreamHand以解决视频中手部运动恢复的遮挡问题 manipulation bi-manual egocentric
18 RoMAN-Flow: Taming Autoregressive Normalizing Flows for Offline Reinforcement Learning in Robotic Manipulation 提出RoMAN-Flow以解决离线强化学习中的采样瓶颈问题 manipulation reinforcement learning offline RL
19 Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis 提出Inter-X++以解决多模态人际互动分析的瓶颈问题 dexterous hand multimodal
20 Towards Surgical World-Action Modeling: A Preliminary Joint Visual-Trajectory Forecasting for Surgical Motion Planning 提出联合视觉-轨迹预测模型以解决外科手术规划问题 motion planning
21 Zero-Shot Color Image Manipulation Localization via Noise Residual Artifact Pattern Analysis 提出零-shot方法以解决图像篡改定位问题 manipulation

🔬 支柱三:空间感知与语义 (Perception & Semantics) (5 篇)

#题目一句话要点标签🔗
22 Point-Based 3D Reconstruction from Sparse Views under Known Illumination 提出基于透明度的点渲染方法以解决稀疏视图下的3D重建问题 3D reconstruction gaussian splatting splatting
23 Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models 提出Stream4D以解决视频生成中的几何漂移问题 3D reconstruction gaussian splatting splatting
24 4DAnyone: Create Anyone in 4D from a Casual Monocular Video 提出4DAnyone以解决单目视频生成4D人类重建问题 gaussian splatting splatting
25 Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation 提出Grounded-Exo2Ego以解决外部视角到内部视角视频生成问题 3D reconstruction egocentric
26 Sparse Light Field Sampling Improves Casual 3D and 4D Reconstruction 提出稀疏光场采样以改善休闲3D和4D重建 3DGS

🔬 支柱四:生成式动作 (Generative Motion) (3 篇)

#题目一句话要点标签🔗
27 Learning to Beat: Phenotype-Guided Latent Flow with Regional Motion Priors for Biventricular Motion Synthesis 提出基于表型引导的区域运动先验以合成双心室运动 motion synthesis motion latent
28 STEP: Score-Based Temporal Energy for Human Pose Video Anomaly Detection 提出STEP框架以解决骨架视频异常检测中的噪声注入问题 physically plausible
29 When Guidance Goes Off-Scale: Recalibrating Diffusion Transformers under Analog Compute-in-Memory Nonidealities 提出重校准方法以解决扩散变换器在模拟计算中的非理想性问题 classifier-free guidance

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
30 G3Ego: Gaze-Guided Graphs for Egocentric Action Understanding 提出G3Ego以解决第一人称动作理解中的实体识别问题 egocentric

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
31 Keep Your Friends Close, and the Right Neighbours Closer: Disaster-Conditioned Kernel-Regularized Graph Attention for Building Damage Classification 提出灾害条件下的图注意力机制以改进建筑损坏分类 spatial relationship

⬅️ 返回 cs.CV 首页 · 🏠 返回主页