cs.CV(2026-08-06)

📊 共 48 篇论文 | 🔗 7 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (18 🔗4) 支柱九:具身大模型 (Embodied Foundation Models) (14 🔗1) 支柱三:空间感知与语义 (Perception & Semantics) (6 🔗1) 支柱一:机器人控制 (Robot Control) (4 🔗1) 支柱八:物理动画 (Physics-based Animation) (3) 支柱七:动作重定向 (Motion Retargeting) (1) 支柱五:交互与反应 (Interaction & Reaction) (1) 支柱六:视频提取与匹配 (Video Extraction) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (18 篇)

#题目一句话要点标签🔗
1 G$^2$ARD-GS: Geometry-Guided Anchor-Regularized Gaussian Splatting Distillation 提出G$^2$ARD-GS以解决稠密LiDAR地图的高存储成本问题 distillation 3D gaussian splatting 3DGS
2 Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture 提出Bar-JEPA以解决条形图数据提取问题 JEPA Joint-Embedding Predictive Architecture joint-embedding predictive architecture
3 Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval 提出UniME-R1以解决多模态检索中的理解偏差问题 reinforcement learning multimodal chain-of-thought
4 Wan-Animate-2: Pushing the Application Boundaries of Character Animation 提出Wan-Animate-2以解决角色动画实时性与精度问题 distillation motion representation character animation
5 LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models 提出LAWM-3D以解决机器人世界模型中的3D感知问题 world model world models foundation model
6 HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models 提出HERA以解决物理预测中的历史证据访问问题 world model world models JEPA
7 MASS: Multiplayer World Models with Authoritative Shared State 提出MAS以解决多玩家环境中的世界模型问题 world model world models
8 Uncertainty-Aware World Model for Aerial Image-Goal Navigation 提出不确定性感知世界模型以解决无人机图像目标导航问题 world model world models
9 Curia-MAE: Multi-Modal Multi-Anatomy MAE Pre-Training for 3D Medical Image Segmentation 提出Curia-MAE以提升3D医学图像分割性能 MAE foundation model
10 SR-JEPA: Learning Predictive Latent State in 3D Scenes 提出SR-JEPA以解决3D场景中缺失实体的预测问题 JEPA Joint-Embedding Predictive Architecture joint-embedding predictive architecture
11 ChronoVision: Temporal Reasoning via Latent State Reconstruction 提出ChronoVision以解决多步时间推理问题 reinforcement learning large language model multimodal
12 The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents 提出Gated Hindsight Distillation以解决GUI代理训练中的信息缺失问题 distillation privileged information
13 Flow-Map Distillation on Relation Manifolds for Image Restoration 提出Flow-Map蒸馏方法以提升图像恢复性能 flow matching distillation
14 Hierarchical Flow Matching for 3D Point Cloud Generation 提出层次流匹配方法以生成高质量3D点云 flow matching
15 Dense-Cast: A lightweight ensemble of deep learning architectures for precipitation nowcasting 提出Dense-Cast以解决降水短期预报问题 MAE spatiotemporal
16 SLED: Scalable Location Encoding via Distillation 提出SLED以解决地理空间数据编码效率低的问题 distillation spatiotemporal multimodal
17 Curia-MAE: Multi-Modal Multi-Anatomy MAE Pre-Training for 3D Medical Image Segmentation 提出Curia-MAE以提升3D医学图像分割性能 MAE foundation model
18 InsertFuse: A Unified Framework for Multi-Category Reference-Guided Image Insertion 提出InsertFuse框架以解决多类别参考引导图像插入问题 flow matching distillation

🔬 支柱九:具身大模型 (Embodied Foundation Models) (14 篇)

#题目一句话要点标签🔗
19 Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models 提出3D CT基础模型基准以评估诊断能力 foundation model
20 STAIL: Semantic Text-Anchored Incremental Learning for Medical Imaging via Large Language Models 提出STAIL以解决医疗影像增量学习中的遗忘问题 large language model
21 KVAE: Family of Tokenizers for Multimodal Generative Models 提出KVAE系列分词器以提升多模态生成模型性能 multimodal
22 SciQNet: Two-Stage Multimodal Adaptation for Scientific Image Quality Assessment 提出SciQNet以解决科学图像质量评估问题 multimodal
23 Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation 提出Vorch-IR以解决多模态身份替换视频生成问题 multimodal
24 CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images 提出CoordRefer以解决3D视觉定位中的坐标框架选择问题 visual grounding
25 UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on 提出UniVVT框架以解决视频虚拟试穿中的几何误差传播问题 large language model multimodal
26 Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs 提出VideoRouter以解决长视频理解中的证据捕捉问题 large language model multimodal
27 Multi-Year Geospatial Reasoning using Interannually-Consistent Historical Predictions as a Free Input Modality 提出基于历史预测的多年度地理空间推理方法以提升作物分类精度 foundation model
28 ConceptADapt: Concept-guided Adaptive Feature Reconstruction with Dynamic Attention for Few-Shot Industrial Anomaly Detection 提出ConceptADapt以解决少样本工业异常检测问题 foundation model
29 One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding 提出Matryoshka框架以解决长视频理解中的帧选择问题 multimodal
30 StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding 提出StreamArena以解决长视频理解中的交互与记忆问题 multimodal
31 TAU-Bench: From Anomaly Instance Tracking to Fine-Grained Video Anomaly Understanding 提出TAU-Bench以解决视频异常理解中的实例跟踪与语义一致性问题 visual grounding
32 Do 3D Medical Foundation Models See Through MRI Artifacts? A Controlled Study of Representation Robustness 评估3D医学基础模型对MRI伪影的鲁棒性 foundation model

🔬 支柱三:空间感知与语义 (Perception & Semantics) (6 篇)

#题目一句话要点标签🔗
33 OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction 提出OmniMech基准以解决工业机械设计中的3D重建问题 3D reconstruction multimodal
34 MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers 提出MAVISEG以解决零-shot开放词汇分割问题 open-vocabulary open vocabulary
35 SCI-CLIP: Segment-Centric Inference with Reference Memory for Training-Free Open-Vocabulary Segmentation 提出SCI-CLIP以解决训练无关的开放词汇分割问题 open-vocabulary open vocabulary
36 Confidence matters: Leveraging Multi-view Geometric Priors for GS-based Reconstruction 提出多视角几何先验以提升3D高斯点云重建质量 3D gaussian splatting 3DGS gaussian splatting
37 Universal Concept Disruption for SAM3 Image Segmentation 提出通用概念干扰以增强SAM3图像分割的对抗鲁棒性 open-vocabulary open vocabulary
38 CDSeg: A Renderable Gaussian Carrier for Image-to-3D Label Transfer 提出CDSeg以解决图像到3D标签转移问题 gaussian splatting splatting

🔬 支柱一:机器人控制 (Robot Control) (4 篇)

#题目一句话要点标签🔗
39 PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models 提出PhyLatent以解决JEPA世界模型中的动态表示问题 MPC model predictive control world model
40 Dual-Attention and Adversarial Transfer Networks for Sim-to-Real Cross-Orientation Wireless Sensing 提出双重注意力与对抗转移网络以解决无线传感中的方向变化问题 sim-to-real
41 Vorch-Omni: Multi-Task Orchestration of Sight and Sound 提出Vorch-Omni以解决多任务音视频生成问题 manipulation flow matching
42 Floating Radiance Networks 提出Floating Radiance Networks以解决传统渲染方法的局限性 manipulation

🔬 支柱八:物理动画 (Physics-based Animation) (3 篇)

#题目一句话要点标签🔗
43 Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding 提出EviSelect以解决长视频理解中的动态视觉选择问题 spatiotemporal TAMP
44 EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal 提出EffectLearner以解决复杂视频对象去除问题 spatiotemporal
45 Implicit Neural Speckle Denoising 提出一种无训练框架以解决动态场景中的散斑噪声问题 spatiotemporal

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
46 Learning visual representations for compositional analysis of artworks and photographs 提出基于人类启发的视觉表示学习以分析艺术作品的构图 spatial relationship foundation model

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
47 HOPE: Hand-Object Pressure Estimation from Monocular Videos 提出HOPE框架以解决动态手-物体压力估计问题 HOI egocentric

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
48 GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? 提出GST-Bench以解决视频理解中的全球空间意识问题 egocentric

⬅️ 返回 cs.CV 首页 · 🏠 返回主页