cs.CV(2026-08-31)

📊 共 40 篇论文 | 🔗 10 篇有代码

🎯 兴趣领域导航

支柱三:空间感知与语义 (Perception & Semantics) (14 🔗4) 支柱九:具身大模型 (Embodied Foundation Models) (13 🔗3) 支柱二:RL算法与架构 (RL & Architecture) (12 🔗3) 支柱一:机器人控制 (Robot Control) (1)

🔬 支柱三:空间感知与语义 (Perception & Semantics) (14 篇)

#题目一句话要点标签🔗
1 BRF-GS: Hyperspectral Bidirectional Reflectance Factor Modeling and Image Generation Based on 3D Gaussian Splatting 提出BRF-GS以解决多角度高光谱反射图像生成问题 3D gaussian splatting 3DGS gaussian splatting
2 OPUS: A Simple yet Effective Unified Framework for Open-Vocabulary Detection 提出OPUS框架以简化开放词汇检测任务 open-vocabulary open vocabulary foundation model
3 SMG: Semantic Motion Graph for Monocular Dynamic Gaussian Splatting 提出语义运动图以解决单目动态高斯点云建模问题 gaussian splatting splatting
4 ATGS: Anchored Temporal Gaussian Splatting for Long Volumetric Video Representation 提出ATGS以解决长序列体积视频重建中的不稳定性问题 gaussian splatting splatting
5 Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling 提出Lucida以解决真实场景建模中的实例几何精度问题 scene reconstruction sam 3D SAM 3D
6 CapFrame: Text-Instructed Viewpoint Grounding in 3D Gaussian Scenes via Geometric Pseudo Labels 提出CapFrame以解决3D场景中的文本指令视角定位问题 3D gaussian splatting 3DGS gaussian splatting
7 ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation 提出ScenePilot以解决文本驱动3D室内场景生成问题 open-vocabulary open vocabulary multimodal
8 ObjectSplat: Improving Mesh Fidelity and Interactivity for 3D Scenes via Object-Level Mesh Splatting 提出ObjectSplat以解决3D场景重建中的对象级结构缺失问题 splatting
9 RealOOB: A Definition-Consistent Real-World Oriented Occlusion Boundary Benchmark 提出RealOOB基准以解决现有遮挡边界估计的不足 monocular depth scene understanding
10 AI-enabled Low-Cost 3D Maize Ear Morphometry Platform at Breeding Scale 提出低成本3D玉米穗形态测量平台以解决高通量表型分析问题 NeRF neural radiance field
11 Failure or Drift? Evaluating Monocular SLAM under Synthetic and Real-World Corruptions 评估单目SLAM在合成与真实世界干扰下的表现 visual SLAM
12 Real-Time Video Anomaly Detection Using YOLO Pose Estimation and CLIP-Based Semantic Scoring 提出轻量级框架以实现实时视频异常检测 optical flow
13 FlowVVTON: Flow-Guided Mask-Free Video Virtual Try-On 提出FlowVVTON以解决视频虚拟试穿中的遮挡与不一致问题 optical flow
14 Amortized Anchor Refinement for Deployable Continuous-Time 4D Gaussian Reconstruction 提出可部署的连续时间4D高斯重建方法以解决XR头戴设备计算限制问题 scene flow

🔬 支柱九:具身大模型 (Embodied Foundation Models) (13 篇)

#题目一句话要点标签🔗
15 SurgSkill-Bench: A Benchmark for Multimodal Surgical Skill Assessment 提出SurgSkill-Bench以解决外科技能评估的自动化问题 multimodal
16 TAMI: Temporally Aligned, Missingness-Aware, and Interpretable Multimodal Fusion for Mental Health Assessment in Older Adults with Mild Cognitive Impairment 提出TAMI框架以解决老年人轻度认知障碍的心理健康评估问题 multimodal
17 A Composition-Aware Pretraining Framework for Geospatial Foundation Models 提出一种组合感知预训练框架以解决地理空间模型的单一概念问题 foundation model
18 Lot Machine: Multimodal Lot Extraction from Auction Catalogs 提出多模态管道以自动提取拍卖目录中的结构化信息 multimodal
19 SeqAlign3DVG: A Sequence-Aligned Benchmark and Voxel Reasoning Framework for 3D Visual Grounding 提出SeqAlign3DVG以解决图像基础3D视觉定位的时序对齐问题 visual grounding
20 Semantic-Spatial Discriminability Enhancement for Generalized Visual Grounding 提出SSDE框架以解决多目标视觉定位中的语义与空间可分性问题 visual grounding
21 Doc-REFRAG: Rethinking Multimodal Document Retrieval-Augmented Generation 提出Doc-REFRAG以解决多模态文档检索增强生成问题 multimodal
22 Fine-Grained Multi Image Object Hallucination Benchmark 提出MIOH基准以解决多图像对象幻觉问题 large language model multimodal
23 Cost-efficient Active Learning for Referring Image Segmentation and Grounding 提出成本有效的主动学习方法以解决视觉定位中的标注瓶颈问题 foundation model visual grounding
24 From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents 提出VESTA以解决长视频代理的证据获取问题 multimodal
25 MEOM: Multi-View Expected-OKS Maximization for Human Pose Triangulation 提出MEOM以解决多视角人类姿态三角测量问题 multimodal
26 VisER: Visual Evidence and Reliance for Object Hallucination Detection in LVLMs 提出VisER以解决大规模视觉语言模型中的物体幻觉检测问题 visual grounding
27 DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection 提出DICS以解决视觉指令选择中的数据一致性问题 instruction following

🔬 支柱二:RL算法与架构 (RL & Architecture) (12 篇)

#题目一句话要点标签🔗
28 MR-JEPA: A General Purpose Video Foundation Model for Cardiac MRI 提出MR-JEPA以解决心脏MRI视频数据处理问题 JEPA MAE spatiotemporal
29 VCAR: Training-Free 3DGS Segmentation via View Completeness and Axis-Aware Boundary Refinement 提出VCAR以解决3D高斯分割中的训练开销与边界模糊问题 distillation 3D gaussian splatting 3DGS
30 Modality Disentangled Learning for Incomplete Multimodal Emotion Recognition: A Primitive Memory Distillation Perspective 提出Primitive Memory Distillation框架以解决多模态情感识别中的缺失模态问题 teacher-student distillation multimodal
31 Aligning Multi-Trajectory Supervision with Policy Optimization for VLA Driving 提出多轨迹监督与策略优化对齐框架以提升VLA驾驶性能 imitation learning distillation vision-language-action
32 VisLens: Single-Pass Interpretable Visual Search for Multimodal LLMs 提出VisLens以解决多模态大语言模型的细粒度视觉搜索问题 reinforcement learning large language model multimodal
33 Efficient and High-Quality Depth Estimation via Pixel-Space Diffusion with Linear Attention 提出Lapis以解决高分辨率深度估计的计算效率问题 linear attention depth estimation monocular depth
34 Can Video World Models Track Unobserved World States? 提出视频世界模型以追踪未观察到的世界状态 world model world models Mamba
35 Multimodal Shared Latent Representation of Narration, Microscope and iOCT Images for Phase Recognition in Vitreoretinal Surgery 提出多模态共享潜在表示以解决玻璃体视网膜手术阶段识别问题 contrastive learning multimodal
36 PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Video Representation Learning 提出PRISM以解决跨视角视频表示学习中的语义纠缠问题 representation learning egocentric
37 Rad-R: A Raw-ADC Radar Dataset and Capture-Invariant SSM for Hardware-Fault Diagnosis 提出Rad-R数据集以解决汽车毫米波雷达硬件故障诊断问题 Mamba SSM
38 DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution 提出DreamX-Creator以解决音视频生成同步问题 reinforcement learning multimodal
39 Motion-Saliency Complementary Masked Modeling for Point Cloud Video Understanding 提出MoSaiC以解决点云视频理解中的动态场景建模问题 representation learning scene understanding

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
40 SegWave: Wavelet-Driven Segmentation of Tampered Regions 提出SegWave以解决图像篡改检测问题 manipulation

⬅️ 返回 cs.CV 首页 · 🏠 返回主页