cs.CV(2026-08-19)

📊 共 28 篇论文 | 🔗 8 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (9 🔗2) 支柱九:具身大模型 (Embodied Foundation Models) (8 🔗4) 支柱三:空间感知与语义 (Perception & Semantics) (7) 支柱六:视频提取与匹配 (Video Extraction) (2 🔗1) 支柱八:物理动画 (Physics-based Animation) (2 🔗1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (9 篇)

#题目一句话要点标签🔗
1 Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI 提出视觉-语言模型以解决自我中心视频理解问题 representation learning egocentric human-to-robot
2 CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes 提出CL4D以解决动态场景中的视觉语言推理问题 contrastive learning metric depth text-to-motion
3 DynCur-Geo: Dynamic Curiosity Reward Shaping for Multimodal Active Geo-Localization 提出DynCur-Geo以解决多模态主动地理定位中的探索与收敛平衡问题 reward shaping multimodal
4 GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery 提出GrabVG以解决无人机图像中的视觉定位问题 distillation visual grounding
5 PCQA-R1: Advancing Generalized 3D Point Cloud Quality Assessment with Reinforcement Learning 提出PCQA-R1以解决无参考点云质量评估问题 reinforcement learning multimodal chain-of-thought
6 Falcon Perception-HD: High Density Perception via Reinforcement Learning 提出Falcon Perception-HD以解决高密度场景下的感知问题 reinforcement learning reward design open-vocabulary
7 Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models 提出SparsePR以解决视频生成中的稀疏注意力问题 world model world models
8 FD-CanKD: Frequency-Decoupled Cross-Attention Distillation as a Refinement Prior for Compact Object Detectors 提出FD-CanKD以解决紧凑型目标检测器的知识蒸馏问题 teacher-student distillation
9 VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation 提出VA-Judger以解决视频音频生成中的人类偏好反馈问题 reinforcement learning chain-of-thought

🔬 支柱九:具身大模型 (Embodied Foundation Models) (8 篇)

#题目一句话要点标签🔗
10 OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios 提出OmniHandwritingOCR以解决手写OCR评估的不足问题 large language model multimodal visual grounding
11 Subgroup performance analysis of adaptation strategies for chest X-ray foundation models 研究适应策略对胸部X光模型子群公平性的影响 foundation model
12 When Two Tracers Disagree: An Investigation of Multimodal Fusion for Clinical PET/CT Segmentation 提出多模态融合方法以改善临床PET/CT肿瘤分割 multimodal
13 A Critical Synthesis of Uncertainty Quantification and Foundation Models for Semantic Segmentation 提出不确定性量化方法以提升语义分割模型的可靠性 foundation model
14 Teach a Molmo2Fish: Towards interactive fish tracking with natural language guidance 提出Molmo2Fish以解决鱼类追踪中的自然语言指导问题 large language model multimodal
15 MR-IQA-2: Faithful Image Quality Reflection via Fine-Grained Credit Assignment 提出MR-IQA-2以解决图像质量评估中的信实性问题 large language model multimodal
16 Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts 提出迭代微调传统OCR管道以提升复杂历史梵文手稿的转录准确性 large language model
17 When Safety Overrides Vision: Exploring Dynamics between Vision Influence and Safety Alignment in Vision-Language Models 探讨安全对视觉语言模型生成行为的影响 multimodal

🔬 支柱三:空间感知与语义 (Perception & Semantics) (7 篇)

#题目一句话要点标签🔗
18 GS-VLA: Plug-and-Play Viewpoint Canonicalization for Frozen VLA Policies via Gaussian Splatting 提出GS-VLA以解决视觉-语言-动作政策中的视角偏移问题 gaussian splatting splatting vision-language-action
19 CoMVS-GS: Collaborative Multi-View Stereo and 3D Gaussian Splatting for Surface Reconstruction 提出CoMVS-GS以解决3D重建中的几何不一致问题 3D gaussian splatting 3DGS gaussian splatting
20 RVLoss: Runoff Vote Loss for Self-Supervised LiDAR Scene Flow Estimation 提出RVLoss以解决自监督LiDAR场景流估计中的运动一致性问题 scene flow
21 Evaluation of Image Matching Methods for Visual Odometry on UAVs 评估图像匹配方法以解决无人机视觉里程计问题 visual odometry
22 COSTA: A Cluster-Centric Paradigm for Annotation-Free Open-Set Semantic Segmentation of Aerial Point Clouds with Domain Shifts 提出COSTA以解决航空点云开放集语义分割问题 open-vocabulary open vocabulary
23 ForeSightGuide: An Anticipatory Framework toward Accurate and Low-Redundancy Guidance for the Visually Impaired 提出ForeSightGuide以解决视觉障碍者导航信息过载问题 scene understanding
24 CamWorldQA: Perceptual Quality Assessment of Camera-Controlled World Video Generation 提出CamWorldQA以解决摄像机控制视频生成的质量评估问题 optical flow

🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)

#题目一句话要点标签🔗
25 EgoHRV: Continuous Heart Rate Variability Estimation from Egocentric Systems for Autonomic Response and Skill Assessment 提出EgoHRV以解决自我中心系统中心率变异性估计问题 egocentric egocentric vision PULSE
26 Generalized Audio-Driven Synthesis of Precise Drummer Motion 提出生成扩散框架以解决音频驱动的鼓手动作合成问题 motion matching character animation

🔬 支柱八:物理动画 (Physics-based Animation) (2 篇)

#题目一句话要点标签🔗
27 X-LMC: Cross-View Spatiotemporal Collateral Circulation Scoring from DSA 提出X-LMC以解决DSA下的LMC自动评分问题 spatiotemporal
28 StateTrace: An Object-Centric Framework for Hidden-State Spatiotemporal Reasoning in Long Videos 提出StateTrace以解决长视频中的隐状态时空推理问题 spatiotemporal

⬅️ 返回 cs.CV 首页 · 🏠 返回主页