cs.CV(2026-09-02)

📊 共 26 篇论文 | 🔗 8 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (11 🔗5) 支柱三:空间感知与语义 (Perception & Semantics) (6 🔗1) 支柱二:RL算法与架构 (RL & Architecture) (5) 支柱一:机器人控制 (Robot Control) (3 🔗2) 支柱六:视频提取与匹配 (Video Extraction) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (11 篇)

#题目一句话要点标签🔗
1 MARS: What Retrieval Signals Are Hidden in Multimodal Large Language Models for Text-Video Retrieval? 提出MARS框架以解决文本-视频检索中的信息压缩问题 large language model multimodal
2 LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory 提出LookStep以解决视觉语言导航中的资源效率问题 VLN large language model multimodal
3 Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment 提出TIC-Bench以解决多模态模型评估中的文本-图像深度交互问题 multimodal
4 Rendering-in-the-Loop: An Execution-Driven Agent for Interactive Web Development 提出RILA以解决交互网页开发中的功能验证问题 large language model foundation model multimodal
5 Morphology signal in whole slide image foundation models can automatically triage slides 提出基于形态信号的模型以自动筛选全切片图像 foundation model
6 ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding 提出ShallowStream以解决流媒体视频理解中的计算开销问题 large language model multimodal
7 Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness 研究表明多模态大语言模型系统性高估面部吸引力 large language model multimodal
8 YesTrack: Referring Multi-Object Tracking via MLLM-based Yes/No Verification 提出YesTrack以解决多目标跟踪中的语言表达匹配问题 large language model multimodal
9 Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding 提出轻量适应方法以解决多光谱和SAR图像理解问题 foundation model instruction following
10 Detecting Object Hallucinations in Large Vision-Language Models via Cross-Modal Attention Drifts and Mask-Based Verification 提出CADMP框架以解决大规模视觉-语言模型中的对象幻觉问题 visual grounding
11 Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework 提出因果驱动评估框架以解析VLM生成过程 multimodal

🔬 支柱三:空间感知与语义 (Perception & Semantics) (6 篇)

#题目一句话要点标签🔗
12 Adapting a Foundation Model for Lunar Surface Height Estimation 提出基于深度学习的月球表面高度估计方法 depth estimation monocular depth Depth Anything
13 TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models 提出TempoGround以解决流媒体视觉定位中的一致性问题 open-vocabulary open vocabulary visual grounding
14 InceptionGS: Generative Bootstrapping for Large-Scale Gaussian Splatting under Unstructured View Sampling 提出InceptionGS以解决大规模场景数字化中的视角稀缺问题 gaussian splatting splatting
15 CC-4DGS: Computational Deformation and Point-Cloud Compression for Storage-Efficient Dynamic Gaussian Splatting 提出CC-4DGS以解决动态高斯点云存储效率问题 gaussian splatting splatting
16 WiFlow: Estimating Optical Flow using WiFi Channel State Information 提出WiFlow以解决光流估计中的隐私与环境影响问题 optical flow
17 Query Rewriting for Complex Object Segmentation in 4D Gaussian Representations 提出查询重写方法以解决4D高斯表示中的复杂对象分割问题 scene understanding

🔬 支柱二:RL算法与架构 (RL & Architecture) (5 篇)

#题目一句话要点标签🔗
18 Spatially Aware World Action Model via Geometric Latent Diffusion 提出空间感知世界行动模型以解决3D信息缺失问题 policy learning world model world models
19 SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models 提出SolarWM以解决长视频世界模型构建的挑战 world model world models distillation
20 World-Coherent Decoding: Self-Verifying Test-Time Planning for World Action Models 提出世界一致解码以解决机器人决策不确定性问题 world action model world action models
21 Balancing Frequencies and Pixels in Flow Matching 提出Focal Log-Frequency Loss以解决像素空间流匹配中的频谱不平衡问题 flow matching
22 CA-OPD: Confidence-Aware On-Policy Distillation for Structured Visual Prediction 提出CA-OPD以解决自回归视觉语言模型中的错误累积问题 distillation

🔬 支柱一:机器人控制 (Robot Control) (3 篇)

#题目一句话要点标签🔗
23 Towards Zero-Shot Transfer Across Embodiments For Driving VLAs 提出多数据集训练与BEV-Forcing以解决驾驶VLA的零-shot迁移问题 manipulation cross-embodiment vision-language-action
24 From Detection to Localization: A Unified Forensics Framework for Fully Synthetic and Tampered Images 提出统一框架以解决图像真实性检测与定位问题 manipulation
25 Evidence-Guided Detection, Localization and Explanation for Text-Centric Image Forensics 提出证据引导的检测、定位与解释系统以解决文本中心图像取证问题 manipulation

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
26 AutoCompass: Accurate Visual Localization on Public Maps by Learning from Weak Labels 提出AutoCompass以解决视觉定位中的标签噪声问题 egocentric

⬅️ 返回 cs.CV 首页 · 🏠 返回主页