cs.CV(2026-08-04)

📊 共 43 篇论文 | 🔗 12 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (17 🔗5) 支柱二:RL算法与架构 (RL & Architecture) (12 🔗4) 支柱三:空间感知与语义 (Perception & Semantics) (8 🔗2) 支柱七:动作重定向 (Motion Retargeting) (2) 支柱四:生成式动作 (Generative Motion) (1) 支柱一:机器人控制 (Robot Control) (1) 支柱五:交互与反应 (Interaction & Reaction) (1) 支柱六:视频提取与匹配 (Video Extraction) (1 🔗1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (17 篇)

#题目一句话要点标签🔗
1 From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology 提出多分辨率金字塔变换器以解决病理图像分析中的分辨率限制问题 large language model foundation model multimodal
2 ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs 提出ParVL框架以优化多模态大语言模型的计算分配 large language model multimodal
3 OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models 提出OmniPack以解决多模态大语言模型的高计算开销问题 large language model multimodal
4 LocAnyMed: Vision-Language Grounding for Multimodal Medical Images 提出LocAnyMed以解决多模态医学图像的视觉语言定位问题 multimodal visual grounding
5 CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation 提出CorePath以解决乳腺核心针活检诊断挑战 foundation model multimodal
6 LDU-Bench: Multimodal LLM Evaluation for Lithography Defect Understanding under Layout-Varying Circuit Backgrounds 提出LDU-Bench以解决光刻缺陷理解中的多任务评估问题 large language model multimodal
7 Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis 提出多模态框架以实现高效植物根系表型分析 multimodal
8 When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware 提出视觉标记优化策略以加速多模态推理 multimodal
9 Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding 提出Hi-Token以解决视觉定位中的坐标表示问题 visual grounding
10 Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models 提出时间采样门控机制以解决视频大语言模型的偏差问题 large language model
11 GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models 提出GSTEP以解决视频大语言模型中的冗余视觉标记问题 large language model
12 SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models 提出SlimVLM以解决视觉语言模型的高计算开销问题 large language model multimodal
13 CIGTSurv: Clinical Information Guided Tri-modal Survival Prediction with Local Prototype Association and Global Feature Alignment 提出CIGTSurv以解决临床信息在生存预测中的低利用问题 foundation model multimodal
14 UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space 提出UHP检测以解决LVLM模型幻觉检测问题 multimodal
15 Attention is Case-Sensitive 提出字母大小写敏感性以优化注意力分配 large language model
16 MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification 提出MT-Web2Code以解决多轮区域重建与局部修改问题 multimodal
17 Frequency-Decorrelated Temporal Ensembles for EEG--fNIRS Imagined-Handwriting Decoding 提出FRED系统以解决EEG-fNIRS想象手写解码问题 multimodal

🔬 支柱二:RL算法与架构 (RL & Architecture) (12 篇)

#题目一句话要点标签🔗
18 Residual Flow Matching with Dynamic Cross-Interaction for 3D Multi-Person Motion Prediction 提出基于残差流匹配的动态交互机制以提升3D多人动作预测精度 flow matching motion prediction
19 Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing 提出DinoSplat-OV以解决遥感语义分割的训练成本问题 contrastive learning gaussian splatting splatting
20 DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack 提出DRIFT以解决流匹配VLA模型的对抗攻击问题 flow matching vision-language-action VLA
21 Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent 提出Video-DeepResearch以解决多模态视频理解问题 imitation learning spatiotemporal multimodal
22 Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning 提出能力感知环境选择与分层难度课程以优化多模态代理学习 curriculum learning multimodal
23 Compass: Degradation-Simulated Reciprocal Learning with Lightweight Needle RWKV for Multimodal Crack Segmentation under Missing Modalities 提出Compass以解决多模态裂缝分割中的缺失模态问题 distillation multimodal
24 CrossScope: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction 提出CrossScope以解决多观察者协作系统的未来预测问题 world model world models
25 CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation 提出CROSS以解决遥感图像分割中的定位漂移与语义偏差问题 contrastive learning distillation
26 Efficient Video Dataset Distillation via Cluster-Guided Prototype Blending 提出ProtoBlend以解决视频数据集蒸馏效率问题 distillation
27 Self-Supervised Representation-Guided Generative Dataset Distillation 提出自监督表示引导的生成数据集蒸馏方法以提升模型性能 distillation
28 Channel-wise Dynamic Knowledge Distillation via Adaptive Sample Generation for Action Recognition 提出自适应样本生成的通道动态知识蒸馏方法以提升动作识别性能 distillation
29 Low-Dimensional High-Leverage Subspace Optimization: Beyond Full-Parameter Coupled Training for Neural Network Quantization 提出低维高杠杆子空间优化以解决神经网络量化问题 teacher-student distillation

🔬 支柱三:空间感知与语义 (Perception & Semantics) (8 篇)

#题目一句话要点标签🔗
30 3DGSI-Assessor: A Large-Scale Dataset and An LMM-based Method for 3D Gaussian Splatting Image Quality Assessment 提出3DGSI-Assessor以解决3D高斯点云图像质量评估问题 3D gaussian splatting 3DGS gaussian splatting
31 Perceptual Anchoring: Prototype-Guided Text Calibration for Training-free Open-Vocabulary Semantic Segmentation 提出原型引导文本校准方法以解决开放词汇语义分割问题 open-vocabulary open vocabulary
32 XiDepth: a Lightweight and Efficient Network for Self-supervised Monocular Depth Estimation 提出XiDepth以解决自监督单目深度估计的轻量化问题 depth estimation monocular depth
33 Global Graph-Validated Optimization for VLM-based 3D Indoor Scene Generation 提出图验证优化方法以解决开放词汇3D室内场景生成问题 open-vocabulary open vocabulary physically plausible
34 SUV: Future Scene Understanding as Video Generation for End-to-End Driving 提出SUV框架以解决未来场景理解问题 scene understanding foundation model
35 SGFormer: Structure-Guided Transformer for Robust Local Feature Matching 提出SGFormer以解决特征匹配中的注意力分散问题 3D reconstruction feature matching
36 SLAMFormer-$\infty$: Infinite SLAM Transformer for Unbounded Frontend and Backend Processing 提出SLAMFormer-$\infty$以解决无限SLAM处理问题 scene reconstruction
37 PolyLayout: Multi-room Manhattan Layout Estimation 提出PolyLayout以解决多房间曼哈顿布局估计问题 scene understanding

🔬 支柱七:动作重定向 (Motion Retargeting) (2 篇)

#题目一句话要点标签🔗
38 Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding 提出Geo-Embed以解决城市理解中的多模态嵌入问题 spatial relationship multimodal visual grounding
39 TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding 提出TDVR框架以解决零样本3D视觉定位中的文本歧义与视角不足问题 spatial relationship visual grounding chain-of-thought

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
40 Learning Biomechanically Plausible Human Motion from Sparse Radar Point Clouds 提出基于雷达点云的人体运动生物力学估计方法 physically plausible human motion motion prediction

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
41 LiteMVS: Efficient Multi-View Stereo with Foundation Distillation and Expert Aggregation 提出LiteMVS以解决多视角立体视觉中的深度估计问题 manipulation representation learning distillation

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
42 Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation 提出表面关键点轨迹以解决多物体和关节人机交互生成问题 human-object interaction HOI OMOMO

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
43 From Routes to Steps: Separating Semantic Progress from Local Execution in Vision-and-Language Navigation 提出Route2Step以解决视觉语言导航中的进度跟踪与执行错误问题 egocentric VLN

⬅️ 返回 cs.CV 首页 · 🏠 返回主页