cs.CV(2026-07-28)

📊 共 33 篇论文 | 🔗 6 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (16 🔗3) 支柱二:RL算法与架构 (RL & Architecture) (7 🔗1) 支柱三:空间感知与语义 (Perception & Semantics) (6 🔗2) 支柱一:机器人控制 (Robot Control) (2) 支柱八:物理动画 (Physics-based Animation) (1) 支柱六:视频提取与匹配 (Video Extraction) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (16 篇)

#题目一句话要点标签🔗
1 SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models 提出SepPrune框架以高效修剪多模态大语言模型 large language model multimodal
2 Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition 提出PRISM-AH框架以解决视频级别的模糊性与犹豫性识别问题 large language model multimodal
3 Food Image Segmentation with LLM-Derived Ingredient Labels and Multimodal Fusion 提出多模态模块以提升食品图像分割性能 large language model multimodal
4 Towards Reliable Stain Transfer: An Iterative Data-Model Co-Optimization Framework Based on Multimodal Expert-Guided Assessment 提出DMCoStain框架以解决异质生物标志物的染色转移问题 multimodal instruction following
5 VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening 提出VetClaw以解决兽医疾病筛查问题 multimodal
6 LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection 提出LaP-Forensics以解决深度伪造检测中的视觉伪影问题 multimodal
7 Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models 提出CRoMa以解决病理基础模型的稳健性问题 foundation model
8 Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation 提出代理智能以解决医学领域多步骤临床任务的挑战 large language model foundation model multimodal
9 CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition 提出CLBench-V以解决多模态上下文学习评估问题 multimodal
10 FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding 提出FORGE以解决长视频理解中的信息选择问题 large language model multimodal
11 MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities 提出MODUS以解决多模态建模的局限性问题 multimodal
12 Fine-Grained Food Image Understanding via Target-Aware Data Alignment 提出目标感知数据对齐方法以解决细粒度食品图像理解问题 multimodal
13 Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications 提出基于指令的图像编辑方法以解决现有编辑系统的局限性 large language model
14 Visual prompt engineering for video models 提出视觉提示工程以提升视频模型的推理性能 foundation model
15 Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation 提出Argus-Unified以解决多模态理解与生成的高成本问题 multimodal
16 CD-RMOT-Bench: Benchmarking the Cross-Domain Referring Multi-Object Tracking 提出CD-RMOT-Bench以解决跨域语言引导多目标跟踪问题 language conditioned

🔬 支柱二:RL算法与架构 (RL & Architecture) (7 篇)

#题目一句话要点标签🔗
17 OrganLens: Organ-Specific Representation Learning for CT Foundation Models 提出OrganLens以解决CT图像中器官特定表示学习问题 representation learning distillation foundation model
18 Medical world models in healthcare: foundations, applications, and challenges for trustworthy clinical translation 提出医疗世界模型以解决临床翻译中的信任问题 world model world models multimodal
19 Wonder: Video World Model Done Better 提出Wonder以实现实时可控的视频世界探索 world model world models distillation
20 Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing 提出GeoMTVR以解决超高分辨率遥感图像的多工具视觉推理问题 reinforcement learning large language model multimodal
21 Parallel Decoding Distillation for Fast Image and Video Generation 提出并行解码蒸馏方法以加速图像和视频生成 flow matching distillation
22 Beyond Background Bias: Saliency-Driven Prototype Alignment for Dataset Distillation 提出基于显著性驱动的原型对齐方法以解决数据集蒸馏问题 distillation
23 Pictura: Perspective-View Self-Play at Scale for Driving 提出Pictura以解决自我对抗训练中的观察差距问题 PPO egocentric

🔬 支柱三:空间感知与语义 (Perception & Semantics) (6 篇)

#题目一句话要点标签🔗
24 Few-Shot Open-Vocabulary Remote Sensing Segmentation via Textual Inversion 通过文本反转实现少样本开放词汇遥感分割 open-vocabulary open vocabulary
25 WHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token Mixing 提出WHTMix以解决高效立体深度估计问题 depth estimation stereo depth
26 PanoLess: Environment Reconstruction from Partial Reflective Views 提出PanoLess以解决部分反射视图环境重建问题 gaussian splatting splatting
27 The LAIA Dataset: Labelled Attention for Intelligent Automobiles 提出LAIA数据集以解决自动驾驶模型可解释性问题 optical flow
28 DensFiLM: Density-Conditioned Video Saliency for Crowd Scenes 提出DensFiLM以解决人群场景视频显著性建模问题 optical flow
29 RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection 提出RDVSv2以解决RGB-D视频显著目标检测数据集不足问题 optical flow

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
30 Schrödinger's Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics 提出GARFIELD以解决场景运动预测中的不确定性问题 motion planning motion generation
31 Explicit Layer Modeling for Video Object Insertion and Layer Decomposition 提出TriLayer与DBL-Diffusion以解决视频对象插入与层分解问题 manipulation

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
32 I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models 提出I2VShield以解决图像到视频模型的主动防御问题 spatiotemporal multimodal

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
33 HOME: Robust Hough-space Matching Method for Structured and Textureless Videos 提出HOME以解决结构化和无纹理视频中的匹配问题 feature matching

⬅️ 返回 cs.CV 首页 · 🏠 返回主页