cs.CV(2026-09-01)

📊 共 38 篇论文 | 🔗 9 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (15 🔗4) 支柱二:RL算法与架构 (RL & Architecture) (9 🔗2) 支柱三:空间感知与语义 (Perception & Semantics) (6 🔗3) 支柱一:机器人控制 (Robot Control) (3) 支柱六:视频提取与匹配 (Video Extraction) (2) 支柱七:动作重定向 (Motion Retargeting) (2) 支柱四:生成式动作 (Generative Motion) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (15 篇)

#题目一句话要点标签🔗
1 S$^2$Prune: Spatially Structured Visual Token Pruning for Multimodal Large Language Models 提出S$^2$Prune以解决多模态大语言模型的视觉标记剪枝问题 large language model multimodal
2 SinkPruner: Sink-Free Visual Token Pruning for Multimodal Large Language Models 提出SinkPruner以解决多模态大语言模型的视觉标记冗余问题 large language model multimodal
3 Forbid Your Attention: Fooling Multimodal Large Language Models by Selectively Removing Intrinsic Focus in Spectral Domain 提出相位感知对抗攻击框架以增强多模态大语言模型的鲁棒性 large language model multimodal
4 Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System 提出任务解耦架构以提升统一多模态模型的理解与生成协同能力 multimodal
5 Semantic-Guided Multimodal Preprocessing for Vision Transformer-Based Clear Cell Renal Cell Carcinoma Grading 提出语义引导的多模态预处理方法以解决CCRCC分级问题 multimodal
6 Multimodal RGB-Infrared Combination for UAV-Based Wildfire Segmentation: A Comparative Study on FLAME3 提出RGB-红外融合方法以提升无人机野火分割精度 multimodal
7 CMRVision: A Foundation Model for Cardiac MR Image Analysis 提出CMRVision以解决心脏磁共振图像分析问题 foundation model
8 Agentic Multimodal Models for Environmental Hyperspectral Unmixing 提出一种基于多模态模型的框架以改进环境高光谱解混问题 multimodal
9 FTU-Seek: Foundation Model-Guided Hard-Negative Learning for Sparse Functional Tissue Unit Segmentation 提出FTU-Seek以解决稀疏功能性组织单元分割问题 foundation model
10 Do Satellites See Commuters? A Critical Benchmark of Vision Foundation Models 比较四种卫星视觉编码器以优化通勤OD生成 foundation model
11 BrainDiff: Longitudinal Report Generation for Multimodal Brain MRI 提出BrainDiff以解决多模态脑MRI的纵向报告生成问题 multimodal
12 IntroConformal: Conformal Factuality Guarantees for Large Vision-Language Models via Introspective Signals 提出IntroConformal以解决大规模视觉语言模型的事实准确性问题 multimodal
13 One Prompt Is Enough: Watermark Laundering Through Foundation Image Models 提出水印洗涤方法以解决基础图像模型的隐形水印问题 foundation model
14 ADGNet: Asymmetric Dual-text Guided Network for Infrared Small Target Detection 提出ADGNet以解决红外小目标检测中的背景干扰问题 multimodal
15 Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures 提出无模型渲染上限以解决视觉语言模型的感知与推理分离问题 multimodal

🔬 支柱二:RL算法与架构 (RL & Architecture) (9 篇)

#题目一句话要点标签🔗
16 What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models 通过层级分析揭示视频基础模型的时空表示特征 JEPA spatiotemporal foundation model
17 Streaming4D: Accelerate 4D World Models via Block-wise Video Generation and Incremental Reconstruction 提出Streaming4D以解决4D生成中的高延迟问题 world model world models 3D reconstruction
18 Panda Diplomacy: Foundation Model Pre-training across Particle Imaging Detectors for High Energy and Nuclear Physics 提出点云自蒸馏框架以提升粒子成像探测器的预训练效果 distillation foundation model
19 On the Design Fundamentals of Pixel Text Representation Learning 提出Pixel Linguist II以解决视觉文本表示学习的多项挑战 representation learning visual grounding
20 ReFlowSET: Representation-Aligned Latent Flow Matching for SAR-to-EO Image Translation 提出ReFlowSET以解决SAR到EO图像翻译中的编码器选择问题 flow matching foundation model
21 ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives 提出ViTAMINS以提升自监督视觉变换器的表示质量 JEPA contrastive learning distillation
22 Solaris: Towards Interfaces That Are Generated, Not Coded 提出Solaris以生成动态用户界面,解决传统编码限制 world model world models distillation
23 Feed-Forward Multi-view Multi-person Reconstruction with Contrastive Human-Aware 3D Representation 提出一种新方法以解决复杂环境下的多视角多人体重建问题 contrastive learning SMPL
24 H3-World: Turning Language Understanding into World Control 提出H3-World以实现语言理解驱动的世界控制 world model world models

🔬 支柱三:空间感知与语义 (Perception & Semantics) (6 篇)

#题目一句话要点标签🔗
25 Seeing the World and the Self from Egocentric Video 提出RESELF框架以解决自我中心视频的3D感知问题 depth estimation scene reconstruction motion generation
26 DualDiff3D: Dual Structure-Appearance Diffusion Priors for Reliability-Enhanced 3D Gaussian Splatting 提出DualDiff3D以解决3D重建中视角不足导致的质量问题 3D gaussian splatting 3DGS 3D reconstruction
27 Monocular Depth Estimation from a Single Image: Progress and Opportunities 综述单目深度估计的进展与未来机遇 visual SLAM depth estimation monocular depth
28 EvoGS: Modeling Deformation Evolution for Dynamic Gaussian Splatting 提出EvoGS以解决动态场景中高效重建问题 3D gaussian splatting 3DGS gaussian splatting
29 VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM 提出VOIM以解决训练无关的开放词汇3D实例映射问题 open-vocabulary open vocabulary occupancy grid
30 On-the-Fly3R: Towards Robust Online 3D Reconstruction with Feed-Forward 3R Models for Large-Scale UAV Scenarios 提出On-the-Fly3R以解决大规模无人机场景下的3D重建问题 3D reconstruction

🔬 支柱一:机器人控制 (Robot Control) (3 篇)

#题目一句话要点标签🔗
31 Design and Implementation of a Kalman Filter-Infused Algorithm for Tilt Estimation 提出基于卡尔曼滤波的倾斜角度估计算法以解决传感器噪声问题 motion tracking
32 HELIOS: From midnight to noon, continuous outdoor urban scene relighting 提出HELIOS以解决户外场景照明转换问题 manipulation distillation
33 CameraEditor: Camera-Controlled Image Editing via Video-Prior Sequential Modeling 提出CameraEditor以解决图像编辑中的相机参数控制问题 manipulation

🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)

#题目一句话要点标签🔗
34 CQF-HMR: Continuous Quaternion Flows for Probabilistic 3D Human Mesh Recovery from a Single Image 提出连续四元数流以解决单图像3D人类网格恢复问题 human mesh recovery HMR SMPL
35 TempCloze: Can Video-LLMs Identify the Missing Middle? 提出TempCloze以解决视频语言模型的时间推理问题 egocentric

🔬 支柱七:动作重定向 (Motion Retargeting) (2 篇)

#题目一句话要点标签🔗
36 Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models 提出Dyn-3D以解决视觉语言模型中的自我运动模糊问题 motion estimation
37 Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking 提出PLANET以解决单目视频多目标跟踪中的深度信息缺失问题 spatial relationship

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
38 Physically Plausible Video Generation via Visual-Semantic Chain-of-Events Conditioning 提出基于事件链的物理可信视频生成方法以解决自然语言条件不足问题 classifier-free guidance physically plausible chain-of-thought

⬅️ 返回 cs.CV 首页 · 🏠 返回主页