cs.CV(2026-08-21)

📊 共 43 篇论文 | 🔗 10 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (19 🔗3) 支柱二:RL算法与架构 (RL & Architecture) (10 🔗2) 支柱三:空间感知与语义 (Perception & Semantics) (8 🔗4) 支柱四:生成式动作 (Generative Motion) (2 🔗1) 支柱七:动作重定向 (Motion Retargeting) (1) 支柱五:交互与反应 (Interaction & Reaction) (1) 支柱一:机器人控制 (Robot Control) (1) 支柱六:视频提取与匹配 (Video Extraction) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (19 篇)

#题目一句话要点标签🔗
1 Just Noticeable Difference Modeling for Token Compression in Vision-Language-Action Models 提出Action-JND以优化视觉语言行动模型中的令牌压缩 vision-language-action VLA foundation model
2 A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving 提出基于VLA的多模态协作交互以解决自主驾驶决策不可靠问题 vision-language-action VLA multimodal
3 Recognition-Conditioned Reasoning: A Training-Free Multimodal-LLM Pipeline for Fine-Grained Micro-Action Understanding 提出无训练的多模态LLM管道以解决微动作理解问题 large language model multimodal
4 EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking 提出EviRank以解决多模态图像重排序问题 multimodal chain-of-thought
5 CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models 提出CertVLA以解决视觉-语言-动作模型的物理攻击问题 vision-language-action VLA
6 Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation 提出Vis-Poison以解决多模态检索增强生成中的视觉知识中毒问题 large language model multimodal
7 Masking Is Not Enough: Generative Restoration for Multimodal De-Identification in Medical AI 提出ClinX框架以解决医疗AI中的多模态去标识化问题 multimodal
8 Latent Ordinal Evidence, Misaligned Outputs: Inference-Time Ordinal Lens Alignment for Multimodal LLMs 提出Ordinal Lens Alignment以解决多模态LLMs的输出对齐问题 multimodal
9 Kinematic Knowledge Maps for Pattern Alignment: Structured Latent Representational Learning in Multimodal Gait Analysis 提出ScoliDetect以解决多模态步态分析中的对齐问题 multimodal
10 SuppreSensing: Expert-Guided Feature Recalibration and Discrepancy Augmentation for Multimodal Object Detection 提出SuppreSensing以解决多模态遥感目标检测中的语义异质性问题 multimodal
11 EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue 提出EmotionDialogCN以解决多模态对话数据集情感标注不足问题 multimodal
12 MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model 提出MV2GF以解决多视角行人检测中的几何泛化问题 foundation model
13 Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds 提出轻量级视觉支架以提升视觉语言模型空间推理能力 multimodal visual grounding
14 MigrationNarrate: A Dataset for Detection of Migration Narratives in YouTube Videos 提出MigrationNarrate数据集以解决移民叙事检测问题 large language model multimodal
15 OccluRank: Controllable Occlusion-Aware Layout-to-Image Generation by Adding Just an Ordinal Rank 提出OccluRank以解决布局到图像生成中的遮挡顺序问题 large language model multimodal
16 AGIDefect-4K: A Richly Annotated Dataset for AI-Generated Image Defect Detection, Localization and Explanation 提出AGIDefect-4K数据集以解决AI生成图像缺陷检测问题 large language model multimodal
17 OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs 提出OmniAssistBench以解决交互式视频助手评估问题 large language model
18 KoViDoRe: Korean Visual Document Retrieval 提出KoViDoRe以解决韩文视觉文档检索问题 multimodal
19 Identity-Preserving Text-to-Video Generation via Agentic Enhancement and Semantic Repair 提出AESR框架以解决身份保留视频生成中的问题 instruction following

🔬 支柱二:RL算法与架构 (RL & Architecture) (10 篇)

#题目一句话要点标签🔗
20 COMET: Contrastive Motion-Enhanced Temporal Reasoning for Video Multimodal Large Language Models 提出COMET以解决视频多模态大语言模型的运动时间理解问题 distillation large language model multimodal
21 Semantically Compatible Knowledge Distillation for Cross-Domain Object Detection with Vision Foundation Models 提出语义兼容知识蒸馏框架以解决跨域目标检测问题 teacher-student distillation foundation model
22 WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving 提出WA-JEPA以解决自主驾驶中的未来预测问题 flow matching JEPA Joint-Embedding Predictive Architecture
23 CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents 提出CIVA以解决视觉世界模型代理的攻击问题 world model world models dreamer
24 Enhancing Localized Reasoning for Long Video Understanding via Efficient Segment-to-Video Supervision 提出Segment-to-Video监督以提升长视频理解能力 reinforcement learning reward design large language model
25 Difficulty-Calibrated Interpolation Paths for Conditional Flow Matching 提出难度校准插值路径以优化条件流匹配 flow matching classifier-free guidance
26 A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration 提出A2DINOv3以解决多模态目标检测中的信息冗余问题 representation learning scene understanding foundation model
27 Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning 提出Re$^3$Cap以解决图像描述中的推理不足问题 reinforcement learning
28 Anchoring Instruction Outside Mask: Exact Reference Caching for Efficient In-Context Diffusion Transformers 提出静态文本锚点以解决多参考图像生成效率问题 distillation instruction following
29 DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion 提出DiGS-Avatar以解决单图像3D人类重建问题 teacher-student 3D reconstruction

🔬 支柱三:空间感知与语义 (Perception & Semantics) (8 篇)

#题目一句话要点标签🔗
30 Stream3Dv2: Geometric-Semantic Fusion Enhanced Streaming Zero-Shot 3D Scene Understanding 提出Stream3Dv2以解决流媒体RGB-D输入处理和噪声分割问题 scene understanding open-vocabulary open vocabulary
31 Lift, Associate, and Fuse: A Decision-Centric Framework for 2D-to-3D Foundation Model Transfer 提出LAF框架以解决2D到3D模型转移中的决策问题 open-vocabulary open vocabulary foundation model
32 M2Depth: Unifying Monocular Depth Foundation Priors with Multi-View Stereo 提出M2Depth以解决多视图立体视觉中的深度预测问题 monocular depth geometric consistency foundation model
33 TopoSurfel: Closing the Loop between Gaussian Surfels and Meshes for Surface Reconstruction 提出TopoSurfel以解决高保真表面重建问题 3D gaussian splatting 3DGS gaussian splatting
34 Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer 提出MVAP-G以解决VGGT模型的安全漏洞问题 3D reconstruction VGGT foundation model
35 MotionPhys: Detecting AI-Generated Videos via Physical Consistency of Optical-Flow Trajectories 提出MotionPhys以解决AI生成视频物理一致性检测问题 optical flow
36 AffordAny: Open-World 3D Affordance Grounding from Monocular RGB Images via Vision-Language-Guided Geometric Reasoning 提出AffordAny以解决开放世界3D功能定位问题 affordance
37 VisTa3D: A Dataset and Benchmark for Thin Object Reconstruction from Vision, Tactile, and 3D Point Clouds 提出VisTa3D数据集以解决薄物体重建问题 3D reconstruction

🔬 支柱四:生成式动作 (Generative Motion) (2 篇)

#题目一句话要点标签🔗
38 ArtiMo: Agent-Driven Articulated Mesh Animation 提出ArtiMo以解决文本驱动的关节网格动画问题 motion generation
39 Robust Validation to Geometric Perturbations for Autonomous Pose Estimation 提出全局利普希茨优化以解决自主姿态估计的鲁棒性问题 physically plausible

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
40 Gaussian-Mixture Latent Flow for Stochastic 3D Human Motion Prediction 提出高斯混合潜在流模型以解决人类运动预测中的不确定性问题 human motion human motion prediction motion prediction

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
41 Identity-Aware Human-Object Interaction Motion Captioning 提出身份感知的人物-物体交互动作字幕生成方法以解决身份缺失问题 human-object interaction HOI

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
42 GAP-SAM: A Global Artifact Prior for Generalizable AI-Generated Image Manipulation Localization 提出GAP-SAM以解决AI生成图像编辑定位的OOD性能问题 manipulation

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
43 Dorsal Hand Images for Immersive (XR) and Privacy-preserving Age Assurance and Child Safety 提出手背图像以解决XR环境中的隐私保护年龄验证问题 egocentric

⬅️ 返回 cs.CV 首页 · 🏠 返回主页