cs.CV(2026-07-30)

📊 共 53 篇论文 | 🔗 9 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (18 🔗5) 支柱二:RL算法与架构 (RL & Architecture) (14 🔗1) 支柱三:空间感知与语义 (Perception & Semantics) (12 🔗2) 支柱一:机器人控制 (Robot Control) (4 🔗1) 支柱四:生成式动作 (Generative Motion) (2) 支柱五:交互与反应 (Interaction & Reaction) (1) 支柱七:动作重定向 (Motion Retargeting) (1) 支柱八:物理动画 (Physics-based Animation) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (18 篇)

#题目一句话要点标签🔗
1 Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models 提出趋势感知修剪以解决多模态大语言模型的视觉令牌过滤问题 large language model multimodal
2 MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models 提出MMOOC基准以解决多模态大语言模型的上下文评估问题 large language model multimodal
3 Witness Evidence Portfolios: Single-Prefill Risk Detection for Closed Multimodal Answers 提出见证证据组合以解决闭合多模态答案的风险检测问题 large language model multimodal
4 MIND: Multimodal Intent-Driven Network via Diffusion Transformers for Medical Image Fusion 提出MIND以解决医学图像融合中的意图驱动问题 multimodal
5 Negative controls reveal volume-driven confounding in radiomics and imaging foundation model features 提出READII-2-ROQC以解决影像组学中的体积驱动混淆问题 foundation model
6 Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures 提出病理基础模型作为有丝分裂图像检测编码器 foundation model
7 Benchmarking Foundation and Large Language Models for Few-Shot Medical Image Segmentation 提出FAME基准以评估少样本医学图像分割方法 large language model
8 DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis 提出多样化架构以解决医学图像分析中的概念检测与标题生成问题 foundation model
9 FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval 提出FiRE以解决复杂图像检索中的细粒度上下文建模问题 large language model multimodal
10 LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference 提出LAST框架以解决边缘云协作MLLM推理中的视觉令牌修剪问题 foundation model multimodal
11 Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA 提出Thinking-Once以解决高分辨率视觉问答中的证据获取问题 large language model multimodal
12 LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA 提出LoMeVQA以解决纵向医学视觉问答问题 large language model multimodal
13 ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation 提出ROAD框架以降低3D形状生成的训练成本 foundation model
14 ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding 提出ObjectStream以解决流媒体视频理解中的记忆管理问题 large language model
15 Scaling Vision-Language Models Is Not Enough to Mitigate Bias 大规模研究揭示视觉-语言模型偏见问题的复杂性 multimodal
16 TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment 提出几何锚定的舌头合成方法以解决面部重演中的舌头动态问题 foundation model
17 MeshFM: 2D Features Are All You Need for 3D Shape Understanding 提出MeshFM以解决3D形状理解中的2D特征利用问题 foundation model
18 Drawing-Recode: Annotation Grounding for Parametric CAD Code Generation from Raster 2D CAD Drawings 提出Drawing-Recode以解决光栅2D CAD图纸的参数化CAD代码生成问题 large language model

🔬 支柱二:RL算法与架构 (RL & Architecture) (14 篇)

#题目一句话要点标签🔗
19 OPLD: On-Policy Latent Distillation for Multimodal Reasoning 提出OPLD以解决多模态推理中的抽象思维问题 distillation multimodal chain-of-thought
20 VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation 提出视觉归因蒸馏以解决多模态知识转移问题 distillation multimodal
21 Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation 提出一种新方法以构建和验证灾害响应多模态数据集 distillation multimodal
22 AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans 提出基于世界模型的分割框架以解决耳部CT图像细粒度分割问题 world model world models latent dynamics
23 RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation 提出RefineSVG以解决图像到SVG生成中的几何漂移问题 reinforcement learning large language model multimodal
24 Learning to Understand Body Language from Flight through Robust 3D Avatar Placing 提出Drones2BodyLanguage数据集以解决无人机社交智能问题 world model world models monocular depth
25 PhiZero: A World Model Built Around Physical Language 提出PhiZero以解决物理世界建模中的隐式动态问题 world model world models
26 ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow 提出ShadowDancer以解决视频世界模型中的动作控制问题 world model world models
27 Beacon: Knowing When and How to Perform Agentic Visual Reasoning 提出Beacon模型以提升多模态大语言模型的视觉推理能力 reinforcement learning large language model multimodal
28 One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting 提出单补丁文本检测方法以解决场景文本定位精度问题 reinforcement learning large language model multimodal
29 SPFM-Net: Semantic-Prior-Guided Frequency-Constrained Mamba for Invisible Watermark Attack 提出SPFM-Net以解决隐形水印攻击效果与视觉保真度的权衡问题 Mamba masked autoencoder
30 ViP-Rig: Visual-Prompted Controllable Rigging 提出ViP-Rig以解决动画任务中骨骼控制不足的问题 VIP
31 Learning Color Grading, No Photo Sharing: Federated Aesthetic Preference Learning for Personalized Image Enhancement 提出FedPAIE框架以解决个性化图像增强中的隐私问题 preference learning
32 Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion 提出集中隐式偏好优化以解决视频生成中的时间稀疏伪影问题 DPO direct preference optimization

🔬 支柱三:空间感知与语义 (Perception & Semantics) (12 篇)

#题目一句话要点标签🔗
33 Beyond Visual Ambiguity: Guiding Robust Monocular Depth Estimation in Challenging Scenarios via Detailed Long Captions 提出CapDepth以解决单目深度估计中的视觉歧义问题 depth estimation monocular depth spatial relationship
34 ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA 提出ViewMind3D以解决训练依赖的3D问答问题 3D reconstruction embodied AI large language model
35 MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians 提出MonoVoc以解决轻量级单目开放词汇3D场景理解问题 scene understanding open-vocabulary open vocabulary
36 Endo-NeRF++: Uncertainty-Aware Neural Rendering with Multi-Resolution Hash Encoding for Dynamic Surgical Scene Reconstruction 提出Endo-NeRF++以解决动态外科场景重建中的不确定性问题 NeRF scene reconstruction
37 4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans 提出4DHumanDiff以解决360度动态人类生成问题 gaussian splatting splatting scene reconstruction
38 S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image 提出S-Avatar以解决单图生成3D头像一致性问题 3D gaussian splatting 3DGS gaussian splatting
39 AdaAnchor4D: Anchor-Conditioned Spatiotemporal Feature Aggregation for Monocular UAV 4D Reconstruction 提出AdaAnchor4D以解决单目无人机动态场景重建问题 scene reconstruction spatiotemporal
40 Convolutional Neural Shading for High-Quality 3D Reconstruction from Multi-View Images 提出卷积神经着色以解决多视图图像高质量3D重建问题 3D reconstruction neural radiance field
41 Filling the Pareto-Optimal Front for Affordance Segmentation on Embedded Devices Using RGB-D Cameras 提出深度信息融合方法以优化可穿戴机器人中的可供性分割 affordance
42 Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars 提出SpiD框架以解决单图像生成头部头像问题 3D gaussian splatting gaussian splatting splatting
43 TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting 提出TARS以解决3D视频重拍中的视角控制问题 3D reconstruction spatiotemporal
44 TSOG: A Format For Temporally And Spatially Ordered Gaussians 提出TSOG格式以高效表示动态4D高斯内容 gaussian splatting splatting

🔬 支柱一:机器人控制 (Robot Control) (4 篇)

#题目一句话要点标签🔗
45 ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine 提出ACE-Data-0以解决人类中心数据捕获的瓶颈问题 locomotion manipulation dexterous manipulation
46 EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE 提出EgoGenesis以解决稀缺的自我中心数据问题 manipulation dual-arm egocentric
47 FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification 提出FaithEyes以解决工具使用不可靠问题 manipulation multimodal
48 Can Vision-Language Models Reason about AI Edits in Images? 提出基于强化学习的VLM框架以检测AI篡改图像 manipulation reinforcement learning

🔬 支柱四:生成式动作 (Generative Motion) (2 篇)

#题目一句话要点标签🔗
49 MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing 提出MPIE-Bench以解决多人物交互编辑中的解剖学问题 penetration multi-person interaction
50 Articulated Object Reconstruction from Rest-State Observation 提出静态状态重建方法以解决关节物体重建问题 physically plausible geometric consistency

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
51 Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer 提出系统性框架以提升手-物交互建模能力 HOI human-to-robot foundation model

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
52 ENCORE: Event-Assisted Complementary Motion Refinement for Learned Video Compression 提出ENCORE框架以解决学习视频压缩中的运动建模问题 motion representation TAMP

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
53 Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers 提出Chimera以解决高分辨率图像和长视频生成问题 spatiotemporal multimodal

⬅️ 返回 cs.CV 首页 · 🏠 返回主页