cs.CV(2026-08-10)

📊 共 39 篇论文 | 🔗 6 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (16 🔗2) 支柱九:具身大模型 (Embodied Foundation Models) (13 🔗4) 支柱三:空间感知与语义 (Perception & Semantics) (4) 支柱一:机器人控制 (Robot Control) (2) 支柱七:动作重定向 (Motion Retargeting) (2) 支柱五:交互与反应 (Interaction & Reaction) (1) 支柱六:视频提取与匹配 (Video Extraction) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (16 篇)

#题目一句话要点标签🔗
1 UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation 提出UniMoFlow以解决3D人类动作编辑中的指令驱动问题 flow matching text-to-motion motion generation
2 World Tokens: Enhancing Embodied Policies with Training-Time World Modeling 提出World Tokens以提升具身策略的训练时世界建模能力 world model world models spatiotemporal
3 Multi-Submap Implicit Neural SLAM with Local-to-Global Loop Closure for Large-Scale Scene Reconstruction 提出多子图隐式神经SLAM以解决大规模场景重建问题 distillation NeRF neural radiance field
4 RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation 提出RecoverFly框架以解决无人机视觉语言导航中的失败问题 reinforcement learning behavior cloning vision-language-action
5 Did the Grid Erase the Event? EndoClock for Auditing Medical World-Model Pipelines 提出EndoClock以解决医疗世界模型同步问题 world model world models PULSE
6 Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning 提出LDR以解决视频生成模型动态建模不足问题 world model world models latent dynamics
7 Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots 提出CVPD框架以实现自我蒸馏解决视觉盲点问题 distillation large language model multimodal
8 ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly Detection 提出ADOPD框架以解决工业异常检测中的参考依赖问题 distillation large language model multimodal
9 Foundation Models are Implicit Deepfake Detectors 提出利用基础模型进行深伪检测的新方法 representation learning foundation model
10 Sekai2: From World Exploration to Interactive World Modeling 提出Sekai2以解决长视频生成与交互建模问题 world model world models
11 UniDFKD: A Unified Semantic Prior Framework for Architecture-Agnostic Data-Free Knowledge Distillation 提出UniDFKD框架以解决数据无关知识蒸馏中的架构依赖问题 teacher-student distillation
12 RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation 提出REST框架以解决高效图像生成中的蒸馏问题 reinforcement learning distillation
13 Marrying Optimal Transport and ODEs for Unified Continuous-Time 4D Reconstruction and Tracking 提出Uni4R框架以解决4D重建与点跟踪问题 flow matching TAMP
14 TeaMatch: Teachable Cross-Modal Representation Learning for 2D-3D Matching 提出TeaMatch以解决2D-3D匹配中的可靠性问题 representation learning
15 CodecArena: Codec Quality Assessment via Visual Reinforcement Learning 提出CodecArena以解决视频编码质量评估问题 reinforcement learning
16 SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs 提出SignLlama以解决无注释手语翻译中的视觉特征优先问题 distillation large language model

🔬 支柱九:具身大模型 (Embodied Foundation Models) (13 篇)

#题目一句话要点标签🔗
17 Multimodal Model Diffing for Feature Discovery and Control 提出MMDiff框架以发现和控制多模态特征 large language model multimodal
18 Visual Distortion Detection in UGC Images Using Large Multimodal Models 提出VIGIL以解决UGC图像中的视觉失真检测问题 large language model multimodal
19 DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning 提出DistMoE以解决分布式视觉指令调优问题 large language model multimodal instruction following
20 LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection 评估基础模型在面部呈现攻击检测中的局限性 foundation model
21 Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework 提出A2I-Set和AudioCanvas以解决音频到图像生成的挑战 multimodal
22 Beyond Uniform Restoration: Empowering All-in-One Restoration with Pixel-Level Multimodal Guidance 提出像素级多模态引导的全能图像修复框架MGN-AIR multimodal
23 Triple Expert Learning from Noisy Labels for Semi-Supervised Vision Foundation Model Adaptation 提出TriNoL框架以解决伪标签噪声对视觉基础模型适应的影响 foundation model
24 Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs 提出动态实例分离子空间以解决LVLM中的幻觉问题 large language model multimodal
25 SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision 提出SI-Edit以解决草图指导的局部图像编辑精度问题 large language model multimodal
26 SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping 提出SwissCrop25以解决作物映射模型的通用性问题 foundation model
27 MemeMind: Reference-Guided Trace Construction for Offline Context Optimization 提出MemeMind以解决离线上下文优化中的失败回滚问题 visual grounding
28 PatchHead: Learning Spatial Patch Evidence for Generalizable AI-Generated Image Detection 提出PatchHead以解决AI生成图像检测的泛化问题 foundation model
29 RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing 提出RAGMesh与FaME-G2E以解决长文本驱动的3D人脸生成与编辑问题 multimodal

🔬 支柱三:空间感知与语义 (Perception & Semantics) (4 篇)

#题目一句话要点标签🔗
30 RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection 提出RefineAny3D以解决单目3D检测中的深度精度问题 open-vocabulary open vocabulary foundation model
31 View-Adaptive Renderer for View-Consistent 2D-to-3D Generation 提出视点自适应渲染器以解决2D到3D生成中的视图一致性问题 3D reconstruction NeRF neural radiance field
32 Right Answer, Wrong Heat: Explanation-Aware Evaluation and Thermal-Grounded Feedback for MLLMs on Infrared Images 提出解释感知评估框架以提升红外图像MLLM的解释能力 scene understanding large language model multimodal
33 Sign Language Recognition Using Original and Synthetic Depth Image Based Point Cloud Data Models 提出基于深度图像的点云数据模型以提升手语识别性能 Depth Anything

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
34 FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search 提出FaLCon以解决文本驱动的行人异常搜索问题 sim2real large language model multimodal
35 Overcoming Data Scarcity and Confidentiality in Hardware Assurance via Synthetic Generation 提出隐私保护管道以解决硬件保障中的数据稀缺与机密性问题 sim-to-real

🔬 支柱七:动作重定向 (Motion Retargeting) (2 篇)

#题目一句话要点标签🔗
36 Towards Collaborative Joint Perception and Prediction: Framework, Baseline Evaluation, and Deployment Perspectives 提出协作联合感知与预测框架以解决感知误差和视觉遮挡问题 motion prediction
37 Illusion or Integrity? Geometrical Consistency Metric for AIGC Video Quality Evaluation 提出GeoCon-Bench以解决AIGC视频质量评估中物理一致性不足的问题 geometric consistency

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
38 Efficient Human-Contact Representation for Human-Scene Interaction 提出稀疏接触表示以解决人类场景交互效率问题 human-scene interaction

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
39 EgoHieraLoc: A Cortically Inspired Hierarchical Segmentation-Guided Framework for Egocentric Visual Query Localization 提出EgoHieraLoc以解决模糊边界下的视觉查询定位问题 egocentric

⬅️ 返回 cs.CV 首页 · 🏠 返回主页