cs.CV(2026-07-29)

📊 共 27 篇论文 | 🔗 3 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (9 🔗1) 支柱九:具身大模型 (Embodied Foundation Models) (9) 支柱三:空间感知与语义 (Perception & Semantics) (5 🔗1) 支柱一:机器人控制 (Robot Control) (2 🔗1) 支柱四:生成式动作 (Generative Motion) (1) 支柱六:视频提取与匹配 (Video Extraction) (1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (9 篇)

#题目一句话要点标签🔗
1 SpatialQ: Understanding 3D Gaussian Splatting Scene Quality via Visual-based MLLM 提出多模态质量评估框架以解决3D Gaussian Splatting场景质量评估问题 representation learning 3D gaussian splatting 3DGS
2 JEPADepth: Masked Predictive Representation Learning for Self-Supervised Monocular Depth Estimation 提出JEPADepth以解决自监督单目深度估计问题 JEPA Joint-Embedding Predictive Architecture joint-embedding predictive architecture
3 StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation 提出StatePlay以解决游戏世界模型中状态一致性问题 world model world models
4 Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection 提出Veritas++以解决AIGI检测中的感知瓶颈问题 distillation large language model
5 SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation 提出SCALPEL以解决医学多模态表示学习中的对齐问题 representation learning large language model multimodal
6 R-SLPR: Region-based Small-to-Large Point-cloud Registration with Contrastive Learning 提出R-SLPR以解决小规模点云与大规模点云配准问题 MAE contrastive learning
7 SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence 提出SciFigAlign以解决科学图形评估问题 MAE multimodal
8 DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation 提出DistillAlign以解决自回归视频蒸馏中的分布对齐问题 distillation
9 Long-Tailed 3D Point Cloud Dataset Distillation 提出长尾3D点云数据集蒸馏方法以解决数据不平衡问题 distillation

🔬 支柱九:具身大模型 (Embodied Foundation Models) (9 篇)

#题目一句话要点标签🔗
10 Progressive Multimodal Alignment for Continual Instruction Tuning 提出渐进式多模态对齐以解决持续指令调优中的遗忘问题 large language model multimodal
11 See2Think: Do Multimodal Models Really Use Intermediate Visual States? 提出See2Think框架以评估多模态模型对视觉状态的依赖性 large language model multimodal
12 Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution 提出解耦视觉处理框架以提高多模态适应效率 large language model multimodal
13 Anatomy Contextualized Adaption of CT Foundation Models 提出解剖上下文适应框架以提升CT基础模型的解剖对齐能力 foundation model
14 Visual Credit Audit for Multimodal Spatial Reasoning 提出视觉信用审计方法以解决多模态空间推理问题 multimodal
15 Multimodal fusion of visual and morphometric features for avian bone classification 提出多模态框架以解决鸟类骨骼分类问题 multimodal
16 Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering 提出交叉分支语义引导以解决统一多模态模型的语义对齐问题 multimodal
17 Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach 提出CDPI方法以解析异构多模态大语言模型融合中的知识转移机制 large language model multimodal
18 Prior Directions: Why GUI Grounding Gets Locked in the Past 提出Prior Directions以解决视觉语言模型中的锁定问题 visual grounding

🔬 支柱三:空间感知与语义 (Perception & Semantics) (5 篇)

#题目一句话要点标签🔗
19 Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications 提出ByDeWay-V2以解决多模态LLM空间推理不足问题 depth estimation monocular depth open-vocabulary
20 3DGBGS: 3D Granular Ball Gaussian Splatting for Compact Novel View Synthesis 提出3DGBGS以解决现有视图合成方法的适应性不足问题 3DGS gaussian splatting splatting
21 Physically Real-time Infrared Attack against Optical Flow Estimation Networks 提出物理实时红外攻击以增强光流估计网络的鲁棒性 optical flow
22 Genie Sim PanoWorld: An Infinite Indoor 3D World Generation Pipeline via Panoramic Scene Modeling and Simulation 提出Genie Sim PanoWorld以解决单视角全景图生成3D场景的问题 3D reconstruction embodied AI
23 VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion 提出VidMap以解决视频中结构重建的鲁棒性问题 monocular depth scene understanding

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
24 TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM 提出TurboVLA以解决现有VLA模型的计算和内存开销问题 manipulation vision-language-action VLA
25 ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection 提出ScratchSim以解决工业表面划痕检测数据稀缺问题 domain randomization

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
26 TPD: Temporal Prior Decoupling for Text-to-Video Diffusion Models 提出TPD以解决文本到视频生成中的时间优先抑制问题 classifier-free guidance

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
27 HumanCLAW: Can Vision-Language Models Act Through a Body? 提出HumanCLAW框架以评估视觉语言模型的身体动作能力 egocentric

⬅️ 返回 cs.CV 首页 · 🏠 返回主页