cs.CV(2026-08-26)

📊 共 33 篇论文 | 🔗 8 篇有代码

🎯 兴趣领域导航

支柱二:RL算法与架构 (RL & Architecture) (14 🔗2) 支柱九:具身大模型 (Embodied Foundation Models) (7 🔗2) 支柱三:空间感知与语义 (Perception & Semantics) (6 🔗2) 支柱一:机器人控制 (Robot Control) (3) 支柱四:生成式动作 (Generative Motion) (1 🔗1) 支柱七:动作重定向 (Motion Retargeting) (1) 支柱六:视频提取与匹配 (Video Extraction) (1 🔗1)

🔬 支柱二:RL算法与架构 (RL & Architecture) (14 篇)

#题目一句话要点标签🔗
1 4DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting 提出4DGS-WAM以解决传统WAM在空间结构建模中的不足 world model world models world action model
2 Code World Model: Coding Agent as World Brain 提出Code World Model以解决视频世界模型的局限性问题 world model world models spatiotemporal
3 V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning 提出V-Rubrics以解决视觉证据不足的问题 reinforcement learning multimodal instruction following
4 MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations 提出MLLMCLIP以解决视觉语言模型的组合性问题 distillation large language model multimodal
5 AdaptiveEmbed: Sample-Adaptive Multi-Vector Representation for Multimodal Retrieval 提出Sample-Adaptive Multi-Vector Representation以解决多模态检索中的固定容量问题 contrastive learning multimodal
6 CloSeR: Unified Relational Distillation from Closed-Set Teachers for Category Discovery 提出CloSeR框架以解决类别发现中的知识蒸馏问题 distillation foundation model
7 CrossMambaTuning: Synergistic Spatial and Cross-Layer Adaptation for Machine Vision Compression 提出CrossMambaTuning以解决机器视觉压缩中的跨层适应问题 Mamba state space model
8 4DStreamCtrl: Interactive Video Generation with Online 4D Control 提出4DStreamCtrl以实现实时4D视频生成与控制 world model world models spatiotemporal
9 Embedding NDRE Trajectories into Contrastive Learning for Label-Free, Physiology-Aware Crop-Stress Staging and DSS Outputs 提出EigenCL以解决作物压力检测不足问题 contrastive learning
10 Socialized Detector Learning: Trajectory-Guided and Reciprocal Distillation for Heterogeneous Object Detectors 提出社会化检测器学习以解决异构目标检测器知识碎片化问题 distillation
11 DeCO: Discriminative Evidence Composition for Fine-Grained Dataset Distillation 提出DeCO以解决细粒度数据集蒸馏中的局部证据缺失问题 distillation
12 Where to Look Matters: On-Policy Self-Distillation for Long-Video Understanding 提出Clue-OPSD框架以提升长视频理解精度 distillation
13 VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 提出VBVR-Pro以解决可扩展视觉推理训练问题 reinforcement learning spatiotemporal
14 Asymmetric Cross-Modal Fine-Grained Visual Categorization: ACF-Net and the BirdPro Benchmark 提出ACF-Net以解决不对称跨模态细粒度视觉分类问题 representation learning optical flow

🔬 支柱九:具身大模型 (Embodied Foundation Models) (7 篇)

#题目一句话要点标签🔗
15 Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios 提出Video-IFBench以解决视频理解中的指令遵循问题 large language model multimodal instruction following
16 A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training 提出视觉依赖感知框架以解决多模态无监督持续后训练问题 multimodal
17 UltraPIPS: Improving model perception in B-mode ultrasound with foundation models 提出UltraPIPS以解决B超图像感知模型不足问题 foundation model
18 PANDA - Prototype-Anchored Alignment for Partially Unpaired Multimodal Learning, with Applications to Alzheimers MRI and TCGA Pathology 提出PANDA框架以解决多模态医学预测中的不完全配对问题 multimodal
19 Precipitation Downscaling Using Foundation Model-Conditioned Diffusion 提出基于扩散模型的降水下分辨率方法以解决气候模型偏差问题 foundation model
20 RSFusionDet: Underwater RGB-Sonar Multimodal Object Detection 提出RSFusionDet以解决水下多模态物体检测问题 multimodal
21 Not All Attention Heads Contribute to Critical Visual Token Selection: Head-Aware Pruning Matters More 提出ProViP以解决视觉语言模型推理效率问题 large language model

🔬 支柱三:空间感知与语义 (Perception & Semantics) (6 篇)

#题目一句话要点标签🔗
22 PIVOT: A Multi-Trajectory Dataset and Testbed for Pose, Intrinsics, and Novel Viewpoint Evaluation in Real-World 3D Reconstruction 提出PIVOT以解决真实场景3D重建评估问题 3D gaussian splatting 3DGS 3D reconstruction
23 Gaussian Splatting Underwater: A Controlled Cross-Regime Study 研究水下环境中的高斯散射以提升3D重建效果 3DGS 3D reconstruction gaussian splatting
24 PAGS: Autofocusing Photoacoustic Tomography via Speed-of-Sound-Adaptive Gaussian Splatting 提出PAGS以解决光声计算断层成像中的声速异质性问题 gaussian splatting splatting
25 Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming 提出PoseOFF以解决低延迟人类动作预测问题 optical flow motion representation
26 Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models 提出LiDAR-SAM2以解决4D LiDAR标注数据不足问题 scene understanding foundation model
27 TAU-Agent: An Agentic Retrieval-Augmented Framework for Traffic Anomaly Understanding 提出TAU-Agent以解决交通异常理解问题 open-vocabulary open vocabulary

🔬 支柱一:机器人控制 (Robot Control) (3 篇)

#题目一句话要点标签🔗
28 StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models 提出StreamPI以解决VLA模型在时间建模中的局限性 manipulation vision-language-action VLA
29 V-Link: Recovering Lost Visual Representations in Action DiT for Vision-Language-Action Models 提出V-Link以解决VLA模型中视觉表示丢失问题 humanoid manipulation vision-language-action
30 PointRL: Learning Point-Level Vision-Language Grounding from Verifiable Annotation Evidence 提出PointRL框架以解决视觉语言指向行为学习问题 manipulation reinforcement learning visual grounding

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
31 InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture Control 提出InteractGesture以解决共语手势生成中的空间控制问题 motion latent VQ-VAE

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
32 TDFNet: Tri-projection Deformable Fusion Network for Panoramic Salient Object Detection 提出TDFNet以解决全景显著目标检测中的几何失真问题 geometric consistency

🔬 支柱六:视频提取与匹配 (Video Extraction) (1 篇)

#题目一句话要点标签🔗
33 Moving Beyond More Views: Redundancy-Aware Ego-Exo Fusion for Proficiency Estimation 提出冗余感知的自我-外部融合方法以提升能力评估 egocentric

⬅️ 返回 cs.CV 首页 · 🏠 返回主页