cs.CV(2026-08-24)

📊 共 35 篇论文 | 🔗 6 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (12 🔗2) 支柱二:RL算法与架构 (RL & Architecture) (10 🔗2) 支柱三:空间感知与语义 (Perception & Semantics) (8) 支柱一:机器人控制 (Robot Control) (2 🔗1) 支柱六:视频提取与匹配 (Video Extraction) (2) 支柱四:生成式动作 (Generative Motion) (1 🔗1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (12 篇)

#题目一句话要点标签🔗
1 Training-Free Pseudo-Fusion for Composed Image Retrieval with Diffusion Models and Multimodal Large Language Models 提出PeFuse以解决无训练的复合图像检索问题 large language model multimodal
2 Action-Aligned Retrieval with Pairwise Multimodal Reranking for Text-Based Person Anomaly Search 提出ActPair以解决文本驱动的人物异常搜索问题 large language model multimodal
3 Dual-Grained Agent Memory and Shapley Context Attribution for Multimodal Agentic Learner 提出DG-Mem框架以增强多模态学习模型的推理能力 large language model multimodal
4 FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding 提出FOVEA以解决多模态解码中的视觉证据适应问题 multimodal visual grounding
5 DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts 提出DF-MoE以解决深度伪造检测的泛化问题 multimodal
6 MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge 提出无训练框架以解决音频引导的视频目标分割问题 large language model foundation model multimodal
7 Toward a Foundation Plug-and-Play Prior for Computed Tomography Reconstruction via a Multimodal Diffusion Model 提出多模态扩散模型以解决CT重建中的先验问题 multimodal
8 ByteAction: Byte-space Action Recognition Foundation Model 提出ByteAction以解决压缩图像流中的动作识别问题 foundation model
9 Optimize Surgical Triplet Recognition: A Knowledge-Driven Mixture-of-Experts Solution 提出知识驱动的专家混合框架以优化外科三元组识别 large language model multimodal
10 DRAgent: Discriminative Reasoning Agent for Referring Expression Segmentation 提出DRAgent以解决语言表达指称分割中的定位偏差问题 large language model multimodal
11 Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation 提出Object-Uni以解决对象实例空间理解与可控生成问题 large language model multimodal
12 Towards Comprehensive Basketball Understanding 提出BasketballBench与BasketballSkills以解决篮球理解的多模态挑战 multimodal

🔬 支柱二:RL算法与架构 (RL & Architecture) (10 篇)

#题目一句话要点标签🔗
13 FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors 提出FixAnything以解决3D渲染伪影问题 DPO direct preference optimization 3DGS
14 GeoWAM: Visual Geometry World Action Models for Autonomous Driving 提出GeoWAM以解决自主驾驶中的场景动态建模问题 world model world models world action model
15 EchoWM: Open and Enterable Omnimodal World Models 提出EchoWM以解决多模态生成媒体的导航与同步问题 world model world models
16 BenthicFlow: Generating Extensible Underwater Environments via Flow Matching 提出BenthicFlow以解决水下环境3D场景理解问题 flow matching scene understanding
17 Following Motion for Sequential Modeling in Video Frame Interpolation 提出运动引导的选择性状态空间模型以解决视频帧插值问题 Mamba SSM state space model
18 Contextrast++: Robust Multi-Scale Contextual Contrastive Learning for Semantic Segmentation 提出Contextrast++以解决语义分割中的上下文捕捉与类别不平衡问题 representation learning contrastive learning
19 Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents 提出VideoRover以解决开放世界视频理解问题 reinforcement learning multimodal
20 BenthicDINO: Physics-Informed Self-Distillation for View-Invariant Side-Scan Sonar Representations 提出BenthicDINO以解决侧扫声纳图像中的视角不变性问题 distillation
21 Hyperbolic Hierarchical Clustering for Visual Representation Learning 提出ClusterMixer以解决视觉模型的可解释性问题 representation learning
22 Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds 提出JoyAI-Echo-1.5以解决长视频生成中的一致性与交互性问题 world model world models

🔬 支柱三:空间感知与语义 (Perception & Semantics) (8 篇)

#题目一句话要点标签🔗
23 AquaFlow: A Monocular Gaussian Splatting SLAM for Underwater Streaming Reconstruction 提出AquaFlow以解决水下场景重建中的视觉退化问题 3D gaussian splatting 3DGS gaussian splatting
24 LagrangeGS: Non-Conservative Lagrangian System on Dynamic 3D Gaussian Splatting 提出LagrangeGS以解决动态3D高斯点云物理一致性问题 3D gaussian splatting 3DGS gaussian splatting
25 NemoSplat: Feed-Forward 4D Gaussian Splatting for Media-Aware Underwater Reconstruction 提出NemoSplat以解决水下动态重建中的光学衰减问题 gaussian splatting splatting foundation model
26 Learning Spherical Occupancy Profiles for Multi-View 3D Reconstruction and Generation 提出球形占用轮廓以解决多视角3D重建与生成问题 3D reconstruction classifier-free guidance
27 Seeing the Unseen: Semantic-in-Gaussian for Sparse-View 3D Generalization 提出SeeU框架以解决稀疏视图下3D重建问题 3D gaussian splatting 3DGS gaussian splatting
28 SiZeUp: Fast 3D Proxy from Aerial Images via Depth Ordinal Loss 提出SiZeUp以解决大规模城市3D建模问题 monocular depth metric depth feature matching
29 Spotter: Efficient Urban Visual Localization via Geo-Referenced Facade Landmarks in GPS-Degraded Environments 提出Spotter以解决城市环境中视觉定位精度不足的问题 visual odometry stereo depth
30 Neighbor-Aware View Synthesis for Restoring Missing Views in Light-Field Camera Arrays 提出邻域感知视图合成以恢复光场相机阵列中的缺失视图 depth estimation

🔬 支柱一:机器人控制 (Robot Control) (2 篇)

#题目一句话要点标签🔗
31 Progressively Learning Heterogeneous Skills in a Unified Latent Space 提出HetSkills框架以在统一潜在空间中逐步学习异构技能 motion tracking distillation text-to-motion
32 Mover360: Controllable Object Manipulation in 360° Panoramic Images 提出Mover360以解决360°全景图像中的可控物体操作问题 manipulation

🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)

#题目一句话要点标签🔗
33 Results of the 1st Asynchronous CASTLE Challenge at the Joint Egocentric Vision Workshop in Conjunction with CVPR 2026 总结首届异步CASTLE挑战赛的贡献与结果 egocentric egocentric vision
34 Motion-Based Tokenization for Cross-Dataset Egocentric Gaze Modeling 提出基于运动的标记化方法以解决跨数据集自我中心注视建模问题 egocentric multimodal

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
35 Spatiotemporally Decoupled Autoregressive Diffusion Model for Human Motion Generation 提出DeMoDiff以解决人类动作生成中的表示与控制问题 motion synthesis motion generation human motion

⬅️ 返回 cs.CV 首页 · 🏠 返回主页