cs.CV(2026-08-03)

📊 共 68 篇论文 | 🔗 18 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (24 🔗7) 支柱二:RL算法与架构 (RL & Architecture) (17 🔗3) 支柱三:空间感知与语义 (Perception & Semantics) (16 🔗4) 支柱七:动作重定向 (Motion Retargeting) (3 🔗1) 支柱一:机器人控制 (Robot Control) (3 🔗1) 支柱六:视频提取与匹配 (Video Extraction) (2) 支柱四:生成式动作 (Generative Motion) (1 🔗1) 支柱五:交互与反应 (Interaction & Reaction) (1 🔗1) 支柱八:物理动画 (Physics-based Animation) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (24 篇)

#题目一句话要点标签🔗
1 Few-Shot Concept Prompt Learning for Segmentation Foundation Models via Visual Grounding 提出少样本概念提示学习以解决医学图像分割问题 foundation model visual grounding
2 Illuminating Visual Identity in Universal Multimodal Embeddings 提出统一视觉身份辨别方法以提升多模态嵌入能力 large language model multimodal
3 Generative AI and Foundation Models in Medical Image 探讨生成式AI与基础模型在医学影像中的应用 large language model foundation model
4 UEmbed: Unified Sparse and Dense Multimodal Embeddings 提出UEmbed以解决多模态稀疏与密集嵌入统一问题 multimodal
5 Implicit Neural Representations for Multimodal Longitudinal Image Imputation and Interpolation 提出条件隐式神经表示以解决多模态MRI图像插补问题 multimodal
6 A General-Purpose VLM Can Teach an Astronomy Foundation Model to Better Recognize Galaxy Morphology 提出通用VLM以提升天文学基础模型的星系形态识别能力 foundation model
7 Deep Multimodal Fusion Detection through Spatial Mask and Channel Fusion 提出基于注意力驱动的互补重采样框架以提升跨模态目标检测 multimodal
8 MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing 提出MIEScore以解决多源图像编辑评估问题 large language model multimodal instruction following
9 SPECTRA: Band-Routed Embedding and Stage-Wise LoRA for Cross-Sensor Fine-Tuning of Geospatial Foundation Models 提出SPECTRA以解决地理基础模型的光谱不匹配和适应成本问题 foundation model
10 FAU at ImageCLEF 2026 Task on Multimodal Reasoning Robust Candidate Scoring and Concise Multilingual Visual Answering 提出多模态推理系统以解决视觉问答任务 multimodal
11 Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning 提出空间-频谱视觉锚学习以解决MLLM视觉感知退化问题 large language model foundation model multimodal
12 MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving 提出MoRAL以解决边缘计算平台上VLM的空间推理问题 multimodal chain-of-thought
13 ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs 提出ET-Prune以解决文本丰富输入的视觉令牌修剪问题 large language model multimodal
14 SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation 提出SVGEval以解决SVG生成评估的可靠性问题 multimodal visual grounding
15 Decoupling semantics from vision: A framework for faithful visual-text compression evaluation 提出新评估框架以解决视觉文本压缩评估问题 large language model multimodal
16 GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation 提出GEOID-Flood以解决洪水分割数据集不足问题 foundation model
17 Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression 提出基于消息的压缩方法以解决视觉语言模型的效率问题 multimodal
18 Invisible Ink Threats: Adversarial Goals Behind Legitimate Tasks in Computer-Use Agents 提出II-Bench以应对计算机使用代理中的隐形墨水威胁 large language model
19 LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks 提出LongHorizon-Harness以解决长时间任务状态管理问题 large language model
20 CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation 提出CultureVidBench以解决文本到视频生成中的文化理解问题 multimodal
21 Parameter-Dynamic Adaptive Fusion and Calibration Network for RGBT Tracking 提出PAFCNet以解决RGBT跟踪中的参数固定问题 multimodal
22 IDraw: Artist Verification from Digital Drawing Images 提出IDraw以解决数字绘画作者验证问题 multimodal
23 Learning Where to Look and How to Judge: Resolution-agnostic Image Quality Assessment with Quality-aware Saliency 提出ReLIQS以解决无参考图像质量评估中的分辨率依赖问题 multimodal
24 UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization 提出UniSim-SLAM以解决SLAM中的几何不一致性问题 foundation model

🔬 支柱二:RL算法与架构 (RL & Architecture) (17 篇)

#题目一句话要点标签🔗
25 DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents 提出DeepVoyager-VL以解决多模态长时间搜索问题 reinforcement learning large language model multimodal
26 DF$^3$: World Modeling via Decoder-Free Feature Forecasting in Autonomous Navigation 提出DF$^3$以解决自主导航中的世界建模问题 world model world models foundation model
27 AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning 提出AdaThinkV以解决视频推理中的令牌效率问题 reinforcement learning large language model multimodal
28 GeoCore-9B: Towards Geo-Aware Generative Foundation Models in Earth Observation 提出GeoCore-9B以解决地球观测生成模型的局限性 flow matching foundation model
29 WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity 提出WorldExam基准以评估世界模型的内在反应能力 world model world models
30 ReMiX-MAE: Learning Missing-Channel Cross-Modal Representations from RGB-Only Clinical Facial Videos for Sympathetic-Mediated Pain Assessment 提出ReMiX-MAE以解决临床面部视频缺失通道问题 masked autoencoder MAE multimodal
31 ASTRA: Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment 提出ASTRA以解决动态场景重建中的时间异步问题 MAE gaussian splatting splatting
32 StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field 提出StyleForge以解决室内家具风格一致性问题 preference learning large language model multimodal
33 SPARE: Structural Parameter-Free Affinity Regularization for Flow Matching 提出SPARE以解决去噪扩散变换器训练收敛慢的问题 flow matching classifier-free guidance
34 VR3D: View-Robust 3D Representation Learning for Aerial-Ground Person Re-Identification 提出VR3D框架以解决空中与地面行人重识别中的视角变化问题 representation learning 3D reconstruction
35 PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs 提出PhyCheck以解决视频语言模型对物理规律理解不足的问题 world model world models large language model
36 Loop-Mamba: A Loop Mamba with Degradation-Aware and Shared Memory for Old Photo Restoration 提出Loop-Mamba以解决老照片修复中的多重退化问题 Mamba
37 VC-Tooler: Learning Compositional and Adaptive Visual Tool Use 提出VC-Tooler以解决视觉工具使用的适应性与组合性问题 reinforcement learning multimodal
38 SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models 提出SPATIALQUERY以解决视觉语言模型的空间推理问题 MAE chain-of-thought
39 Linear Multi-Timescale Retention as a Memory-Efficient Vision-Language Bridge 提出线性多时间尺度保留模块以解决视觉语言模型的内存瓶颈问题 linear attention scene understanding
40 Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression 提出SPIRAL框架以解决视觉文本压缩中的路径不一致问题 DPO distillation
41 PartMat: Material-Aware 3D Part Decomposition with a Single Global Latent 提出PartMat以解决3D物体材料边界分解问题 reinforcement learning flow matching

🔬 支柱三:空间感知与语义 (Perception & Semantics) (16 篇)

#题目一句话要点标签🔗
42 UniqueSplat: View-conditioned 3D Gaussian Splatting for Generalizable 3D Reconstruction 提出UniqueSplat以解决视角适应性不足的3D重建问题 3D gaussian splatting 3D reconstruction gaussian splatting
43 DerainSplat: Feed-Forward Clean 3D Gaussian Splatting from Sparse Rainy Views 提出DerainSplat以解决稀疏雨景下3D场景重建问题 3D gaussian splatting 3DGS gaussian splatting
44 DecoupleGS: Interactive 3D Gaussian Splatting for End-to-End Autonomous Driving Testing 提出DecoupleGS以解决E2E自动驾驶测试中的动态场景问题 3D gaussian splatting 3DGS gaussian splatting
45 StreamSplat: Streaming Feed-Forward 3D Gaussian Splatting 提出StreamSplat以解决在线3D场景更新问题 3D gaussian splatting 3DGS gaussian splatting
46 D^2-4DGS: Dual-Depth Guided Sparse-Camera 4D Gaussian Splatting 提出D^2-4DGS以解决稀疏相机动态4D高斯渲染问题 monocular depth metric depth gaussian splatting
47 GIFT: Geometry-Invariant Fine-Tuning for Non-Lambertian Monocular Depth Estimation 提出GIFT以解决非朗伯面单目深度估计问题 depth estimation monocular depth foundation model
48 FAST-GS: Frequency Aware Space-time Gaussian Splatting for Photorealistic Dynamic Novel View Synthesis 提出频率感知时空高斯点云以解决动态场景合成问题 3D reconstruction gaussian splatting splatting
49 EOVSAM: Efficient Open-Vocabulary Segmentation with SAM 3 in One Pass 提出EOVSAM以解决SAM 3计算开销大的问题 open-vocabulary open vocabulary
50 PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph 提出PhotoHOI以解决从单张RGB照片合成3D手-物体交互的问题 open-vocabulary open vocabulary affordance
51 G-Skin: Learning to Bind 3D Gaussians with Generative Visual Priors 提出G-Skin以解决3D高斯模型绑定与动画生成问题 3D gaussian splatting gaussian splatting splatting
52 InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis 提出InfiniSplat以解决单图像3D高斯渲染问题 3D gaussian splatting 3DGS gaussian splatting
53 GSRAIN: Physically Calibrated High-/Low-Frequency Rainfall Synthesis for 3D Gaussian Driving Scenes 提出GSRAIN以解决自主驾驶场景中降雨模拟的物理可控性问题 3D gaussian splatting 3DGS gaussian splatting
54 Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery 提出DiffeoAfford框架以降低腹腔镜手术中外科医生的认知负担 affordance
55 CLEAR: Conflict-aware Learning via Evidence-guided Adaptive Routing for Unified Sparse-View 3D Gaussian Super-Resolution 提出CLEAR以解决稀疏视图3D高斯超分辨率问题 3D gaussian splatting gaussian splatting splatting
56 CHOW-SLAM: Compact Hybrid Representation with Complementary Overlap Window Optimization for RGB-D SLAM 提出CHOW-SLAM以解决RGB-D SLAM中的空间与时间约束问题 NeRF neural radiance field scene reconstruction
57 MoCRA: Mixture of Compositional Rank-1 Atoms for 4K All-in-One Video Restoration 提出MoCRA以解决4K视频恢复中的多重退化问题 optical flow

🔬 支柱七:动作重定向 (Motion Retargeting) (3 篇)

#题目一句话要点标签🔗
58 Sen-Cap: Sensor-Flexible and Noise-Resilient Human Motion Capture via LiDAR-Camera Integration 提出Sen-Cap以解决多模态传感器对齐及噪声鲁棒性问题 human motion
59 Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations 提出超越形态的运动转移方法以解决视频动画问题 motion representation
60 Extended Field of View Analysis for VideoGAN-based Trajectory Generation 提出基于VideoGAN的轨迹生成方法以解决交通行为复杂性问题 spatial relationship

🔬 支柱一:机器人控制 (Robot Control) (3 篇)

#题目一句话要点标签🔗
61 Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models 提出进化计算引导的跨任务攻击框架以增强视觉语言模型的鲁棒性 trajectory optimization multimodal
62 Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow 提出生成检测器以解决开放集视觉文本取证问题 manipulation flow matching
63 SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models 提出SpatioLM以解决视觉语言模型的空间推理问题 manipulation

🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)

#题目一句话要点标签🔗
64 Dynamic Resolution Routing for Efficient Egocentric Grounding 提出SmartRes框架以解决自我中心视觉定位中的高分辨率输入问题 egocentric Ego4D large language model
65 HiResNets: Native Full-HD Video Recognition with Foveal Residual Streams 提出HiResNets以解决高分辨率视频识别中的内存瓶颈问题 egocentric

🔬 支柱四:生成式动作 (Generative Motion) (1 篇)

#题目一句话要点标签🔗
66 GenPrior: Unleashing Text-to-Motion Generative Priors for Zero-Shot Skeleton-based Action Recognition 提出GenPrior以解决零样本骨架动作识别中的语义-运动差距问题 text-to-motion

🔬 支柱五:交互与反应 (Interaction & Reaction) (1 篇)

#题目一句话要点标签🔗
67 UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation 提出UniMoCa以解决多人物视频生成中的运动与相机控制问题 multi-person interaction human motion

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
68 Global-Scale Self-Supervised Spatiotemporal Learning for NDVI Time-Series Reconstruction 提出GloSSR框架以解决NDVI时间序列重建问题 spatiotemporal

⬅️ 返回 cs.CV 首页 · 🏠 返回主页