cs.CV(2026-08-28)

📊 共 32 篇论文 | 🔗 7 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (12 🔗2) 支柱二:RL算法与架构 (RL & Architecture) (9 🔗2) 支柱三:空间感知与语义 (Perception & Semantics) (6 🔗2) 支柱四:生成式动作 (Generative Motion) (3) 支柱六:视频提取与匹配 (Video Extraction) (2 🔗1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (12 篇)

#题目一句话要点标签🔗
1 AIM: Anchor Identity Features, Then Match for Multimodal Large Language Model Unlearning 提出AIM方法以解决多模态大语言模型的身份遗忘问题 large language model multimodal
2 Visual Token Coding for Video Multimodal Large Language Models 提出视觉令牌编码以解决视频多模态大语言模型的令牌压缩问题 large language model multimodal
3 Iron: Intent-Aligned and Retrospective Dual Learning Framework for Enhancing Generalist Virtual Agents 提出Iron框架以解决虚拟代理任务自动化中的挑战 embodied AI generalist agent large language model
4 Semantic Head Specialization Guides Hybrid ViT Attention for Multimodal LLMs 提出语义头专业化指导混合ViT注意力以优化多模态LLM multimodal
5 Explainable Diabetic Retinopathy Classification Using Vision Foundation Models 提出可解释的糖尿病视网膜病变分类框架以提高筛查准确性 foundation model
6 EXPOSE: Explainable and Domain-Robust Embeddings from Pathology Vision Foundation Models using Sparse Autoencoders 提出EXPOSE框架以解决病理图像领域迁移问题 foundation model
7 Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding 提出并行管道解码以解决视频时空定位效率问题 large language model multimodal
8 Temporal Tree of Thought: Reasoning-Guided Visual Cue Search for Long-Video Understanding 提出Temporal Tree of Thought以解决长视频理解问题 large language model multimodal
9 Cut-ViT: Task-Specific Model Pruning via Gram Anchoring Subspace Consistency 提出Cut-ViT以解决视觉基础模型剪枝的鲁棒性和任务特异性问题 foundation model
10 Memory-efficient GPU pipelines for real-time non-line-of-sight reconstruction 提出内存高效的GPU管道以解决实时非视线重建问题 TAMP
11 Ada-TokenCom: Rate-Adaptive Token Communications via Large-Model-Driven Token Compression and Generation 提出Ada-TokenCom以解决低比特率语义通信问题 multimodal
12 Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models 提出动态对齐补偿以缓解大型视觉语言模型的幻觉问题 multimodal

🔬 支柱二:RL算法与架构 (RL & Architecture) (9 篇)

#题目一句话要点标签🔗
13 WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild 提出WilLaGS以解决不受约束场景下的3D重建与外观合成问题 teacher-student 3D gaussian splatting 3DGS
14 GAAT: Geometry-Aware Alignment Transformer for Multimodal UAV Perception 提出GAAT以解决无人机多模态感知中的对齐问题 contrastive learning scene understanding foundation model
15 uScenes: A Multimodal RGB and 3D Sonar Dataset for Underwater Robot Perception 提出uScenes数据集以解决水下机器人感知问题 representation learning scene understanding multimodal
16 WALDO: One-Shot Exemplar-Conditioned Object Detection in Cluttered Scenes 提出WALDO以解决复杂场景中的单次示例条件物体检测问题 world model world models JEPA
17 Token-Budget Distillation: Transferring Full-Token Semantics to Compressed Video Vision-Language Models 提出Token-Budget Distillation以解决视频VLM适应性高成本问题 teacher-student distillation
18 Denoising-Aware Temporal Point Cloud Completion for 3D Crop Architecture Recovery and Phenotypic Trait Extraction 提出Denoising-Aware方法以解决3D植物重建中的噪声与遮挡问题 Mamba MAE 3D reconstruction
19 CommerceVibe: Learning to Design E-Commerce Creatives as Executable Visual Code via Dual-Feedback Reinforcement Learning 提出CommerceVibe以解决电商创意生成中的结构化与可编辑性问题 reinforcement learning
20 Relational Knowledge Distillation Brings DNN Representations Close Enough to Humans to Be Aligned Without Supervision 提出关系知识蒸馏方法以实现DNN与人类表征的无监督对齐 distillation
21 Dual-Stream Semantic Guidance with Prototype Anchor Calibration for Source-Fully-Free Adaptation of Vision-Language Models 提出双流语义引导以解决视觉语言模型的源完全自由适应问题 teacher-student distillation

🔬 支柱三:空间感知与语义 (Perception & Semantics) (6 篇)

#题目一句话要点标签🔗
22 From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation 提出DEX以解决鱼眼相机深度估计和开放词汇分割问题 depth estimation monocular depth open-vocabulary
23 Non-Uniform Quantisation for 3DGS Compression 提出非均匀量化方案以解决3DGS压缩问题 3D gaussian splatting 3DGS gaussian splatting
24 A-PAIR: A Benchmark and Identity-Consistent Grounding Framework for Air-Ground Cross-View Referring Person Detection 提出A-PAIR框架以解决空地视角下的个体一致性检测问题 open-vocabulary open vocabulary
25 Video Generative Models as Geometry Learner 提出GeoNeXt以解决几何估计任务中的数据效率问题 monocular depth
26 Lossy Event Compression: From Event Stream Distortion to Task Performance 提出基于任务驱动的事件流压缩方法以解决带宽挑战 optical flow
27 ZipMVS: Multi-View Stereo with Compressed Cost Volumes 提出ZipMVS以解决多视角立体重建中的内存消耗问题 3D reconstruction

🔬 支柱四:生成式动作 (Generative Motion) (3 篇)

#题目一句话要点标签🔗
28 Cross-Spectral Dense Correspondence for Multimodal Spectral Medical Imaging 提出跨光谱密集对应方法以解决多模态医学成像问题 physically plausible HSI multimodal
29 GraspHOI: Full-Body 3D Human-Object Reconstruction with Finger-Level Grasps from a Single In-the-Wild Image 提出GraspHOI以解决单图像下全身3D人机交互重建问题 contact-aware penetration human-object interaction
30 SignRR: Retrieve and Refine Real Motion for Sign Language Production 提出SignRR以解决手语生成中的运动一致性问题 VQ-VAE

🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)

#题目一句话要点标签🔗
31 RASA: Disentangled Spatial-Motional Priors for Cross-Identity Character Animation 提出RASA框架以解决跨身份角色动画中的空间与运动控制问题 SMPL character animation
32 GAN-Based Semantic Communication for Image Transmission in IoV 提出基于GAN的语义通信框架以提升车联网图像传输效率 feature matching

⬅️ 返回 cs.CV 首页 · 🏠 返回主页