cs.CV(2026-08-27)

📊 共 58 篇论文 | 🔗 10 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (21 🔗3) 支柱二:RL算法与架构 (RL & Architecture) (18 🔗3) 支柱三:空间感知与语义 (Perception & Semantics) (10 🔗3) 支柱六:视频提取与匹配 (Video Extraction) (3) 支柱五:交互与反应 (Interaction & Reaction) (2 🔗1) 支柱八:物理动画 (Physics-based Animation) (2) 支柱七:动作重定向 (Motion Retargeting) (1) 支柱一:机器人控制 (Robot Control) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (21 篇)

#题目一句话要点标签🔗
1 Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models 提出视觉信息引导的并行解码方法以提升多模态生成质量 large language model multimodal
2 Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning 提出Aphanta框架以优化多模态推理中的图像编辑过程 large language model multimodal
3 From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation 提出MedREAL以解决医学图像分析中的像素级定位问题 large language model multimodal
4 Unsupervised Adaptation of 3D CT Foundation Models for 3D CBCT Segmentation 提出无监督领域适应框架以解决3D CBCT分割问题 foundation model
5 Data-efficient crack quantification in lithium-ion cathodes using foundation model transfer 提出基于基础模型转移的高效裂纹量化方法以解决锂离子电池阴极老化问题 foundation model
6 How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space 提出自监督框架以探索AI的艺术美学结构 multimodal
7 Anatomy-Guided Foundation Model Adaptation with Within-Case Prototype Supervision for Standard Plane Detection in Fetal Ultrasound Blind Sweeps 提出AnatoProto以解决胎儿超声盲扫中的标准平面检测问题 foundation model
8 FAN-LoRA: A Fourier-Adaptive Nonlinear Low-Rank Adaptor for Medical Foundation Model Domain Adaptation 提出FAN-LoRA以解决医学图像领域适应问题 foundation model
9 HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence 提出HUG-VIS以解决人本视觉智能的多模态理解与生成问题 multimodal
10 CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection 提出CODE以解决开放世界物体检测中的语义模糊问题 foundation model multimodal
11 AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability 提出AesCanvas以解决图像美学评估中的上下文适用性问题 large language model multimodal
12 Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information 提出视觉检索头以提升视觉语言模型的图像定位能力 large language model
13 DINOcular: Self-Supervised Visuospatial Representations 提出DINOcular框架以解决RGB-D视觉空间表示学习问题 foundation model
14 PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference 提出PACE以解决视觉语言模型推理速度慢的问题 large language model
15 Temporal Sensitivity Analysis of Tessera Embeddings 提出时间敏感性分析以优化土地覆盖映射 foundation model
16 Magpie: Real-Time World Renderer for Interactive Games 提出Magpie以解决互动游戏实时渲染问题 foundation model
17 AraMS-28k: The Largest Publicly Released Line-Level Dataset of Historical Arabic Manuscripts with Margin and Insertion-Anchor Annotations 提出AraMS-28k数据集以促进历史阿拉伯手稿的识别与分析 multimodal
18 Systematic Literature Review of Machine Learning Models and Applications for Text Recognition 系统评估机器学习模型以提升文本识别能力 multimodal
19 Fully Unleashing the Multimodal Attacker: Meta-Adaptive Jailbreaking of Vision-Language Models 提出Meta-Adaptive Multimodal Jailbreaking以解决视觉语言模型的安全性问题 multimodal
20 ShiftSplit-AD: Separating Domain Shift from Defects in Foundation-Feature Visual Anomaly Detection 提出ShiftSplit-AD以解决视觉异常检测中的领域偏移问题 foundation model
21 Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning 提出Code-as-World以解决物理推理中的世界表示问题 multimodal

🔬 支柱二:RL算法与架构 (RL & Architecture) (18 篇)

#题目一句话要点标签🔗
22 Cross-Architecture Knowledge Distillation from a Vision Foundation Model to a Lightweight Visual State Space Model for Tea Leaf Disease Classification 提出跨架构知识蒸馏方法以提升茶叶病害分类精度 SSM state space model distillation
23 Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models 提出逐步容量增长方法以优化视觉变换器编码器的任务适应性 world model world models JEPA
24 Video-OPSD: Exploiting Privileged Visual Evidence for On-Policy Self-Distillation in Video Large Language Models 提出Video-OPSD以解决视频大语言模型的自蒸馏问题 distillation large language model
25 LLaVAFlow: Preserving Latent Alignment Flow for Parameter-Efficient Multimodal Fine-Tuning 提出LLaVAFlow以解决多模态大语言模型遗忘问题 distillation large language model multimodal
26 Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher 提出Self-OPD以解决流匹配模型中的教师依赖问题 flow matching distillation large language model
27 Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs 提出Echo-GRPO以解决视频推理中的蒸馏训练问题 reinforcement learning distillation large language model
28 SpatialCrafter: Single Image World Modeling with Generative 3D Proxies 提出SpatialCrafter以解决图像到场景生成中的一致性问题 world model world models
29 PAWBench: How Far Are We from Probabilistically Aligned World Modeling? 提出PAWBench以评估视频生成模型的概率对齐能力 world model world models
30 R2M-Bench: Evaluating Revisit Memory via Relative Consistency in Interactive Video World Models 提出R2M-Bench以解决视频世界模型记忆评估问题 world model world models
31 FU-Mamba: A Frequency-Enhanced Dynamic Scanning Framework for Oralscan Image Segmentation 提出FU-Mamba框架以解决Oralscan图像分割问题 Mamba SSM state space model
32 Knowledge Distillation Driven Semantic NOMA with GAN Refinement for 6G Robotic Vehicle Networks 提出基于知识蒸馏的语义NOMA框架以解决6G机器人车辆网络中的干扰问题 distillation
33 Video-FLAIR: Not Whether to Reason, But How 提出Video-FLAIR以优化多模态推理策略 reinforcement learning multimodal
34 Generative Semantic Scene Completion 提出生成语义场景补全方法以解决户外LiDAR数据稀疏问题 flow matching semantic map
35 LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics 提出LeVJEPA以解决视频预训练计算成本高的问题 JEPA visual pre-training
36 LiveVVT: High-Fidelity Video Virtual Try-On in Real Time 提出LiveVVT以解决视频虚拟试穿中的延迟和计算开销问题 flow matching distillation
37 SpatialCrafter: Single Image World Modeling with Generative 3D Proxies 提出SpatialCrafter以解决图像到场景生成中的一致性问题 world model world models
38 PAWBench: How Far Are We from Probabilistically Aligned World Modeling? 提出PAWBench基准以评估视频生成模型的概率对齐能力 world model world models
39 LiveVVT: High-Fidelity Video Virtual Try-On in Real Time 提出LiveVVT以解决视频虚拟试穿中的延迟和计算开销问题 flow matching distillation

🔬 支柱三:空间感知与语义 (Perception & Semantics) (10 篇)

#题目一句话要点标签🔗
40 KISS-GS: 3D Gaussian Splatting Compression Kept Simple 提出KISS-GS以简化3D高斯点云压缩问题 3D gaussian splatting 3DGS gaussian splatting
41 CGS-SLAM: Collaborative Gaussian Splatting based SLAM for Multi-Agent Reconstruction 提出CGS-SLAM以解决多智能体重建中的RGB-D输入限制问题 monocular depth 3DGS gaussian splatting
42 UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization 提出UniGeo以解决文本引导的无人机地理定位问题 semantic mapping semantic map large language model
43 Text-to-seed generation: Training-free open-vocabulary seeded semantic segmentation via re-purposing diffusion as text-guided seed generator 提出Text-to-Seed框架以解决开放词汇语义分割问题 open-vocabulary open vocabulary foundation model
44 Per-View Gaussian Predictions Enable Training-Free Distractor Filtering in Feed-Forward 3DGS 提出训练无关的过滤方法以解决3D重建中的干扰物问题 3D gaussian splatting 3DGS 3D reconstruction
45 CoGeo-GS: Concept-Driven and Geometry-Aware Multi-Object Removal in 3D Scenes 提出CoGeo-GS以解决3D场景中的多物体去除问题 monocular depth 3D gaussian splatting 3DGS
46 Glass Surface Detection Grounded in 3D Visual Geometry 提出基于3D视觉几何的玻璃表面检测方法以解决透明性挑战 scene understanding VGGT
47 RECAP-Forcing: Retaining Content Appearances for Long Video Generation 提出RECAP-Forcing以解决长视频生成中的记忆挑战 optical flow
48 ABCD: Alpha-Composited Block Coordinate Descent: Constant-VRAM Training for Large Radiance Fields 提出ABCD框架以解决大规模辐射场训练中的内存限制问题 3D gaussian splatting 3DGS gaussian splatting
49 Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction 提出ABot-Recon以解决长视频流3D重建问题 3D reconstruction

🔬 支柱六:视频提取与匹配 (Video Extraction) (3 篇)

#题目一句话要点标签🔗
50 UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City 提出UrbanGround以解决城市环境中多模态语言模型的导航问题 first-person view large language model multimodal
51 Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning 提出CMPM基准以解决中文多面板表情包的视觉-语言推理问题 HuMoR multimodal
52 Order Matters: A Chinese Multi-Panel Meme Benchmark for Vision-Language Reasoning 提出CMPM基准以解决多面板中文表情包的视觉-语言推理问题 HuMoR multimodal

🔬 支柱五:交互与反应 (Interaction & Reaction) (2 篇)

#题目一句话要点标签🔗
53 Multi-Person Human Motion Forecasting in Complex Scenes 提出对象条件社交扩散模型以解决复杂场景中的多人运动预测问题 HOI multi-person interaction human motion
54 Reconstructing Humans and Objects in Interaction using Large Reconstruction Models 提出MILO框架以解决3D人机交互重建问题 human-object interaction HOI embodied AI

🔬 支柱八:物理动画 (Physics-based Animation) (2 篇)

#题目一句话要点标签🔗
55 Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning 提出多指令多镜头长视频编辑框架以解决视频编辑一致性问题 spatiotemporal large language model
56 Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation 提出TetherMem以解决长视频生成中的场景进展问题 spatiotemporal

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
57 Beyond Atomic Layouts: Compositional Design Understanding with Vision-Language Models 提出CoDeLayout与MASON以解决复杂布局理解问题 spatial relationship multimodal

🔬 支柱一:机器人控制 (Robot Control) (1 篇)

#题目一句话要点标签🔗
58 VidParse: Online Parsing of Egocentric Procedures Like a Pro 提出VidParse以解决自我中心视频解析中的复杂性问题 manipulation human-object interaction egocentric

⬅️ 返回 cs.CV 首页 · 🏠 返回主页