cs.CV(2026-08-25)

📊 共 25 篇论文 | 🔗 7 篇有代码

🎯 兴趣领域导航

支柱九:具身大模型 (Embodied Foundation Models) (9 🔗2) 支柱二:RL算法与架构 (RL & Architecture) (6 🔗2) 支柱一:机器人控制 (Robot Control) (3 🔗1) 支柱四:生成式动作 (Generative Motion) (2 🔗1) 支柱六:视频提取与匹配 (Video Extraction) (2 🔗1) 支柱七:动作重定向 (Motion Retargeting) (1) 支柱八:物理动画 (Physics-based Animation) (1) 支柱三:空间感知与语义 (Perception & Semantics) (1)

🔬 支柱九:具身大模型 (Embodied Foundation Models) (9 篇)

#题目一句话要点标签🔗
1 VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference 提出VisCache以解决视觉大语言模型推理效率问题 large language model multimodal
2 LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training 提出LAION-BVD以解决多模态预训练数据不足问题 multimodal
3 Graph-Supervised Hierarchical Clinical Alignment for Radiology Report Generation with Large Language Models 提出图监督层次临床对齐以解决放射学报告生成问题 large language model
4 MoTE: Mixture of Task Experts for Multi-Task Video Understanding 提出MoTE以解决多任务视频理解中的专家路由问题 large language model multimodal
5 ViSculpt: Visual-Centric Agentic Geometry Editing 提出ViSculpt以解决3D几何编辑的复杂性问题 large language model multimodal
6 Boot-and-Feedback Framework for Generalist-Expert Model Collaboration in Breast Ultrasound Diagnosis 提出Boot-and-Feedback框架以解决乳腺超声诊断中的模型协作问题 large language model multimodal
7 SandwichQuant: Which Parameters Matter Before and After Quantization? 提出SandwichQuant以优化量化前后参数调整 large language model
8 WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report 提出WeMM-Embedding以解决多模态嵌入问题 multimodal
9 Luce: Relightable Gaussians for 3D Asset Generation 提出Luce以解决高保真图像到3D生成问题 multimodal

🔬 支柱二:RL算法与架构 (RL & Architecture) (6 篇)

#题目一句话要点标签🔗
10 Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training 提出Game2World引擎以解决游戏视频训练数据问题 world model world models multimodal
11 TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation 提出TurboT2VA以解决大规模文本到视频音频生成的效率问题 distillation multimodal
12 On-Policy Self-Distillation in Diffusion Models 提出DiffusionOPSD以解决扩散模型与人类偏好对齐问题 reinforcement learning distillation
13 Joint-Embedding Prediction of Masked Point Tubes for Self-Supervised Learning on 4D Point Cloud Videos 提出JEPA框架以解决4D点云视频自监督学习问题 JEPA representation learning spatiotemporal
14 IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves 提出IDeaL方法以解决无数据多教师蒸馏问题 distillation
15 Representation Learning in Diffusion and Flow-based Model: An Application Aspect 提出三层进阶框架以提升生成模型的表示学习能力 representation learning

🔬 支柱一:机器人控制 (Robot Control) (3 篇)

#题目一句话要点标签🔗
16 LeFlow: Generative Latent Flow Planning for World Models 提出LeFlow以解决世界模型中的规划效率问题 trajectory optimization world model world models
17 NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation 提出NeoWorld-Pro以解决单目图像到交互场景构建问题 manipulation scene reconstruction physically plausible
18 VizAnchor: Decoding Manipulation Intent from Tampering Visualizations via Dual-Anchor Reasoning 提出VizAnchor以解决数据可视化操控理解问题 manipulation

🔬 支柱四:生成式动作 (Generative Motion) (2 篇)

#题目一句话要点标签🔗
19 SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling 提出SeMoCo以解决运动语言建模中的语义编码问题 text-to-motion language-conditioned motion motion generation
20 Metadata-Aware Adaptation of a Generative Foundation Model for Conditional CMR Synthesis 提出基于元数据的生成模型以改善心脏磁共振图像合成 classifier-free guidance foundation model

🔬 支柱六:视频提取与匹配 (Video Extraction) (2 篇)

#题目一句话要点标签🔗
21 From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms 提出统一框架以提升智能眼镜的第一人称智能平台能力 egocentric egocentric vision multimodal
22 EgoErrorVQA: Assess Egocentric Comprehension Capabilities through Procedural Errors for Ego-Agentic AI 提出EgoErrorVQA以解决程序性理解能力评估问题 egocentric

🔬 支柱七:动作重定向 (Motion Retargeting) (1 篇)

#题目一句话要点标签🔗
23 Event-Based Motion Estimation via Oriented Distance Fields 提出定向距离场运动估计以解决事件驱动运动估计的低延迟问题 motion estimation

🔬 支柱八:物理动画 (Physics-based Animation) (1 篇)

#题目一句话要点标签🔗
24 Mind the Student: Behavioral and Contextual Cues for Automated Engagement Prediction in Online Learning 提出多模态框架以解决在线学习中学生参与度预测问题 spatiotemporal multimodal

🔬 支柱三:空间感知与语义 (Perception & Semantics) (1 篇)

#题目一句话要点标签🔗
25 KLTNet: Learning Sparse Feature Tracking for Robust and Accurate Monocular Visual-Inertial Odometry 提出KLTNet以解决稀疏特征跟踪的鲁棒性与准确性问题 VIO optical flow

⬅️ 返回 cs.CV 首页 · 🏠 返回主页