| 1 |
4DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting |
提出4DGS-WAM以解决传统WAM在空间结构建模中的不足 |
world model world models world action model |
|
|
| 2 |
Code World Model: Coding Agent as World Brain |
提出Code World Model以解决视频世界模型的局限性问题 |
world model world models spatiotemporal |
|
|
| 3 |
V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning |
提出V-Rubrics以解决视觉证据不足的问题 |
reinforcement learning multimodal instruction following |
|
|
| 4 |
MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations |
提出MLLMCLIP以解决视觉语言模型的组合性问题 |
distillation large language model multimodal |
|
|
| 5 |
AdaptiveEmbed: Sample-Adaptive Multi-Vector Representation for Multimodal Retrieval |
提出Sample-Adaptive Multi-Vector Representation以解决多模态检索中的固定容量问题 |
contrastive learning multimodal |
|
|
| 6 |
CloSeR: Unified Relational Distillation from Closed-Set Teachers for Category Discovery |
提出CloSeR框架以解决类别发现中的知识蒸馏问题 |
distillation foundation model |
✅ |
|
| 7 |
CrossMambaTuning: Synergistic Spatial and Cross-Layer Adaptation for Machine Vision Compression |
提出CrossMambaTuning以解决机器视觉压缩中的跨层适应问题 |
Mamba state space model |
✅ |
|
| 8 |
4DStreamCtrl: Interactive Video Generation with Online 4D Control |
提出4DStreamCtrl以实现实时4D视频生成与控制 |
world model world models spatiotemporal |
|
|
| 9 |
Embedding NDRE Trajectories into Contrastive Learning for Label-Free, Physiology-Aware Crop-Stress Staging and DSS Outputs |
提出EigenCL以解决作物压力检测不足问题 |
contrastive learning |
|
|
| 10 |
Socialized Detector Learning: Trajectory-Guided and Reciprocal Distillation for Heterogeneous Object Detectors |
提出社会化检测器学习以解决异构目标检测器知识碎片化问题 |
distillation |
|
|
| 11 |
DeCO: Discriminative Evidence Composition for Fine-Grained Dataset Distillation |
提出DeCO以解决细粒度数据集蒸馏中的局部证据缺失问题 |
distillation |
|
|
| 12 |
Where to Look Matters: On-Policy Self-Distillation for Long-Video Understanding |
提出Clue-OPSD框架以提升长视频理解精度 |
distillation |
|
|
| 13 |
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning |
提出VBVR-Pro以解决可扩展视觉推理训练问题 |
reinforcement learning spatiotemporal |
|
|
| 14 |
Asymmetric Cross-Modal Fine-Grained Visual Categorization: ACF-Net and the BirdPro Benchmark |
提出ACF-Net以解决不对称跨模态细粒度视觉分类问题 |
representation learning optical flow |
|
|