| 1 |
UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation |
提出UniMoFlow以解决3D人类动作编辑中的指令驱动问题 |
flow matching text-to-motion motion generation |
|
|
| 2 |
World Tokens: Enhancing Embodied Policies with Training-Time World Modeling |
提出World Tokens以提升具身策略的训练时世界建模能力 |
world model world models spatiotemporal |
|
|
| 3 |
Multi-Submap Implicit Neural SLAM with Local-to-Global Loop Closure for Large-Scale Scene Reconstruction |
提出多子图隐式神经SLAM以解决大规模场景重建问题 |
distillation NeRF neural radiance field |
✅ |
|
| 4 |
RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation |
提出RecoverFly框架以解决无人机视觉语言导航中的失败问题 |
reinforcement learning behavior cloning vision-language-action |
|
|
| 5 |
Did the Grid Erase the Event? EndoClock for Auditing Medical World-Model Pipelines |
提出EndoClock以解决医疗世界模型同步问题 |
world model world models PULSE |
|
|
| 6 |
Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning |
提出LDR以解决视频生成模型动态建模不足问题 |
world model world models latent dynamics |
|
|
| 7 |
Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots |
提出CVPD框架以实现自我蒸馏解决视觉盲点问题 |
distillation large language model multimodal |
|
|
| 8 |
ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly Detection |
提出ADOPD框架以解决工业异常检测中的参考依赖问题 |
distillation large language model multimodal |
✅ |
|
| 9 |
Foundation Models are Implicit Deepfake Detectors |
提出利用基础模型进行深伪检测的新方法 |
representation learning foundation model |
|
|
| 10 |
Sekai2: From World Exploration to Interactive World Modeling |
提出Sekai2以解决长视频生成与交互建模问题 |
world model world models |
|
|
| 11 |
UniDFKD: A Unified Semantic Prior Framework for Architecture-Agnostic Data-Free Knowledge Distillation |
提出UniDFKD框架以解决数据无关知识蒸馏中的架构依赖问题 |
teacher-student distillation |
|
|
| 12 |
RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation |
提出REST框架以解决高效图像生成中的蒸馏问题 |
reinforcement learning distillation |
|
|
| 13 |
Marrying Optimal Transport and ODEs for Unified Continuous-Time 4D Reconstruction and Tracking |
提出Uni4R框架以解决4D重建与点跟踪问题 |
flow matching TAMP |
|
|
| 14 |
TeaMatch: Teachable Cross-Modal Representation Learning for 2D-3D Matching |
提出TeaMatch以解决2D-3D匹配中的可靠性问题 |
representation learning |
|
|
| 15 |
CodecArena: Codec Quality Assessment via Visual Reinforcement Learning |
提出CodecArena以解决视频编码质量评估问题 |
reinforcement learning |
|
|
| 16 |
SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs |
提出SignLlama以解决无注释手语翻译中的视觉特征优先问题 |
distillation large language model |
|
|