| 22 |
Cross-Architecture Knowledge Distillation from a Vision Foundation Model to a Lightweight Visual State Space Model for Tea Leaf Disease Classification |
提出跨架构知识蒸馏方法以提升茶叶病害分类精度 |
SSM state space model distillation |
|
|
| 23 |
Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models |
提出逐步容量增长方法以优化视觉变换器编码器的任务适应性 |
world model world models JEPA |
|
|
| 24 |
Video-OPSD: Exploiting Privileged Visual Evidence for On-Policy Self-Distillation in Video Large Language Models |
提出Video-OPSD以解决视频大语言模型的自蒸馏问题 |
distillation large language model |
|
|
| 25 |
LLaVAFlow: Preserving Latent Alignment Flow for Parameter-Efficient Multimodal Fine-Tuning |
提出LLaVAFlow以解决多模态大语言模型遗忘问题 |
distillation large language model multimodal |
|
|
| 26 |
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher |
提出Self-OPD以解决流匹配模型中的教师依赖问题 |
flow matching distillation large language model |
|
|
| 27 |
Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs |
提出Echo-GRPO以解决视频推理中的蒸馏训练问题 |
reinforcement learning distillation large language model |
|
|
| 28 |
SpatialCrafter: Single Image World Modeling with Generative 3D Proxies |
提出SpatialCrafter以解决图像到场景生成中的一致性问题 |
world model world models |
✅ |
|
| 29 |
PAWBench: How Far Are We from Probabilistically Aligned World Modeling? |
提出PAWBench以评估视频生成模型的概率对齐能力 |
world model world models |
|
|
| 30 |
R2M-Bench: Evaluating Revisit Memory via Relative Consistency in Interactive Video World Models |
提出R2M-Bench以解决视频世界模型记忆评估问题 |
world model world models |
|
|
| 31 |
FU-Mamba: A Frequency-Enhanced Dynamic Scanning Framework for Oralscan Image Segmentation |
提出FU-Mamba框架以解决Oralscan图像分割问题 |
Mamba SSM state space model |
✅ |
|
| 32 |
Knowledge Distillation Driven Semantic NOMA with GAN Refinement for 6G Robotic Vehicle Networks |
提出基于知识蒸馏的语义NOMA框架以解决6G机器人车辆网络中的干扰问题 |
distillation |
|
|
| 33 |
Video-FLAIR: Not Whether to Reason, But How |
提出Video-FLAIR以优化多模态推理策略 |
reinforcement learning multimodal |
|
|
| 34 |
Generative Semantic Scene Completion |
提出生成语义场景补全方法以解决户外LiDAR数据稀疏问题 |
flow matching semantic map |
|
|
| 35 |
LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics |
提出LeVJEPA以解决视频预训练计算成本高的问题 |
JEPA visual pre-training |
|
|
| 36 |
LiveVVT: High-Fidelity Video Virtual Try-On in Real Time |
提出LiveVVT以解决视频虚拟试穿中的延迟和计算开销问题 |
flow matching distillation |
|
|
| 37 |
SpatialCrafter: Single Image World Modeling with Generative 3D Proxies |
提出SpatialCrafter以解决图像到场景生成中的一致性问题 |
world model world models |
✅ |
|
| 38 |
PAWBench: How Far Are We from Probabilistically Aligned World Modeling? |
提出PAWBench基准以评估视频生成模型的概率对齐能力 |
world model world models |
|
|
| 39 |
LiveVVT: High-Fidelity Video Virtual Try-On in Real Time |
提出LiveVVT以解决视频虚拟试穿中的延迟和计算开销问题 |
flow matching distillation |
|
|