| 13 |
FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors |
提出FixAnything以解决3D渲染伪影问题 |
DPO direct preference optimization 3DGS |
|
|
| 14 |
GeoWAM: Visual Geometry World Action Models for Autonomous Driving |
提出GeoWAM以解决自主驾驶中的场景动态建模问题 |
world model world models world action model |
|
|
| 15 |
EchoWM: Open and Enterable Omnimodal World Models |
提出EchoWM以解决多模态生成媒体的导航与同步问题 |
world model world models |
|
|
| 16 |
BenthicFlow: Generating Extensible Underwater Environments via Flow Matching |
提出BenthicFlow以解决水下环境3D场景理解问题 |
flow matching scene understanding |
✅ |
|
| 17 |
Following Motion for Sequential Modeling in Video Frame Interpolation |
提出运动引导的选择性状态空间模型以解决视频帧插值问题 |
Mamba SSM state space model |
|
|
| 18 |
Contextrast++: Robust Multi-Scale Contextual Contrastive Learning for Semantic Segmentation |
提出Contextrast++以解决语义分割中的上下文捕捉与类别不平衡问题 |
representation learning contrastive learning |
|
|
| 19 |
Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents |
提出VideoRover以解决开放世界视频理解问题 |
reinforcement learning multimodal |
|
|
| 20 |
BenthicDINO: Physics-Informed Self-Distillation for View-Invariant Side-Scan Sonar Representations |
提出BenthicDINO以解决侧扫声纳图像中的视角不变性问题 |
distillation |
|
|
| 21 |
Hyperbolic Hierarchical Clustering for Visual Representation Learning |
提出ClusterMixer以解决视觉模型的可解释性问题 |
representation learning |
|
|
| 22 |
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds |
提出JoyAI-Echo-1.5以解决长视频生成中的一致性与交互性问题 |
world model world models |
✅ |
|