| 1 |
Twin Rollouts: Noise-Coupled Counterfactual Branching in Interactive Video World Models |
提出噪声耦合双重回放以解决交互视频世界模型的反事实生成问题 |
world model world models spatiotemporal |
|
|
| 2 |
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models |
提出DreOPD以解决流匹配模型的优化冲突问题 |
reinforcement learning flow matching distillation |
|
|
| 3 |
Multimodal Federated Learning under Dual-Axis Modality Missingness |
提出Flux框架以解决双轴模态缺失问题的多模态联邦学习 |
representation learning multimodal |
✅ |
|
| 4 |
SR-OPSD: Self-Referenced On-Policy Self-Distillation |
提出SR-OPSD以解决现有自蒸馏方法的不稳定性问题 |
reinforcement learning distillation large language model |
|
|
| 5 |
Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation |
提出SKALD框架以解决强化学习中的奖励信号不足问题 |
reinforcement learning distillation |
|
|
| 6 |
Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation |
提出OP²SD以探讨教师行为对自蒸馏的影响 |
distillation privileged information |
|
|
| 7 |
MARA: Flow-Matching-Guided Multi-Agent Resource Allocation for Computational Resource Efficient Learning |
提出MARA以解决多任务计算资源分配问题 |
reinforcement learning flow matching |
|
|
| 8 |
Bayesian Symbolic Regression with Entropic Reinforcement Learning |
提出ERRLESS以解决符号回归中的不确定性量化问题 |
reinforcement learning |
|
|
| 9 |
Learning to Modulate, Not to Cycle: Soft Actor---Critic Recovers Inverter-Style Heat-Pump Control |
提出基于SAC的控制策略以解决热泵压缩机循环问题 |
reinforcement learning PPO SAC |
|
|
| 10 |
WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training |
提出WDL-OPD以解决在线蒸馏的不稳定性问题 |
distillation |
|
|
| 11 |
Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training |
提出TrajVal以解决任务学习性评估不足问题 |
reinforcement learning large language model |
|
|
| 12 |
Dynamic Distribution-Aware Uncertainty Tracking in Vision-Language Representation Learning |
提出动态分布感知的不确定性跟踪以解决视觉语言模型的可靠性问题 |
representation learning |
|
|