| 14 |
What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations |
提出证书门控协议以识别潜在世界模型中的物理参数 |
world model world models multimodal |
|
|
| 15 |
Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method |
提出时序中心化SIGReg以改善多任务LeWorldModel学习 |
behavior cloning diffusion policy world model |
|
|
| 16 |
HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models |
提出HiFloat4格式与Rollout Residual Quantization以解决FP4强化学习后训练精度问题 |
reinforcement learning large language model |
|
|
| 17 |
CalTwin: Towards Calibrated, Shift-Robust Medical World Models via Fisher-Information Regularisation |
提出CalTwin以解决医疗世界模型的校准与鲁棒性问题 |
world model world models latent dynamics |
|
|
| 18 |
DHRCL:Training Code LLMs with Dense Hierarchical Rewards and Curriculum Learning |
提出DHRCL框架以优化代码大语言模型的训练效果 |
reinforcement learning curriculum learning large language model |
|
|
| 19 |
CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation |
提出CoCaRS以解决异构知识蒸馏中的冗余抑制问题 |
SAC teacher-student distillation |
|
|
| 20 |
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution |
提出SkillRise以解决跨任务技能演化问题 |
reinforcement learning large language model |
|
|
| 21 |
Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning |
提出CWAC以解决离线强化学习中的过度估计问题 |
reinforcement learning SAC TD3 |
|
|
| 22 |
DREvo: Distilling Recalibrated Historical Experience for Harness Self-Evolution |
提出DREvo以解决历史经验在自我进化中的有效性问题 |
distillation large language model |
|
|
| 23 |
Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement |
提出IRIS框架以解决个性化语言模型的用户建模问题 |
preference learning large language model |
|
|
| 24 |
Single-Beat Cuffless Blood Pressure Estimation Using Ear-PPG and ECG with a Lightweight Hybrid Learning Framework |
提出轻量级混合学习框架以实现无袖带血压监测 |
MAE PULSE |
✅ |
|