| 1 |
Q-based Variational Inverse Reinforcement Learning |
提出Q基变分逆强化学习以解决人类偏好学习问题 |
reinforcement learning inverse reinforcement learning |
|
|
| 2 |
CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated? |
提出CaliBench以解决视频世界模型的物理校准问题 |
world model world models |
|
|
| 3 |
SCALE: State-Calibrated Latent Embeddings for JEPA Planning in the Right Geometry |
提出SCALE以提升JEPA规划中的状态校准潜在嵌入 |
world model worldmodel world models |
|
|
| 4 |
Le Critique: Privileged Value Functions for LLM Reinforcement Learning |
提出特权价值函数以解决LLM强化学习中的方差问题 |
reinforcement learning large language model |
|
|
| 5 |
PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data |
提出PertMind以解决生物推理训练成本高的问题 |
reinforcement learning large language model |
|
|
| 6 |
TRACE-CASH: Trial-History-Conditioned Reinforcement Learning for Adaptive Configuration Exploration in Time-Series CASH |
提出TRACE-CASH以解决时间序列CASH中的自适应配置探索问题 |
reinforcement learning |
|
|
| 7 |
POI Recommendation with LLM-Augmented Multi-Graph Learning and Contrastive Alignment |
提出LLM-MGCL以解决POI推荐中的冷启动问题 |
contrastive learning multimodal |
|
|
| 8 |
The Trade-off Between Covariate Dependence and Latent Structure in Representation Learning |
提出统一框架以解决潜变量依赖与结构之间的权衡问题 |
representation learning |
|
|
| 9 |
An Analytical-Prior Framework for Data-Efficient Prediction of Sound-Reduction Frequencies in Rectangular Side-Branch Helmholtz Resonators |
提出分析先验框架以提高赫尔姆霍兹共鸣器的声减频率预测效率 |
MAE distillation |
|
|
| 10 |
Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated Value Dynamics |
通过受控引导和调节价值动态提出深度Q学习稳定性解决方案 |
reinforcement learning representation learning |
|
|