| 1 |
Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency |
提出动作条件预测一致性以诊断JEPA世界模型的视觉扰动问题 |
world model world models JEPA |
|
|
| 2 |
Intern-S2-Preview: Scientific Agentic Foundation Model |
提出Intern-S2-Preview以解决科学发现中的多模态理解问题 |
reinforcement learning distillation foundation model |
|
|
| 3 |
Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling |
提出基于亲密距离的奖励模型以解决社交合规导航问题 |
reinforcement learning deep reinforcement learning DRL |
|
|
| 4 |
Intervention-Aware Clinical World Model for Post-Op Outcome Forecasting in Cardiology |
提出干预感知临床世界模型以预测心脏病术后结果 |
world model world models MAE |
|
|
| 5 |
The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use |
提出新的目标函数以解决长时间规划中的瓶颈问题 |
world model worldmodel world models |
|
|
| 6 |
CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation |
提出CardioState-JEPA以解决心脏信号多模态学习问题 |
JEPA Joint-Embedding Predictive Architecture joint-embedding predictive architecture |
|
|
| 7 |
The Impact of Temporal Context Length and Encoding Strategies on Self-Supervised ECG Representation Learning |
提出基于时间上下文长度与编码策略的自监督ECG表示学习方法 |
representation learning foundation model |
✅ |
|
| 8 |
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents |
提出步骤级自蒸馏策略优化以解决深度搜索代理的稀疏奖励问题 |
reinforcement learning teacher-student distillation |
|
|
| 9 |
FlowLOB: Efficient and Controllable Limit Order Book Generation with Flow Matching |
提出FlowLOB以解决限价单簿生成效率与可控性问题 |
flow matching |
|
|
| 10 |
Latent On-Policy Self-Distillation |
提出潜在的在线自蒸馏方法以解决自我进化AI的学习效率问题 |
distillation |
|
|
| 11 |
Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning |
提出在线推断方法以优化分布式强化学习中的量化时间差学习 |
reinforcement learning |
|
|
| 12 |
I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization |
提出I-SDPO以解决自蒸馏策略优化中的偏差问题 |
distillation |
|
|
| 13 |
Dynamic Multi-Depot Vehicle Routing with Online Requests: Event-Driven Transformer--DRL and Rolling-Horizon Benchmarking |
提出事件驱动的动态多仓库车辆调度框架以应对在线请求问题 |
DRL PPO behavior cloning |
|
|
| 14 |
The Query Knows What to Forget: A Second Erase Direction for Linear Attention |
提出查询导向的擦除方向以解决线性注意力中的干扰问题 |
linear attention |
|
|
| 15 |
Contrastive Learning for Interpretable Anomaly Detection at Collider Experiments |
提出ORCA以解决对撞机实验中的可解释异常检测问题 |
contrastive learning |
|
|
| 16 |
Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning |
提出在线推断方法以优化分布式强化学习中的量化时间差学习 |
reinforcement learning |
|
|