| 1 |
StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models |
提出StrategyBench以评估大语言模型中的显式策略诱导能力 |
large language model |
|
|
| 2 |
Modalities Should Talk to Each Other: Dual-Stream Multimodal Learning for Long-Horizon Influenza Forecasting |
提出双流注意力机制以解决流感预测中的多模态融合问题 |
multimodal |
|
|
| 3 |
Proxy reliance in large language model decisions is uncalibrated to predictive evidence |
提出代理依赖性评估方法以解决大语言模型决策不准确问题 |
large language model |
|
|
| 4 |
Let the Bullets Fly: Multimodal Fake News Detection with Temporal-Aligned Generative Danmaku |
提出Genda框架以解决假新闻检测中的时序不一致问题 |
multimodal |
|
|
| 5 |
Performance of a domain-specific large language model in answering patient questions in psychiatry |
提出MIND模型以提升精神科患者教育问答质量 |
large language model |
|
|
| 6 |
Compositional Chain-of-Relations for Faithful Knowledge Graph Question Answering with Large Language Models |
提出关系中心探索框架以解决复杂知识图谱问答问题 |
large language model |
|
|
| 7 |
NetConfArena: An Executable Benchmark for LLM Agents in Closed-Loop Network Configuration |
提出NetConfArena以评估LLM代理在闭环网络配置中的表现 |
large language model foundation model |
|
|
| 8 |
LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications |
提出基于LLM的预测代理以增强时间序列预测能力 |
large language model foundation model |
|
|
| 9 |
LLM-Based Selection of Incongruent Verbal and Nonverbal Behavior for Virtual Humans |
提出基于LLM的虚拟人不一致行为选择方法以增强交互真实感 |
large language model |
|
|
| 10 |
Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty |
提出安全方向惩罚以解决推理引发的失调问题 |
chain-of-thought |
|
|
| 11 |
SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning |
提出自反政策优化框架以提升长时间推理能力 |
large language model |
✅ |
|
| 12 |
Walking on the DARKSIDE |
提出DARKSIDE以增强LLM的逻辑一致性审计 |
large language model |
|
|
| 13 |
EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models |
提出EviSafe框架以解决视觉语言模型安全评估问题 |
multimodal |
|
|
| 14 |
Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance |
提出探测方法以解决大型语言模型的伦理合规问题 |
large language model |
|
|
| 15 |
Automated Construction of FAIR Digital Object Knowledge Graphs from Flat Cultural Heritage Records |
提出自动化构建FAIR数字对象知识图谱的方法以解决文化遗产记录的整合问题 |
large language model |
|
|
| 16 |
Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data |
提出混合SFT以超越下一块推理强化学习的性能 |
chain-of-thought |
|
|
| 17 |
POOL: Propagated Uncertainty Over Lookalikes |
提出POOL框架以解决黑箱语言模型的置信度估计问题 |
large language model |
|
|
| 18 |
AgentWeave: Routing Before Reasoning for Efficient Function Calling in Tool-Rich Language Models |
提出AgentWeave以提高工具丰富语言模型的函数调用效率 |
large language model |
|
|
| 19 |
From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation |
提出NIS-Agent以解决深度研究代理的惯性偏差问题 |
large language model |
|
|
| 20 |
Budget-Constrained Embodied Perception: Four Resource Walls and a Pre-Registered Evaluation of Access-Structured Perception on Open Models at less than 31B |
提出ASP方法以解决预算约束下的多模态感知问题 |
multimodal |
|
|
| 21 |
Toward Effective and Reliable LLM Agents via Dynamic Ontology |
提出OaK框架以动态构建可靠的LLM代理本体 |
large language model |
|
|
| 22 |
Your AI, On a Dial: Controlling Investment Bias in LLMs with a Single Neuron |
提出投资偏见调节器以解决LLMs投资决策偏差问题 |
large language model |
|
|
| 23 |
Minimal Local Simulation Foundations for LLM- and VLM-Driven Agents in 2D and 3D Environments |
提出最小化本地仿真基础以支持LLM和VLM驱动的智能体 |
large language model |
|
|
| 24 |
The Retriever Should Remember: Experience-Amortized Reranking for Long-Term Agent Memory |
提出EARM框架以解决长期代理记忆中的检索效率问题 |
large language model |
|
|
| 25 |
Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf |
探讨AI代理购物中的位置偏差及其影响 |
large language model |
|
|