| 1 |
GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory |
提出GraphMemix以解决多模态代理长期记忆组织问题 |
foundation model multimodal |
|
|
| 2 |
BrailleBench: Investigating Multi-Criteria Braille Comprehension in Large Language Models |
提出BrailleBench以解决盲人和聋盲用户的盲文理解问题 |
large language model |
|
|
| 3 |
Discovering Relationships in Data Lakes Using Large Language Models: An Industrial Case |
提出ColRel方法以解决数据湖中列关系发现问题 |
large language model |
|
|
| 4 |
Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation |
提出Multi2AV-Safety以评估多模态音视频生成的安全性 |
multimodal |
|
|
| 5 |
From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities |
提出过程级基准以评估LLMs的数学智能能力 |
large language model multimodal |
|
|
| 6 |
SymbolLKG: Towards Verifiable Logical Reasoning via Logical Knowledge Graph and Symbolic Solvers |
提出SymbolLKG以解决逻辑推理的可验证性问题 |
large language model chain-of-thought |
|
|
| 7 |
LLMs Can Design Near-Optimal OR Algorithms |
利用大型语言模型设计近似最优的运筹学算法 |
large language model |
|
|
| 8 |
SWE-Prime: Fewer Trajectories, Better Performance |
提出SWE-Prime以优化大语言模型的软件问题解决能力 |
large language model |
|
|
| 9 |
From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench |
提出MCR-Bench以解决多轮代码审查的真实场景问题 |
large language model |
|
|
| 10 |
Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit |
提出Persona-Execution Separation以解决LLM代理执行审计问题 |
large language model |
|
|
| 11 |
Sophistication in GenAI Use: Field Evidence from a Large Firm |
研究生成性AI使用的复杂性及其影响因素 |
large language model |
|
|
| 12 |
Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance |
提出能力框架以提升模型评估合规性 |
chain-of-thought |
|
|
| 13 |
LLMs in Digital EDA: A perspective on shifting roles from Generation to Orchestration |
提出层次化角色模型以优化电子设计自动化中的LLM应用 |
large language model |
|
|
| 14 |
Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents |
提出LoopHarness以解决自主LLM代理的安全性问题 |
large language model |
|
|
| 15 |
LAAF: A Layered Accountability Architecture Framework for LLM Applications |
提出LAAF框架以解决大语言模型应用中的问责问题 |
large language model |
|
|
| 16 |
pro-team at LLMs4OL 2026 Tasks Flagship and Reuse: Retrieval-Augmented Generation and Vocabulary-Constrained Filtering for Ontology Learning |
提出基于检索增强生成的本体学习方法以解决现有挑战 |
large language model |
|
|
| 17 |
Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency |
比较人类与LLM筛选工作流程以优化文献综述中的回忆与工作负载平衡 |
large language model |
|
|
| 18 |
C-Unseen: Weak Signal Detection in Dynamic Temporal Knowledge Graphs via LLM Reasoning |
提出C-Unseen以解决动态时间知识图中的弱信号检测问题 |
chain-of-thought |
|
|
| 19 |
BekchiAI: Measuring, Observing, and Controlling LLM Agents in One Click |
提出BekchiAI以测量和控制大语言模型代理的能力 |
large language model |
|
|
| 20 |
LiveSim: Simulating Environment-Shaped Users in Multi-Agent Live-Stream Ecosystems |
提出LiveSim以解决直播生态系统中用户行为模拟不足的问题 |
large language model |
|
|
| 21 |
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training |
提出边界校准干预转移方法以优化自主LLM后训练 |
large language model |
|
|
| 22 |
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling |
提出AgentJudgeBench基准以评估LLM在工具调用中的可靠性 |
chain-of-thought |
|
|
| 23 |
Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance |
提出基于账本控制的自我协调机制以提升LLM编码性能 |
large language model |
|
|
| 24 |
Nemotron 3.5 Content Safety Moderator: A Compact Multimodal, Multilingual, and Reasoning Enabled Content Safety Moderator |
提出Nemotron 3.5以解决多模态内容安全审核问题 |
multimodal |
|
|
| 25 |
LitCurate: A Configuration-Driven AI-Assisted Framework for Scientific Database Construction with an Application to Lower-Mantle Equation-of-State Data |
提出LitCurate框架以解决科学数据库构建问题 |
large language model |
|
|
| 26 |
LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails |
提出LongGuard以解决长文本安全防护失效问题 |
large language model |
|
|
| 27 |
Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents |
提出LoopHarness以解决自主LLM代理的安全性问题 |
large language model |
|
|