| 1 |
CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction |
提出CACSurv以解决癌症生存预测中的不匹配问题 |
large language model multimodal |
✅ |
|
| 2 |
HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents |
提出HaReCAP以解决长时间任务中的最后一步动作基础问题 |
large language model |
|
|
| 3 |
Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm |
将米尔格拉姆服从实验引入大型语言模型以评估其服从性 |
large language model |
|
|
| 4 |
MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment |
提出MUPA²E框架以解决多模态情感评估问题 |
multimodal |
|
|
| 5 |
ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction |
提出ALPS基准以衡量大型语言模型的有效创造力 |
large language model |
|
|
| 6 |
LAVA: Logic-Aware Validation and Augmentation Framework for Large-Scale Financial Document Auditing |
提出LAVA框架以解决金融文档审计中的验证与增强问题 |
large language model multimodal |
|
|
| 7 |
Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis |
提出拓扑归因距离(TAD)以解决LLM输出可信度问题 |
large language model |
|
|
| 8 |
TDD-Agent: Test-Driven Reasoning for Code Generation |
提出TDD-Agent以解决代码生成中的正确性问题 |
large language model |
|
|
| 9 |
Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies |
提出重建基准以从预出版文献中恢复研究思想 |
large language model |
|
|
| 10 |
Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization |
提出AudioChaps框架以解决音频章节化问题 |
chain-of-thought |
✅ |
|
| 11 |
A Policy Algebra for Trust-Preserving Agentic AI Execution |
提出信任保护的代理AI执行策略代数以解决可靠性问题 |
large language model |
|
|
| 12 |
Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152 |
提出RegulaRAG以解决合规场景生成问题 |
large language model |
|
|
| 13 |
Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs |
提出Ventor-QTest以解决第三方LLM API审核问题 |
large language model |
✅ |
|
| 14 |
AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment |
提出AeroCopilotBench以解决航空领域LLM代理评估问题 |
large language model |
|
|
| 15 |
Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming Assessments: Insights from 2026 |
评估生成式人工智能在面向对象编程评估中的表现 |
large language model |
|
|
| 16 |
Decoupled Temporal Encoding for Generative Recommendation |
提出解耦时间编码以解决推荐系统中的时间动态问题 |
TAMP |
|
|
| 17 |
MUSE: An Interactive Meta-Agent for Understanding and Steering LLM-powered Data Science Systems |
提出MUSE以增强用户对数据科学系统的理解与控制 |
large language model |
|
|