| 1 |
Grounding Without Corrective Control: Truth-Tracking Profiles for Large Language Models |
提出真相追踪配置以解决大型语言模型的纠正控制问题 |
large language model multimodal |
|
|
| 2 |
Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding |
提出mmWave-QA以解决毫米波雷达数据理解问题 |
large language model language conditioned |
|
|
| 3 |
Towards Efficient Multimodal and Multilingual Opinion Extraction for STI: A QLoRA-Based Fine-Tuning Approach |
提出多模态核心观点提取框架以解决信息过载问题 |
large language model multimodal |
|
|
| 4 |
A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images |
提出多模态理解框架以解决科学图像解读难题 |
multimodal |
|
|
| 5 |
Attributing Preprocessing Invariance in Spectral Foundation Models |
提出谱基础模型的预处理不变性评估方法 |
foundation model |
|
|
| 6 |
Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports |
提出Wyvern框架以生成多模态的基础技术报告 |
multimodal |
|
|
| 7 |
Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice |
审计算法以评估大语言模型推荐医生的偏见与透明性 |
large language model |
|
|
| 8 |
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning |
提出Mobius架构以提升知识压缩与推理效率 |
foundation model |
|
|
| 9 |
HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation |
提出HAM-RAG以解决多模态文档生成中的结构性问题 |
multimodal |
✅ |
|
| 10 |
Act2Intention: A Benchmark For Developing Active Mobile Agents Through Inferring User Intention from GUI Actions |
提出Act2Intention框架以解决主动移动代理用户意图推断问题 |
large language model multimodal |
|
|
| 11 |
Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers |
提出Regime-Conditional Verification以解决安全分类器适应性问题 |
large language model |
|
|
| 12 |
Handover of In-Context Learning State Across Session Boundaries |
提出会话状态交接方法以解决大语言模型的上下文限制问题 |
large language model |
|
|
| 13 |
SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning |
提出SheetCompass以解决自动化电子表格推理问题 |
large language model |
|
|
| 14 |
The Past and Future of AI Scientists |
提出AI科学家以解决科学自动化整合问题 |
foundation model |
|
|
| 15 |
LLMs Don't Pay for the Jump |
提出物理成本耦合机制以解决机器推理能力不足问题 |
large language model |
|
|
| 16 |
Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons |
提出Tripwire以解决大语言模型的安全防护问题 |
large language model |
|
|
| 17 |
A Hybrid LLM-Based Framework for Automated Security Annotation Generation in Business Process Models |
提出混合LLM框架以自动生成业务流程模型的安全注释 |
large language model |
|
|
| 18 |
AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs |
提出AnchorBench基准以评估LLMs中的锚定效应 |
large language model |
|
|
| 19 |
TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments |
提出TimeSage-EV以解决动态环境下时间序列分析的有效性问题 |
large language model |
|
|
| 20 |
Mandato: Protocol-Level Enforcement of Digitally Signed Mandates on AI Agent Actions with Cryptographically Chained Audit Trails |
提出Mandato以解决AI代理行动授权不足问题 |
TAMP |
|
|
| 21 |
Scaling Domain Data Repetition in LLM Pretraining |
提出重复高质量领域数据以优化LLM预训练效果 |
large language model |
|
|
| 22 |
MACS: A Hybrid Multi-Agent Framework for Reliable Conversational E-Commerce Recommendation |
提出MACS框架以解决电商推荐中的可靠性问题 |
large language model |
|
|
| 23 |
Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-efficiency |
提出基于LLM的自动程序修复方法以应对复杂错误问题 |
large language model |
|
|
| 24 |
Agent-Orchestration in Autonomous Chip Design |
提出AI组织模型以推动自主芯片设计的智能化 |
large language model |
|
|
| 25 |
Never the Number: Structural Abstention for AI Systems Whose Answers Are Consumed as Fact |
提出结构性拒绝机制以解决AI系统答案可靠性问题 |
large language model |
|
|