| 1 |
SciMIF: Understanding Multimodal Instruction Following in Scientific Domains |
提出SciMIF以评估科学领域中的多模态指令遵循能力 |
large language model multimodal instruction following |
✅ |
|
| 2 |
Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs |
提出多粒度上下文增强的MMKG以提升多模态RAG性能 |
large language model multimodal |
|
|
| 3 |
MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities |
提出MMJailBench以解决多模态大型语言模型的监狱突破脆弱性问题 |
large language model multimodal |
|
|
| 4 |
Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders |
提出稀疏自编码器以实现中微子模型的可解释性 |
foundation model |
|
|
| 5 |
Data Citation for Large Language Models: A Challenge |
提出数据引用机制以解决大型语言模型的引用挑战 |
large language model |
|
|
| 6 |
Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents |
提出EASEL基准以解决多模态智能体的精细视觉工具使用问题 |
multimodal |
|
|
| 7 |
FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review |
提出FinRiskAtlas以解决金融风险审查中的模型评估问题 |
large language model |
|
|
| 8 |
AI Slop and Hallucinations in Vulnerability Assessment: A Survey on Reasoning Failures and Trustworthy Mitigation |
提出主动神经符号验证以解决AI漏洞评估中的信任危机 |
large language model chain-of-thought |
|
|
| 9 |
Narcissus: Program Synthesis Using Context-Aware LLM Approximations |
提出Narcissus以解决编程语言固定任务的合成问题 |
large language model |
|
|
| 10 |
CRAMER: Control via Request-Aware Masking for Editing Recommenders |
提出CRAMER框架以解决推荐系统对用户请求响应不足的问题 |
large language model |
|
|
| 11 |
Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems |
提出SymTrace框架以解决LLM多智能体系统调试问题 |
large language model |
|
|
| 12 |
ToST: A Tree-of-Thought Socratic Teaching Framework for Multi-Path Guidance and Parallel Thinking |
提出ToST框架以解决单路径教学限制问题 |
large language model |
|
|
| 13 |
CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval |
提出CaSKG框架以解决大语言模型技能检索问题 |
large language model |
✅ |
|
| 14 |
MACGen: Toward Functionally Correct and Secure Code Generation via Multi-Agent Collaboration |
提出MACGen以解决安全与功能正确性代码生成问题 |
large language model |
|
|
| 15 |
Q&A or Document-Based? The Effects of Interface Type on How Screen Reader Users Access Interconnected Documents |
比较问答界面与文档界面对盲人用户知识构建的影响 |
large language model |
|
|