| 1 |
ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models |
提出ReFrame以解决多模态大语言模型的安全对齐问题 |
large language model multimodal |
|
|
| 2 |
Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis |
提出多模态推测解码以解决扩散基础并行草拟问题 |
vision-language-action VLA multimodal |
|
|
| 3 |
CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models |
提出CellPath-Bench以评估病理基础模型的细胞表示能力 |
foundation model |
|
|
| 4 |
Large Language Models at the Intersection of Software Engineering and Software Security:An Evidence-Centered Structured Survey and Research Agenda |
提出证据中心结构调查以解决软件工程与安全交集问题 |
large language model |
|
|
| 5 |
Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance |
提出DGEval基准以评估大型语言模型在IMDG合规性中的表现 |
large language model |
|
|
| 6 |
Can Scientific Claims Be Removed from Large Language Models? A Systematic Evaluation of Claim-Level Unlearning |
提出科学声明遗忘任务以解决语言模型知识过时问题 |
large language model |
|
|
| 7 |
Foundation Models for Partial Causal Identification |
提出因果基础模型以解决部分因果识别问题 |
foundation model |
|
|
| 8 |
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming |
提出TLive-Omni以解决电商直播中的多模态理解问题 |
instruction following visual grounding |
|
|
| 9 |
TRACE: Agentic Catalog Enrichment with Multi-source Evidence Grounding |
提出TRACE框架以解决电商目录属性稀缺问题 |
large language model multimodal |
|
|
| 10 |
When Generated Images Look Right and Retrieve Wrong: Coverage-Guided Cross-Scale Re-Indexing for Knowledge-Faithful Generative Perception |
提出CERES框架以解决生成图像检索不一致问题 |
multimodal |
|
|
| 11 |
Neuro-Geospatial Modelling of EEG Affective States Using Literature-Informed Environmental Context |
提出基于文献信息的环境上下文的EEG情感状态建模方法 |
multimodal |
✅ |
|
| 12 |
Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems |
提出可信RAG以解决生成AI系统中的虚假信息和知识污染问题 |
large language model |
✅ |
|
| 13 |
Can Legal AI Know When It Is Wrong? And Do Students Know When It Is? |
提出高置信度错误率以解决法律AI过度自信问题 |
large language model |
|
|
| 14 |
Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models |
提出MOSAIC基准以解决理论心智与社会行为之间的脱节问题 |
multimodal |
|
|
| 15 |
Deep Learning Models Also Recall Features |
提出特征回忆概念以深化深度学习模型理解 |
large language model |
|
|
| 16 |
Structured but Fragile: On the Limits of LLMs in Cybersecurity Decision-Making |
探讨大型语言模型在网络安全决策中的局限性 |
large language model |
|
|
| 17 |
Vibe Coding and Web Application Security: A Twin-Prompt Study |
研究安全意识提示对生成Web应用程序安全性的影响 |
large language model |
|
|
| 18 |
The Logic of Machine Self-Preservation |
探讨机器自我保护逻辑以应对智能体行为问题 |
large language model |
|
|
| 19 |
BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP |
提出BC-Bench以评估ERP领域特定语言中的代理工程 |
multimodal |
|
|
| 20 |
Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence |
提出多轮认证鲁棒性框架以提升LLM安全性 |
large language model |
|
|
| 21 |
Profiling What Matters: Context-Aware Item Profiles from Large-Scale Metadata for LLM Recommenders |
提出CAIRO框架以解决LLM推荐中的项目侧信息利用问题 |
large language model |
|
|
| 22 |
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning |
提出DirEAG以解决数学推理中信心估计不可靠问题 |
large language model |
|
|
| 23 |
VortexChat: An agentic framework for autonomous multi-objective integrated photonic design |
提出VortexChat框架以实现自主的光子器件设计 |
large language model |
|
|
| 24 |
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents |
提出加权记忆树以解决长时间跨度LLM代理的记忆问题 |
large language model |
|
|