| 1 |
BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models |
提出BEAR-Bench以解决多模态模型在专业文档推理中的不足 |
large language model multimodal |
|
|
| 2 |
Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges |
提出多模态对话AI以解决持续交互中的记忆和上下文问题 |
multimodal |
✅ |
|
| 3 |
Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds |
提出几何分析方法以揭示大语言模型的推理机制 |
large language model |
|
|
| 4 |
Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study |
利用多模态语言模型审计TikTok有害内容的跨国研究 |
multimodal |
|
|
| 5 |
Effects of Answer Format Variation on Gender Bias in Large Language Models |
探讨回答格式变化对大型语言模型性别偏见的影响 |
large language model |
|
|
| 6 |
An Investigation of Translationese in the Generations of Multilingual Large Language Models |
研究多语言大语言模型生成的翻译特征以识别翻译语现象 |
large language model |
|
|
| 7 |
Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings |
提出Reflex-Guard以解决LLM提示安全性问题 |
large language model |
|
|
| 8 |
Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints |
审计冻结LLM中的解码生成控制差距以解决几何约束问题 |
large language model |
|
|
| 9 |
Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses |
提出人类可解释性分析框架以评估大型语言模型的潜在结构 |
large language model |
|
|
| 10 |
PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX |
提出PTXBench以优化GPU内核的LLM适配问题 |
large language model |
|
|
| 11 |
Chain-of-Experience for Continual LLM Improvement |
提出Chain-of-Experience以提升大语言模型的持续学习能力 |
large language model |
|
|
| 12 |
Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It |
提出基于表达方式的信念与事实处理方法以提升LLM性能 |
large language model |
✅ |
|
| 13 |
ArborMem: Navigating Interaction States with Memory Forests |
提出ArborMem以解决对话状态记忆管理问题 |
large language model |
|
|
| 14 |
Decomposition Attacks Across Unlinkable Identities: Limits of Stateful Defenses for LLM Services |
提出状态监控机制以应对分解攻击挑战 |
large language model |
|
|