| 1 |
Said Aloud, Read Different: Cross-Modal Instability in Multimodal Models |
提出跨模态不稳定性基准以解决多模态模型一致性问题 |
foundation model multimodal |
|
|
| 2 |
Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding |
提出基于视觉对齐的方法以解决跨语言语音映射问题 |
multimodal visual grounding |
|
|
| 3 |
Dependency-Aware Revocable Decoding for Efficient Diffusion Large Language Model Inference |
提出依赖感知可撤销解码以提升扩散大语言模型推理效率 |
large language model multimodal |
|
|
| 4 |
SCIT: Testing Causal Cache Carriers in Latent Chain-of-Thought Models |
提出SCIT以解决潜在链式思维模型中的因果缓存问题 |
chain-of-thought |
|
|
| 5 |
Prediction of Prediction (PoP): Inter-Layer Activation Fusion for Single-Pass Hallucination Detection in Large Language Models |
提出PoP机制以解决大型语言模型的事实错误检测问题 |
large language model |
|
|
| 6 |
RuleWeaver: Benchmarking Rule-Centered Scenario Reasoning for Large Language Models |
提出RuleWeaver以解决大型语言模型的规则中心场景推理问题 |
large language model |
✅ |
|
| 7 |
Benchmarking Clinical Decision Pathway Adherence in Large Language Models |
提出MEGA-CDP以评估医疗大语言模型的临床决策路径遵循性 |
large language model |
|
|
| 8 |
Surgical Alignment in Knowledge Graph Training for Clinical Diagnosis with Large Language Models |
提出KG信号整合方法以提升临床诊断中的LLM推理能力 |
large language model |
✅ |
|
| 9 |
Double Trouble: Bilingual Pretraining Leaves Language-Conditioned Effects in Shared-Language Representations |
揭示双语预训练模型共享语言表示的隐藏状态差异 |
language conditioned |
|
|
| 10 |
RCMN: Understanding Misleadingness in Influential Public Discourse |
提出RCMN框架以理解公共话语中的误导性问题 |
foundation model multimodal |
|
|
| 11 |
INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment |
提出INTENT-AS-A-TOOL以解决代理性不一致问题 |
large language model chain-of-thought |
✅ |
|
| 12 |
DocTalkBN: A Novel Dataset of Expert Telemedicine Conversations in Bengali |
提出DocTalkBN数据集以解决孟加拉语医疗对话数据匮乏问题 |
large language model multimodal |
|
|
| 13 |
Reasoning about In-Context Samples for Machine-Translation |
提出基于片段推理的机器翻译方法以提升翻译可靠性 |
large language model chain-of-thought |
|
|
| 14 |
How Language Models Organize and Structure Moral Knowledge |
提出线性探针以探讨语言模型中的道德知识结构 |
large language model |
|
|
| 15 |
Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference |
提出PRISM框架以解决大语言模型的人格一致性评估问题 |
large language model |
|
|
| 16 |
CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes |
提出CritICL框架以提升小型语言模型推理效率 |
large language model |
✅ |
|
| 17 |
BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing |
提出BALMS基准以解决长期心理健康感知问题 |
chain-of-thought |
|
|
| 18 |
JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols |
提出JudgeStealer以高效提取LLM评判能力 |
large language model |
|
|
| 19 |
Information-Guided Frontier Decoding: Contextual Utility-Driven Commitment in dMLLMs |
提出信息引导前沿解码以提升多模态语言模型的解码质量 |
multimodal |
|
|
| 20 |
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update |
提出反谄媚策略以平衡理性更新与支持性妥协问题 |
large language model |
|
|
| 21 |
When Tokenizers Fail: Byte-Level Chunking for Zero-Shot Transfer to Low-Resource Languages |
提出适应性层次网络框架以解决低资源语言处理问题 |
zero-shot transfer |
|
|
| 22 |
Surgical Alignment in Knowledge Graph Training for Clinical Diagnosis with Large Language Models |
提出KG信号整合方法以提升临床诊断中的LLM推理能力 |
large language model |
✅ |
|
| 23 |
Memorization Is Not Extraction: Tight Differential-Privacy Bounds and Audit Blind Spots |
提出精确的差分隐私界限以解决记忆与提取问题 |
large language model |
|
|
| 24 |
Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict |
研究音频-视觉LLM中的组合失败问题,提出新分析方法 |
instruction following |
✅ |
|
| 25 |
Load-Bearing Context: The Question Damage Score for Evaluating Context Reliance in Linguistic Reasoning |
提出问题损伤评分以评估语言推理中的上下文依赖性 |
large language model |
|
|
| 26 |
INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning |
提出INSPIRE以解决数学推理模型内化不足问题 |
large language model |
|
|