| 1 |
Where Identity Lives: Localized, Retain-Free Identity Unlearning in Multimodal Large Language Models |
提出PAVA方法以解决多模态大语言模型中的身份信息删除问题 |
large language model multimodal |
|
|
| 2 |
MMDS-Bench: Benchmarking Multimodal Large Language Models on Dynamic Stance in Social Media Interactions |
提出MMDS-Bench以解决社交媒体动态立场分类问题 |
large language model multimodal |
|
|
| 3 |
Reading the News: Adapting Large Language Models to Swedish Journalism Through Continued Pre-Training |
通过继续预训练适应大型语言模型于瑞典新闻领域 |
large language model instruction following |
|
|
| 4 |
LCoT-GV: Graph Attention Networks for Verifying Long Reasoning Chains in Large Language Models |
提出LCoT-GV以验证大型语言模型的长推理链 |
large language model chain-of-thought |
|
|
| 5 |
When LLM Meets Tree Search: A Systematic View of Inference as Search in Large Language Models |
提出树搜索方法以优化大语言模型推理过程 |
large language model chain-of-thought |
|
|
| 6 |
Quantifying and Mitigating Korean Jamo-Level Typographical Vulnerabilities in Large Language Models |
提出Typo-Aware Chain-of-Thought以解决韩文Jamo级别的输入错误问题 |
large language model chain-of-thought |
|
|
| 7 |
ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation |
提出ImageEval 2026以解决阿拉伯多模态评估问题 |
multimodal |
|
|
| 8 |
Every Token Leaves a Ripple in the Stream of Thought: Eliciting Model-Internal Token Saliency for Chain-of-Thought Compression |
提出MIST以解决链式推理压缩中的令牌选择问题 |
chain-of-thought |
|
|
| 9 |
SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators? |
提出SwarmBench以评估大语言模型在多智能体编排中的能力 |
large language model |
|
|
| 10 |
Sequential Trajectories and Simultaneous Blending: Multi-Emotion Modeling for Instruction-Following TTS |
提出HybridEmo以解决多情感控制的TTS问题 |
instruction following |
|
|
| 11 |
Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts |
提出READI基准以解决间接言语行为理解问题 |
multimodal |
|
|
| 12 |
GPAgentBench-2K: Benchmarking Large Language Model Agents in Complex Clinical Action Space |
提出GPAgentBench-2K以解决临床决策中的行动空间限制问题 |
large language model |
|
|
| 13 |
BiG-SURE - Bipartite Graph for Semantic Uncertainty and Reliability Estimation of LLMs |
提出BiG-SURE以解决LLMs的不确定性估计问题 |
large language model multimodal |
|
|
| 14 |
Evidence-Bounded Mental Health Reasoning from Heterogeneous Speech Protocols |
提出基于证据的心理健康推理框架以解决多模态数据的有效性问题 |
multimodal chain-of-thought |
|
|
| 15 |
GUIDE: Guiding Internal Evidence with Language Instructions |
提出GUIDE框架以调控多模态模型的内部证据使用 |
multimodal instruction following |
|
|
| 16 |
OCR-MetaReasoning Benchmark: Evaluating the Meta-Reasoning Ability of MLLMs in Text-Rich Image Understanding |
提出OCR-MetaReasoning基准以评估MLLMs在文本丰富图像理解中的推理能力 |
large language model multimodal |
✅ |
|
| 17 |
Reactivating Test-Time Scaling for Plane Geometry Problem Solving |
提出多轨合成方法以解决平面几何问题的推理挑战 |
multimodal visual grounding |
✅ |
|
| 18 |
SocialReasonBench: A Video-QA Benchmark for Social Reasoning with Counterfactual Narrative Videos |
提出SocialReasonBench以解决视频社交推理评估问题 |
multimodal |
|
|
| 19 |
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions |
提出MineAmongUs以解决多模态社交互动中的欺骗问题 |
multimodal |
|
|
| 20 |
TopoCompress: Long Context Compression via Graph-Wired Semantic Trajectories |
提出TopoCompress以解决长上下文压缩问题 |
large language model |
|
|
| 21 |
An Agentic Retrobiosynthesis Framework with Learned Frontier Selection |
提出基于学习的边界选择框架以优化反向生物合成 |
large language model |
|
|
| 22 |
Two Centuries of Sexism in British Parliament: A Computational Analysis of Women's Representation in the Hansard Corpus |
通过计算分析揭示英国议会两百年的性别歧视现象 |
large language model |
|
|
| 23 |
Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation Steering |
提出CREST以解决专用LLM推理时安全性问题 |
large language model |
✅ |
|
| 24 |
Using Prosody to Predict Syntactic Structure |
提出信息论框架以量化韵律与句法结构的关系 |
multimodal |
|
|
| 25 |
Calibrating Small Language Models for Claim Check-Worthiness Detection |
提出NN-PPI以解决小型语言模型的准确性问题 |
large language model |
|
|
| 26 |
Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Text |
提出隐蔽偏见注入方法以解决合成数据安全问题 |
large language model |
|
|
| 27 |
Enhancing Low-Resource Language Reasoning via High-Resource Language Feature Transfer |
通过高资源语言特征转移提升低资源语言推理能力 |
large language model |
|
|
| 28 |
Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs |
提出符号编辑框架以评估LLMs的逻辑推理能力 |
large language model |
|
|
| 29 |
The Differential Reasoning Router: Operationalizing Cost-Aware LLM Annotation in E-commerce |
提出差异推理路由器以解决电商冷启动标注问题 |
large language model |
|
|
| 30 |
CPR for LLMs: Critical-Point Routing against Catastrophic Forgetting in Domain Adaptation |
提出CPR以解决大语言模型领域适应中的灾难性遗忘问题 |
large language model |
|
|
| 31 |
Low-Resource Preference Adaptation of LLMs via Activation-Based Label Propagation |
提出激活基于标签传播的方法以解决低资源偏好适应问题 |
large language model |
✅ |
|
| 32 |
You Shouldn't Have Asked: A Pragmatics-Inspired Taxonomy for Evaluating LLM Refusals |
提出基于语用学的分类法以评估大型语言模型的拒绝行为 |
large language model |
|
|
| 33 |
Error-Type-Aware Loss Reweighting for Robust Named Entity Recognition with Noisy LLM Labels |
提出错误类型感知损失重加权以解决NER中的噪声标签问题 |
large language model |
|
|
| 34 |
CLIN: an Objective Framework for Evaluating Creativity in Short Persian Literary Text |
提出CLIN框架以评估波斯短文学文本的创造力 |
large language model |
|
|
| 35 |
MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines |
提出Murano框架以解决机制可解释性研究的整合问题 |
large language model |
|
|
| 36 |
More Capable, Less Faithful: A Multilingual Analysis of Mathematical (Un)Solvability Detection in LLMs |
提出多语言基准以解决大语言模型数学可解性检测问题 |
large language model |
|
|
| 37 |
From Final Artifacts to Trajectories: Retrospective Process Supervision for Evidence-Grounded Long-Form Generation |
提出RetroGen框架以解决开放任务中的轨迹数据不足问题 |
large language model |
|
|
| 38 |
Learning to Reason and Use Tools through Unsupervised Fine-Tuning in Task-Oriented Dialog Systems |
通过无监督微调提升任务导向对话系统的推理与工具使用能力 |
large language model |
|
|
| 39 |
AIA$^{2}$: Attribute-Agnostic Imbalance Augmentation for Subgroup Robustness |
提出AIA²框架以解决子群体鲁棒性问题 |
large language model |
✅ |
|
| 40 |
When Errors Become Memories: Causal Pathway Tracing in Multi-Turn Memory-Augmented LLMs |
提出基于结构因果模型的框架以解决记忆增强LLM中的错误传播问题 |
large language model |
|
|
| 41 |
ALTSTEER: Selective Safety Steering for Moving Beyond Hard Refusals to Constructive Alternatives |
提出ALTSTEER以解决安全引导中的拒绝问题 |
large language model |
|
|
| 42 |
CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents |
提出CAST框架以解决长时间工具调用代理的可靠性问题 |
large language model |
|
|
| 43 |
Verification-Aware Training for Speculative Decoding |
提出验证感知训练以提升推测解码效率 |
large language model |
✅ |
|