| 1 |
M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models |
提出M-GATE基准以评估多语言模型的语法和翻译能力 |
large language model |
|
|
| 2 |
VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs |
提出VIBE基准以解决大语言模型输出的情感分析问题 |
large language model |
|
|
| 3 |
GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models |
提出GPTKB 2.0以解决知识库构建中的歧义问题 |
large language model |
|
|
| 4 |
Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili |
分析大型语言模型中的跨语言偏见以提升多语言安全性 |
large language model |
|
|
| 5 |
DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models |
提出DUD框架以解决大语言模型的不确定性量化问题 |
large language model |
|
|
| 6 |
On the Diversity of Analogy Making in Large Language Models |
评估大型语言模型类比生成的多样性以促进创新 |
large language model |
|
|
| 7 |
Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining |
提出可扩展的子文档去重框架以解决大规模预训练中的冗余问题 |
large language model |
|
|
| 8 |
Activation-Guided Neuron Intervention to Induce Alzheimer's-Related Computational Language Phenotypes in a Large Language Model |
提出激活引导神经元干预以揭示阿尔茨海默病相关语言特征 |
large language model |
|
|
| 9 |
Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models |
提出多维评估框架以提升大型语言模型的统计推理能力 |
large language model |
|
|
| 10 |
Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling |
提出层级分析方法以解决语言模型对扰动输入的鲁棒性问题 |
chain-of-thought |
|
|
| 11 |
Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension |
提出多种初始化策略以优化大语言模型词汇扩展 |
large language model |
|
|
| 12 |
MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning |
提出MultiGlobeQA以解决地理空间推理的多语言基准问题 |
large language model |
|
|
| 13 |
Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking |
提出源条件描述长度增益以解决生成性抄袭检测问题 |
large language model |
|
|
| 14 |
How Closely Do LLM Reviews Align with Human Peer Review? |
比较大型语言模型与人类评审的一致性 |
large language model |
|
|
| 15 |
Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation |
提出区分模拟与估计以优化语言模型的人类意见模拟 |
large language model |
|
|
| 16 |
SocietyBench: Forecasting Counterfactual Social-World Evolution |
提出SocietyBench以评估社会事件预测能力 |
large language model |
|
|
| 17 |
WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament |
提出一种无泄漏的前沿LLM评估方法以预测世界杯赛事 |
large language model |
|
|
| 18 |
ConlangBench: Exploring Language Knowledge and Learning in LLMs through Diverse Constructed Languages |
提出ConlangBench以评估大语言模型在构造语言学习中的表现 |
large language model |
|
|
| 19 |
Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference |
提出Alien Abduction游戏以研究LLMs在推理中的信息获取能力 |
large language model |
|
|
| 20 |
Relational Priors as Convergence Pressure in LLM-Based Multi-Agent Systems |
提出关系先验作为LLM-MAS中的收敛压力以优化多智能体系统 |
large language model |
|
|
| 21 |
ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization |
提出ICO框架以优化语义转移越狱攻击效果 |
foundation model |
|
|
| 22 |
What Language Does and What the Evidence Supports: A Functional Role Taxonomy and Evidence Audit of Language Grounding in Embodied Agents |
提出语言功能角色分类以评估语言在具身智能体中的作用 |
foundation model |
|
|
| 23 |
SeqLLM: Augmenting LLMs with Behavioral-Sequence Modeling for High-Stakes Decisions at WeChat Pay |
提出SeqLLM以解决微信支付商户风险控制问题 |
large language model |
|
|