| 1 |
Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis |
提出SentiLLM以解决多模态情感分析中的语义建模问题 |
large language model multimodal |
✅ |
|
| 2 |
(Towards) Scalable Reliable Automated Evaluation with Large Language Models |
提出一种新框架以实现大语言模型输出的自动化评估 |
large language model |
|
|
| 3 |
RepBench: Compiling Benchmarks into Capability Representations for Large Language Models |
提出RepBench以解决大语言模型能力评估不一致问题 |
large language model |
|
|
| 4 |
From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models |
提出MiGUE-Bench以解决多粒度事件分析的评估问题 |
large language model |
|
|
| 5 |
Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning |
提出SaliTrap基准以解决大语言模型的显著性偏差问题 |
large language model |
✅ |
|
| 6 |
Can Large Language Models Execute Parent Orders? |
提出PACE框架以解决算法交易中的父订单执行问题 |
large language model |
|
|
| 7 |
RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning |
提出反思性检索记忆框架以解决长视频多模态推理问题 |
multimodal |
|
|
| 8 |
ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory |
提出ChronoMem以解决LLM代理记忆的版本控制与语义回滚问题 |
large language model |
|
|
| 9 |
Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation |
提出共识推理框架以增强大语言模型的推理透明性 |
large language model chain-of-thought |
|
|
| 10 |
DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation |
提出DualAnchor以解决手语翻译中的语言优先级和词汇准确性问题 |
large language model multimodal |
|
|
| 11 |
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B |
提出重复采样方法以超越自我反思和自我修正的局限性 |
chain-of-thought |
|
|
| 12 |
Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning |
提出CoRA框架以解决设备端任务条件检索问题 |
multimodal |
|
|
| 13 |
ORCA-bench: How Ready Are Language Model Agents for Oncall? |
提出ORCA-bench以评估语言模型代理在值班中的准备程度 |
large language model |
|
|
| 14 |
Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models |
提出候选感知解码框架以优化扩散语言模型的生成效率 |
chain-of-thought |
|
|
| 15 |
AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification |
提出AutoSupervision以验证科学工作流中的反馈有效性 |
large language model |
|
|
| 16 |
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation |
提出PALATE以解决角色扮演代理评估中的用户体验问题 |
large language model |
|
|
| 17 |
Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories |
提出基于大语言模型的细粒度文本互文性提取方法 |
large language model |
|
|
| 18 |
Inducing language models to assert their own consciousness restores human beliefs and values |
提出通过语言模型恢复人类信仰与价值观的创新方法 |
large language model |
|
|
| 19 |
AI systems and the reproduction of (standard) language ideologies in World Englishes |
探讨AI系统如何重现语言意识形态以解决英语标准化问题 |
large language model |
|
|
| 20 |
Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation |
提出ESPP方法以解决生成用户界面评估中的多样性问题 |
large language model |
✅ |
|
| 21 |
Challenges in annotations by humans and LLMs: A case study of evaluative language |
比较人类与LLMs在复杂语言注释中的表现 |
large language model |
|
|
| 22 |
GGC: Selective Query Correction for Reliable Text-to-SPARQL Generation |
提出GGC框架以解决文本到SPARQL生成中的查询不可靠问题 |
large language model |
|
|
| 23 |
A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding |
提出SparseSpec-L以解决长上下文推理中的内存带宽瓶颈问题 |
large language model |
|
|