| 1 |
OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora |
提出OmniPhys以解决物理领域多模态基准缺乏问题 |
large language model multimodal |
✅ |
|
| 2 |
PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans |
提出PlanSightRAG以解决土木标准图纸合规检查问题 |
multimodal |
|
|
| 3 |
Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty |
提出Think-Probe-Respond以解决大语言模型判断研究创意新颖性的问题 |
large language model |
|
|
| 4 |
Distinct dynamics of conceptual and referential disruptions in human reading and large language model processing |
研究概念与指称信息在阅读中的动态差异 |
large language model |
|
|
| 5 |
Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study |
比较生成与编码大型语言模型在ASR评估中的应用 |
large language model |
|
|
| 6 |
Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence |
揭示大语言模型评估中的锚定偏差问题 |
large language model chain-of-thought |
|
|
| 7 |
Skill Issue: Are Skills Language-Invariant in LLMs? |
量化语言模型跨语言技能不一致性以提升多语言性能 |
large language model |
|
|
| 8 |
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution |
提出JIT-Agent以解决智能体工具设计的可扩展性问题 |
foundation model |
|
|
| 9 |
Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Language Models |
提出基于束搜索和自一致性的文本到SQL转换方法以优化小型语言模型 |
large language model |
|
|
| 10 |
Controllable Affective Generation via Latent Vector Steering |
提出EmoVec以解决大语言模型情感生成不足问题 |
large language model |
|
|
| 11 |
Adaptive Triggering for Bias Correction in LLM Reasoning |
提出自适应触发机制以解决LLM推理中的偏见问题 |
chain-of-thought |
|
|
| 12 |
A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks |
提出自演化多智能体框架以防御LLM越狱攻击 |
large language model |
|
|
| 13 |
Unveiling Spectral Mechanisms in Training-Free LLM Text Detection |
提出频谱分析方法以解决训练无关的文本检测问题 |
large language model |
|
|
| 14 |
From Passive Response to Proactive Correction: Enhancing LLM Robustness Against Input Fact Perturbations |
提出DEDUCE框架以增强LLM对输入事实扰动的鲁棒性 |
large language model |
|
|
| 15 |
Localize-Then-Decide Guarantees for LLM Judgments |
提出Localize-Then-Decide框架以解决LLM判断一致性问题 |
large language model |
|
|
| 16 |
GRIP: Granular Reward-Guided Parameter Interpolation for Efficient Reasoning |
提出GRIP框架以实现高效推理的参数插值 |
large language model |
|
|
| 17 |
Cross-Dataset Stability of Expert-Informed Skill Prompting and Fine-Tuning for Chinese Metaphor Identification |
提出专家指导的技能提示以提升中文隐喻识别的跨数据集稳定性 |
large language model |
|
|
| 18 |
VietAIDetector: An Open-Source Zero-Shot Detector for Vietnamese AI-Generated Text |
提出VietAIDetector以解决越南语AI生成文本检测问题 |
large language model |
✅ |
|
| 19 |
Provenance Before Prose: Claim-Locked Reporting |
提出Claim-Locked Reporting以解决统计报告的可重复性问题 |
large language model |
|
|
| 20 |
Groundhog Bit-Flip Attack: Seeding Infinite Generation Loops in Mixture-of-Experts LLMs through Bit Flips |
提出Groundhog Bit-Flip攻击以揭示MoE LLM的脆弱性 |
large language model |
|
|