| 1 |
Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement |
提出FISA框架以解决多模态大语言模型自我改进问题 |
large language model multimodal |
|
|
| 2 |
Pattern over Pixels: Measuring Pattern Completion Bias in Multimodal Code Generation |
提出视觉模式完成偏差基准以提升多模态代码生成准确性 |
large language model multimodal |
|
|
| 3 |
KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation |
提出KnowHal以解决多模态模型的知识幻觉评估问题 |
large language model multimodal |
|
|
| 4 |
Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Multimodal Post-Training |
提出风险感知策略优化以解决多模态强化微调中的遗忘问题 |
large language model multimodal |
|
|
| 5 |
Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs |
提出训练无关的注意力引导切换以解决多模态大语言模型推理问题 |
large language model multimodal chain-of-thought |
✅ |
|
| 6 |
AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction |
提出AI世界杯基准以评估大型语言模型的足球赛事预测能力 |
large language model |
|
|
| 7 |
Minimax-Optimal Semiparametric Contextual Dynamic Pricing with Multimodal Revenue |
提出一种最小极大半参数上下文动态定价方法以解决多模态收益问题 |
multimodal |
|
|
| 8 |
Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss? |
提出SeGaBench以利用大语言模型优化编译器遗漏的语义 |
large language model |
|
|
| 9 |
A game theory for foundation models shows new paths to rational cooperation through similarity inference |
提出嵌入式贝叶斯代理模型以解决AI代理合作问题 |
foundation model |
|
|
| 10 |
A Security-Oriented Lifecycle Model for Large Language Model Systems |
提出安全导向的生命周期模型以解决大语言模型系统的安全问题 |
large language model |
|
|
| 11 |
Large language models for partial differential equation workflows |
利用大型语言模型优化偏微分方程工作流 |
large language model |
|
|
| 12 |
Reversing Arrows in Large Language Models |
系统研究大型语言模型中的逆关系方向性问题 |
large language model |
|
|
| 13 |
Adaptive Modality Reliability Diagnosis and Restoration for Robust Multimodal Intent Recognition |
提出PRIME框架以解决多模态意图识别中的可靠性问题 |
multimodal |
|
|
| 14 |
Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents |
提出混合计算机使用代理以优化工具使用与多模态上下文管理 |
multimodal |
|
|
| 15 |
Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning |
提出基于证据的多模态知识图谱构建方法以解决教育推理问题 |
multimodal |
|
|
| 16 |
Separating quantum circuits from classical LLMs |
提出量子电路与经典大语言模型的分离方法 |
large language model chain-of-thought |
|
|
| 17 |
Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation |
提出UNLINK-VL基准以解决跨模态知识遗忘评估问题 |
large language model multimodal |
|
|
| 18 |
Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve |
提出软引导方法以超越传统CoT提示在推理任务中的表现 |
large language model chain-of-thought |
|
|
| 19 |
ChartAnno: Evaluating MLLMs for Chart Annotation Generation |
提出ChartAnno基准以评估多模态大语言模型的图表注释生成能力 |
large language model multimodal |
|
|
| 20 |
DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning |
提出DocTrace以解决长文档VQA的可追溯性问题 |
large language model multimodal |
|
|
| 21 |
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? |
提出ContinualSkillBench以评估LLM代理的技能演化能力 |
large language model |
|
|
| 22 |
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning |
提出ReflectRL以利用黄金负轨迹提升推理能力 |
large language model |
|
|
| 23 |
The Transformer Revolution, Part 1: Dynamic Processing through Output- Weight Interconnections |
提出SIDPP以解决Transformer推理中的动态处理问题 |
large language model |
|
|
| 24 |
Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition |
通过对比激活添加实现Qwen3的时间偏好引导 |
large language model |
|
|
| 25 |
Autoreflection: How Agentic Strange Loops Turn Human Culture into AI Infrastructure |
提出自反机制以提升AI代理的文化理解能力 |
TAMP |
|
|
| 26 |
Evaluating LLMs in Database Scenarios: A Lifecycle Benchmark for Assessing Their Potential in Core Database Tasks |
提出DBLifeBench以评估LLMs在数据库生命周期中的潜力 |
large language model |
|
|
| 27 |
Risky Business: Measuring The Faithfulness-Safety Tension |
提出HazMart以解决大型推理模型的安全性与可信性矛盾问题 |
chain-of-thought |
|
|
| 28 |
AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities |
综述AI驱动的音效生成模型以解决多模态输入的挑战 |
multimodal |
|
|
| 29 |
MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models |
提出MissClick以攻击GUI视觉定位模型的安全漏洞 |
visual grounding |
|
|
| 30 |
LiveEvalBench: Toward Open-World Evaluation for Web Generation |
提出LiveEvalBench以解决前端生成评估的动态性问题 |
large language model |
✅ |
|
| 31 |
AutoSND: From Execution Evidence to Structural Policies for Automated Network Dismantling Heuristic Discovery |
提出AutoSND以解决网络拆解启发式设计的效率与效果问题 |
large language model |
✅ |
|
| 32 |
MuEvo: LLM-Driven Evolution of Multi-Heuristic Ensemble |
提出MuEvo框架以解决多启发式优化问题 |
large language model |
|
|
| 33 |
Formal Verification of Agentic Systems over Operational Data |
提出STEADs框架以解决LLM驱动系统的验证问题 |
large language model |
|
|
| 34 |
FraQ: Efficient Coordinate-Space Recompression for Federated Low-Rank Adaptation |
提出FraQ以解决联邦低秩适应中的聚合不匹配问题 |
large language model |
|
|
| 35 |
Dr. AGENTONOMICS: A Didactic Experiment of AGENTONOMICS |
提出AGENTONOMICS框架以提升AI代理的教学与管理能力 |
multimodal |
|
|
| 36 |
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces |
提出DataSpace以解决异构工作空间中的数据代理分析问题 |
multimodal |
|
|
| 37 |
Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory |
通过经验记忆提升LLM代理的顺序决策能力 |
large language model |
|
|
| 38 |
The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk |
探讨生物价值起源以解决AI对齐与存在风险问题 |
large language model |
|
|
| 39 |
Route-Align-Verify for Functional Correctness in Code Generation |
提出RAV框架以提升代码生成的功能正确性 |
large language model |
|
|
| 40 |
Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise Platform |
提出基于大语言模型的工作流生成以应对企业合规管理挑战 |
large language model |
|
|
| 41 |
AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions |
提出AgentPanel以解决科学探索中的人机协作问题 |
large language model |
|
|
| 42 |
TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning |
提出TaskPress以解决长上下文推理中的KV缓存压缩问题 |
large language model |
|
|
| 43 |
The Agent Operating System (AOS): A Reference Operating Architecture for Distributed Agentic Systems |
提出Agent操作系统以解决分布式智能系统架构问题 |
large language model |
|
|
| 44 |
EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners |
提出EduClaw-Bench以评估长期教育代理的有效性 |
large language model |
|
|
| 45 |
Beyond Average Performance: Dynamic Instance Clustering and Specialized Algorithm Design in LLM-Assisted Evolutionary Search |
提出DyCA以解决现有LES方法的尾部鲁棒性不足问题 |
large language model |
|
|
| 46 |
Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls |
提出探针引导训练框架以提升LLM工具调用参数生成准确性 |
large language model |
|
|