| 1 |
PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images? |
提出PathVU基准以解决病理图像多尺度理解问题 |
large language model multimodal |
|
|
| 2 |
IndustryForge-27B: A Domain-Enhanced Multimodal Foundation Model for Industrial CAD |
提出IndustryForge-27B以解决工业CAD设计自动化问题 |
foundation model multimodal |
|
|
| 3 |
IFHierBench: Hierarchical Instruction Following for Large Language Models |
提出IFHierBench以解决层次化指令遵循问题 |
large language model instruction following |
|
|
| 4 |
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications |
提出AISPA框架以审计大型语言模型应用中的系统提示 |
large language model foundation model |
|
|
| 5 |
DualG-MRAG: Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation |
提出DualG-MRAG以解决多模态检索增强生成中的复杂推理问题 |
multimodal |
|
|
| 6 |
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports |
提出EndoCLIP以解决内窥镜报告与图像关联不足问题 |
foundation model |
|
|
| 7 |
Asymmetric Communication: Large Language Models and Language Games |
提出不对称沟通框架以重新审视语言模型的属性 |
large language model |
|
|
| 8 |
Distilling Answer Set Programming Theories from Large Language Models |
提出神经符号方法从大型语言模型中蒸馏答案集编程理论 |
large language model |
|
|
| 9 |
MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware |
提出MMLDSum-LLM以解决多模态长文档摘要中的信息遗漏问题 |
multimodal |
|
|
| 10 |
Specification-Guided Synthesis of Deadlock-Free Communication Protocol Refinements with Large Language Models |
提出Syntropy框架以解决通信协议死锁问题 |
large language model |
|
|
| 11 |
One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs |
提出统一多语言多模态安全对齐框架以应对复合攻击问题 |
multimodal |
|
|
| 12 |
Guiding Large Language Models with Genetic Programming-Evolved Heuristic Knowledge for Dynamic Multi-Mode Project Scheduling |
利用遗传编程进化的启发式知识指导大型语言模型进行动态多模式项目调度 |
large language model |
|
|
| 13 |
Using Large Language Models for Idea Generation in Innovation |
利用大型语言模型提升创新产品构思的有效性 |
large language model |
|
|
| 14 |
MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos |
提出MMHBench以解决长视频心理健康理解问题 |
large language model multimodal |
|
|
| 15 |
Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation |
提出简单图像变换以突破现代AI内容审核系统的局限 |
foundation model multimodal |
|
|
| 16 |
Albilich: Steerable Proof-State Orchestration for LLM-Based Mathematical Research with CAS Integration |
提出Albilich以解决数学研究中的长远证明协调问题 |
large language model |
|
|
| 17 |
Strategy, Not Payoffs: A Behavioural Embedding of Normal-Form Games |
提出行为嵌入以解决LLM战略能力转移问题 |
large language model |
|
|
| 18 |
MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems |
提出MANTA框架以解决多智能体系统通信拓扑固定性问题 |
large language model |
|
|
| 19 |
A Fuzzy Rule-based Neuro-Symbolic Approach for Pipe Severity Prediction in Sewer Networks |
提出模糊规则神经符号方法以解决污水管道严重性预测问题 |
large language model |
|
|
| 20 |
HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection |
提出HyperClaim以解决视频虚假信息检测中的多模态推理问题 |
multimodal |
|
|
| 21 |
From Textual Requirements to Microservice Architectures - A Comprehensive Evaluation of LLM-Based Design Synthesis |
基于LLM的设计合成方法解决微服务架构识别问题 |
large language model |
|
|
| 22 |
MemHarness: Memory Is Reconstructed, Not Replayed |
提出MemHarness框架以重构而非重放记忆提升LLM智能体表现 |
large language model |
|
|
| 23 |
LLM-Guided Evolutionary Search for Constraint Model Reformulation to Improve Solver Efficiency |
提出基于LLM的进化搜索以优化约束模型重构 |
large language model |
|
|
| 24 |
Vibe-FDTR: An agent-oriented framework for reproducible frequency-domain thermoreflectance data analysis |
提出Vibe-FDTR框架以解决频域热反射数据分析的复杂性问题 |
large language model |
|
|
| 25 |
An Instrument to Evaluate Governance Proposals: AI Policy Analysis at Scale |
提出AI政策分析框架以评估治理提案 |
large language model |
|
|
| 26 |
DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness |
提出DataClawEval基准以评估工业数据工程代理的能力 |
large language model |
✅ |
|
| 27 |
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs |
提出可信凭证以解决大型语言模型安全性问题 |
large language model |
|
|
| 28 |
ARES: Adaptive Reasoning-Effort Steering for PPA- and Cost-Aware RTL Optimization with LLM Agents |
提出ARES以优化RTL设计的PPA和成本问题 |
large language model |
|
|
| 29 |
MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory |
提出MemTxn以解决持久内存更新和状态恢复问题 |
large language model |
|
|