| 1 |
MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence |
提出MMArch基准以解决多模态推理在建筑工程中的挑战 |
large language model multimodal |
✅ |
|
| 2 |
GeoPhysAdapter: Scale-Matched Geophysical Adaptation for Cross-Domain Landslide Mapping with Vision Foundation Models |
提出GeoPhysAdapter以解决跨域滑坡映射中的错误警报问题 |
foundation model |
✅ |
|
| 3 |
SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge |
提出SafeSceneReason以解决工业安全推理问题 |
multimodal |
|
|
| 4 |
MELLON - Multimodal Enhanced LLM for Online Navigation |
提出MELLON以提升在线导航任务的多模态推理能力 |
multimodal |
|
|
| 5 |
CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits |
提出CircuitReason-1k基准以评估电路的视觉到符号推理能力 |
large language model multimodal |
|
|
| 6 |
Hallucination-Free GUI Grounding via Regression-Free Layout-Aware Matching |
提出无回归布局感知匹配以解决GUI定位问题 |
large language model multimodal |
|
|
| 7 |
CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation |
提出CADEngBench以评估CAD模型的工程行为 |
multimodal |
|
|
| 8 |
MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts |
提出MoRSE以解决多智能体系统中的任务细分与角色适应问题 |
large language model |
|
|
| 9 |
GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis |
提出GENCO以解决稳态电网分析中的多任务问题 |
foundation model |
|
|
| 10 |
Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics |
提出Avalon-ToM-Bench以评估细粒度的心智理论 |
chain-of-thought |
|
|
| 11 |
TSPORec: Token Selection via Preference Optimization for LLM-Based Sequential Recommendation |
提出TSPORec以解决LLM推荐系统中的信息损失问题 |
large language model |
✅ |
|
| 12 |
From Sweep to Seam: Interleaved Cross-Block Post-Training Quantization |
提出交错跨块后训练量化方法以提升模型压缩效果 |
large language model |
|
|
| 13 |
OpenLoopEvolve: A Verifiable Self-Evolution Framework for Loop Policies in Long-Horizon Complex Tasks |
提出OpenLoopEvolve框架以解决长时间复杂任务中的自我演化问题 |
large language model |
|
|
| 14 |
P$^{3}$: Joint Program-and-Proof Planning for Verified Code Generation |
提出P$^{3}$以解决验证代码生成中的效率与有效性问题 |
large language model |
|
|
| 15 |
SafeQL: Search-based Refinement for Safe and Efficient LLM-based Text-to-SQL |
提出SafeQL以解决LLM文本到SQL生成中的安全性与效率问题 |
large language model |
|
|
| 16 |
Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM Inference |
提出KVGov以解决多租户LLM推理中的时序侧信道泄露问题 |
large language model |
|
|
| 17 |
Agentic Router: An Execution-Grounded Continual Learning Approach With Memory |
提出执行基础的持续学习方法以提升CLI操作的成功率 |
large language model |
|
|
| 18 |
From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents |
提出奖励感知动态执行门以优化技能基础LLM代理的执行效率 |
large language model |
|
|
| 19 |
CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment |
提出CIDER数据集以解决隐私偏好对齐问题 |
large language model |
|
|
| 20 |
ChronoState: Hidden Elapsed-Time Conditioning for Temporal-State Action Selection in Frozen-Backbone Language Models |
提出ChronoState以解决语言模型中的时间状态选择问题 |
TAMP |
|
|
| 21 |
RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement |
提出RAVEN-Eval以解决AI视频生成模型评估难题 |
instruction following |
|
|
| 22 |
Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways |
提出跨层安全路径识别方法以解决多语言安全差距问题 |
large language model |
|
|