| 1 |
When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning |
提出提示区域对齐以解决多模态推理中的语义通道差距问题 |
large language model multimodal |
|
|
| 2 |
MarsCast: Transfer Learning of AI Weather Foundation Models to Planetary Atmospheres |
提出MarsCast以解决地球天气模型在火星大气预测中的适用性问题 |
foundation model |
|
|
| 3 |
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings |
提出隐性影响下的链式思维监控评估基准 |
chain-of-thought |
✅ |
|
| 4 |
Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load |
提出基于EU-AI法案的短期负荷预测方法以应对安全关键环境挑战 |
foundation model |
|
|
| 5 |
Hardware Design and Security in the Era of Chiplets and LLMs |
提出统一分析以应对芯片和大语言模型时代的硬件安全挑战 |
large language model |
|
|
| 6 |
The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations |
探讨种族与性别对警察语言使用的影响 |
large language model |
|
|
| 7 |
Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning |
提出梯度免疫机制以应对恶意微调问题 |
large language model |
✅ |
|
| 8 |
From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking |
提出审计化LLM框架以改进足球比分预测 |
large language model |
|
|
| 9 |
SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models |
提出SciCode-Verified以解决科学编码能力评估的缺陷问题 |
instruction following |
|
|
| 10 |
RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists |
提出RepoProbe以解决现有代码理解基准的不足问题 |
large language model |
|
|
| 11 |
Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning |
提出ReCo以解决长链推理中的缓存效率问题 |
chain-of-thought |
|
|
| 12 |
Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports |
提出Active-SWE以解决缺乏问题报告的主动修复挑战 |
large language model |
|
|
| 13 |
Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness |
提出Leak-Resistant Unlearning基准以评估多跳推理一致性与恢复鲁棒性 |
large language model |
|
|
| 14 |
CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models |
提出CARGO-VL以解决视觉语言模型中的反事实仲裁问题 |
multimodal |
|
|
| 15 |
ExeCRE: Execution-Consistency Guided Reliability Estimation for Self-Correcting Code Generation |
提出ExeCRE框架以解决自校正代码生成中的可靠性估计问题 |
large language model |
|
|
| 16 |
Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework |
提出SecureCollaRAG以解决知识腐蚀问题 |
large language model |
|
|