| 1 |
RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents |
提出RetailAgent框架以解决金融市场中的可预测性问题 |
large language model multimodal |
|
|
| 2 |
CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction Tuning |
提出CoRe-MoE以解决持续多模态指令调优中的参数冗余问题 |
large language model multimodal |
✅ |
|
| 3 |
When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI |
研究ASR错误对具身AI安全风险的影响 |
embodied AI |
|
|
| 4 |
Do Medical Vision Models Reason About Anatomy? Probing the Spatial Inductive Biases of Learned Visual Representations |
构建SPAR-Bench以探讨医学视觉模型的解剖推理能力 |
foundation model multimodal zero-shot transfer |
|
|
| 5 |
MAP: A Benchmark on Multimodal Accessibility Planning for Real World Places |
提出MAP基准以解决多模态无障碍规划问题 |
multimodal |
|
|
| 6 |
WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents |
提出WeAgent-MMSearch以解决多模态搜索代理中的视觉信息缺失问题 |
multimodal |
|
|
| 7 |
Automated Analysis Framework for Multilingual Climate-Health Literature Based on Multi-Agent Large Language Model |
提出多代理大语言模型框架以解决气候健康文献分析问题 |
large language model |
|
|
| 8 |
See, Hypothesize, Validate: Multimodal Agentic Framework for Discovering Governing PDEs |
提出MAGE框架以解决偏微分方程发现问题 |
multimodal |
|
|
| 9 |
The Shape of Power: A Multilingual Framework for Social Power Reasoning in Dialogues |
提出多语言框架以解决社会权力推理的跨文化分析问题 |
large language model multimodal |
|
|
| 10 |
Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code |
提出GenIaC-SecBench以评估LLM生成的基础设施代码安全性 |
large language model chain-of-thought |
|
|
| 11 |
Learning a Size-Weight Frontier for Synthetic-Augmented Inference |
提出合成增强推断的大小-权重前沿以解决数据稀缺问题 |
large language model |
|
|
| 12 |
GRACE:Gradient-guided Coreset Selection for LLM Unlearning |
提出GRACE以解决大语言模型的遗忘与保留数据选择问题 |
large language model |
|
|
| 13 |
Stay Within Your Bounds: Distance-Guided Decoding for Guaranteed Context-Free Grammar Compliance |
提出基于推送自动机的解码框架以确保上下文无关文法合规性 |
large language model |
|
|
| 14 |
Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification |
提出基于zk-SNARK的审计框架以解决LLM隐私验证问题 |
large language model |
|
|
| 15 |
LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment |
提出基于大型语言模型的智能体以提升软件和系统安全性 |
large language model |
|
|
| 16 |
LongPIBench: A Long-Context Benchmark for Prompt Injection |
提出LongPIBench以解决长上下文中的提示注入攻击问题 |
large language model |
|
|
| 17 |
Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers |
提出分层LLM防御体系以解决模型安全性问题 |
large language model |
|
|
| 18 |
MAIL: Memory-driven, Adaptive, Incremental, and Literature-grounded Framework for Hypothesis Generation in Chemistry |
提出MAIL框架以解决化学假设生成中的知识导航问题 |
large language model |
|
|
| 19 |
Beyond Task-Only Matching: Personalized Skill Routing with Counterfactual Evaluation |
提出个性化技能路由方法以解决任务匹配不足问题 |
large language model |
|
|
| 20 |
REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features |
提出REINS以增强稀疏自编码器特征的拒绝能力 |
large language model |
|
|
| 21 |
CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks? |
提出CAITLYN以解决大语言模型的注入攻击防御问题 |
large language model |
|
|
| 22 |
When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent Systems |
提出K-GAT以解决多智能体系统中的协作拓扑生成问题 |
large language model |
|
|
| 23 |
AI Alignment through a Game-theoretic Lens: A Survey |
通过博弈论视角提出AI对齐方法以应对复杂人类价值问题 |
large language model |
|
|
| 24 |
FISGuard: Defending Against Membership Inference via Fixed Input Subspaces |
提出FISGuard以解决成员推断攻击问题 |
large language model |
|
|