| 1 |
Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models |
提出趋势感知修剪以解决多模态大语言模型的视觉令牌过滤问题 |
large language model multimodal |
|
|
| 2 |
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models |
提出MMOOC基准以解决多模态大语言模型的上下文评估问题 |
large language model multimodal |
|
|
| 3 |
Witness Evidence Portfolios: Single-Prefill Risk Detection for Closed Multimodal Answers |
提出见证证据组合以解决闭合多模态答案的风险检测问题 |
large language model multimodal |
|
|
| 4 |
MIND: Multimodal Intent-Driven Network via Diffusion Transformers for Medical Image Fusion |
提出MIND以解决医学图像融合中的意图驱动问题 |
multimodal |
|
|
| 5 |
Negative controls reveal volume-driven confounding in radiomics and imaging foundation model features |
提出READII-2-ROQC以解决影像组学中的体积驱动混淆问题 |
foundation model |
|
|
| 6 |
Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures |
提出病理基础模型作为有丝分裂图像检测编码器 |
foundation model |
✅ |
|
| 7 |
Benchmarking Foundation and Large Language Models for Few-Shot Medical Image Segmentation |
提出FAME基准以评估少样本医学图像分割方法 |
large language model |
|
|
| 8 |
DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis |
提出多样化架构以解决医学图像分析中的概念检测与标题生成问题 |
foundation model |
✅ |
|
| 9 |
FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval |
提出FiRE以解决复杂图像检索中的细粒度上下文建模问题 |
large language model multimodal |
|
|
| 10 |
LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference |
提出LAST框架以解决边缘云协作MLLM推理中的视觉令牌修剪问题 |
foundation model multimodal |
|
|
| 11 |
Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA |
提出Thinking-Once以解决高分辨率视觉问答中的证据获取问题 |
large language model multimodal |
|
|
| 12 |
LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA |
提出LoMeVQA以解决纵向医学视觉问答问题 |
large language model multimodal |
✅ |
|
| 13 |
ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation |
提出ROAD框架以降低3D形状生成的训练成本 |
foundation model |
✅ |
|
| 14 |
ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding |
提出ObjectStream以解决流媒体视频理解中的记忆管理问题 |
large language model |
|
|
| 15 |
Scaling Vision-Language Models Is Not Enough to Mitigate Bias |
大规模研究揭示视觉-语言模型偏见问题的复杂性 |
multimodal |
|
|
| 16 |
TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment |
提出几何锚定的舌头合成方法以解决面部重演中的舌头动态问题 |
foundation model |
|
|
| 17 |
MeshFM: 2D Features Are All You Need for 3D Shape Understanding |
提出MeshFM以解决3D形状理解中的2D特征利用问题 |
foundation model |
✅ |
|
| 18 |
Drawing-Recode: Annotation Grounding for Parametric CAD Code Generation from Raster 2D CAD Drawings |
提出Drawing-Recode以解决光栅2D CAD图纸的参数化CAD代码生成问题 |
large language model |
|
|