| 1 |
From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology |
提出多分辨率金字塔变换器以解决病理图像分析中的分辨率限制问题 |
large language model foundation model multimodal |
|
|
| 2 |
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs |
提出ParVL框架以优化多模态大语言模型的计算分配 |
large language model multimodal |
✅ |
|
| 3 |
OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models |
提出OmniPack以解决多模态大语言模型的高计算开销问题 |
large language model multimodal |
|
|
| 4 |
LocAnyMed: Vision-Language Grounding for Multimodal Medical Images |
提出LocAnyMed以解决多模态医学图像的视觉语言定位问题 |
multimodal visual grounding |
✅ |
|
| 5 |
CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation |
提出CorePath以解决乳腺核心针活检诊断挑战 |
foundation model multimodal |
|
|
| 6 |
LDU-Bench: Multimodal LLM Evaluation for Lithography Defect Understanding under Layout-Varying Circuit Backgrounds |
提出LDU-Bench以解决光刻缺陷理解中的多任务评估问题 |
large language model multimodal |
|
|
| 7 |
Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis |
提出多模态框架以实现高效植物根系表型分析 |
multimodal |
|
|
| 8 |
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware |
提出视觉标记优化策略以加速多模态推理 |
multimodal |
|
|
| 9 |
Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding |
提出Hi-Token以解决视觉定位中的坐标表示问题 |
visual grounding |
|
|
| 10 |
Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models |
提出时间采样门控机制以解决视频大语言模型的偏差问题 |
large language model |
|
|
| 11 |
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models |
提出GSTEP以解决视频大语言模型中的冗余视觉标记问题 |
large language model |
|
|
| 12 |
SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models |
提出SlimVLM以解决视觉语言模型的高计算开销问题 |
large language model multimodal |
|
|
| 13 |
CIGTSurv: Clinical Information Guided Tri-modal Survival Prediction with Local Prototype Association and Global Feature Alignment |
提出CIGTSurv以解决临床信息在生存预测中的低利用问题 |
foundation model multimodal |
✅ |
|
| 14 |
UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space |
提出UHP检测以解决LVLM模型幻觉检测问题 |
multimodal |
✅ |
|
| 15 |
Attention is Case-Sensitive |
提出字母大小写敏感性以优化注意力分配 |
large language model |
|
|
| 16 |
MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification |
提出MT-Web2Code以解决多轮区域重建与局部修改问题 |
multimodal |
|
|
| 17 |
Frequency-Decorrelated Temporal Ensembles for EEG--fNIRS Imagined-Handwriting Decoding |
提出FRED系统以解决EEG-fNIRS想象手写解码问题 |
multimodal |
✅ |
|