| 1 |
CytoFormer: A Molecularly Supervised Cell Foundation Model for Histopathology Cell Classification |
提出CytoFormer以解决细胞分类中的标注效率问题 |
foundation model |
|
|
| 2 |
MIRROR: Multimodal Intelligent Radiology Reasoning and Observation Reporter |
提出MIRROR以解决放射科报告生成中的信息失真问题 |
multimodal |
|
|
| 3 |
Beyond Accuracy: Assessing Calibration of Geospatial Foundation Models and Their Sensitivity to Distribution Shifts |
提出评估地理基础模型校准与分布变化敏感性的方法 |
foundation model |
|
|
| 4 |
CoM$^3$eT: A foundation model for medical image analysis through federated, multidimensional context integration |
提出CoM$^3$eT以解决医学图像分析中的多任务学习问题 |
foundation model |
|
|
| 5 |
AnchorScore: A CLIP-Based Diagnostic of MLLM Annotation Difficulty |
提出AnchorScore以解决MLLM注释难度评估问题 |
large language model multimodal |
|
|
| 6 |
HarmTrace: Anchor-Calibrated Decoupled Optimization for Fine-Grained Target Identification in Harmful Memes |
提出HarmTrace以解决有害表情包中的细粒度目标识别问题 |
large language model multimodal |
✅ |
|
| 7 |
Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans |
比较多模态大语言模型与人类视觉搜索的差异 |
large language model multimodal |
|
|
| 8 |
MLLM-Guided Semantic Correction for Text-to-Video Generation |
提出MLLM引导的语义修正方法以解决文本到视频生成中的语义错误问题 |
large language model multimodal |
|
|
| 9 |
Remote-Sensing City Layout Extraction with MLLM |
提出基于多模态大语言模型的城市布局提取方法 |
large language model multimodal |
|
|
| 10 |
PersonaShot: Benchmarking Person-Centric Narrative Continuity in Multi-Shot Video Generation |
提出PersonaShot以解决多镜头视频生成中的叙事连续性问题 |
multimodal |
|
|
| 11 |
Self-Routed Tensor Adapters for Parameter-Efficient Universal Visual Adaptation |
提出自路由张量适配器以解决多领域视觉适应问题 |
foundation model |
✅ |
|
| 12 |
Seeing Before Answering: Training-Free Visual Layer Profiling for Vision-Language Models |
提出视觉层剖析方法以优化视觉语言模型性能 |
multimodal |
|
|
| 13 |
Representation Is Not Enough: Body-Localized Thermal Evidence for Contactless Stress and Craving Sensing in Opioid Use Disorder |
提出FABLE-Therm以解决无接触压力与渴望感知问题 |
foundation model |
|
|