| 1 |
Few-Shot Concept Prompt Learning for Segmentation Foundation Models via Visual Grounding |
提出少样本概念提示学习以解决医学图像分割问题 |
foundation model visual grounding |
|
|
| 2 |
Illuminating Visual Identity in Universal Multimodal Embeddings |
提出统一视觉身份辨别方法以提升多模态嵌入能力 |
large language model multimodal |
✅ |
|
| 3 |
Generative AI and Foundation Models in Medical Image |
探讨生成式AI与基础模型在医学影像中的应用 |
large language model foundation model |
|
|
| 4 |
UEmbed: Unified Sparse and Dense Multimodal Embeddings |
提出UEmbed以解决多模态稀疏与密集嵌入统一问题 |
multimodal |
|
|
| 5 |
Implicit Neural Representations for Multimodal Longitudinal Image Imputation and Interpolation |
提出条件隐式神经表示以解决多模态MRI图像插补问题 |
multimodal |
|
|
| 6 |
A General-Purpose VLM Can Teach an Astronomy Foundation Model to Better Recognize Galaxy Morphology |
提出通用VLM以提升天文学基础模型的星系形态识别能力 |
foundation model |
✅ |
|
| 7 |
Deep Multimodal Fusion Detection through Spatial Mask and Channel Fusion |
提出基于注意力驱动的互补重采样框架以提升跨模态目标检测 |
multimodal |
|
|
| 8 |
MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing |
提出MIEScore以解决多源图像编辑评估问题 |
large language model multimodal instruction following |
✅ |
|
| 9 |
SPECTRA: Band-Routed Embedding and Stage-Wise LoRA for Cross-Sensor Fine-Tuning of Geospatial Foundation Models |
提出SPECTRA以解决地理基础模型的光谱不匹配和适应成本问题 |
foundation model |
✅ |
|
| 10 |
FAU at ImageCLEF 2026 Task on Multimodal Reasoning Robust Candidate Scoring and Concise Multilingual Visual Answering |
提出多模态推理系统以解决视觉问答任务 |
multimodal |
|
|
| 11 |
Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning |
提出空间-频谱视觉锚学习以解决MLLM视觉感知退化问题 |
large language model foundation model multimodal |
✅ |
|
| 12 |
MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving |
提出MoRAL以解决边缘计算平台上VLM的空间推理问题 |
multimodal chain-of-thought |
|
|
| 13 |
ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs |
提出ET-Prune以解决文本丰富输入的视觉令牌修剪问题 |
large language model multimodal |
|
|
| 14 |
SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation |
提出SVGEval以解决SVG生成评估的可靠性问题 |
multimodal visual grounding |
|
|
| 15 |
Decoupling semantics from vision: A framework for faithful visual-text compression evaluation |
提出新评估框架以解决视觉文本压缩评估问题 |
large language model multimodal |
|
|
| 16 |
GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation |
提出GEOID-Flood以解决洪水分割数据集不足问题 |
foundation model |
✅ |
|
| 17 |
Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression |
提出基于消息的压缩方法以解决视觉语言模型的效率问题 |
multimodal |
|
|
| 18 |
Invisible Ink Threats: Adversarial Goals Behind Legitimate Tasks in Computer-Use Agents |
提出II-Bench以应对计算机使用代理中的隐形墨水威胁 |
large language model |
|
|
| 19 |
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks |
提出LongHorizon-Harness以解决长时间任务状态管理问题 |
large language model |
|
|
| 20 |
CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation |
提出CultureVidBench以解决文本到视频生成中的文化理解问题 |
multimodal |
|
|
| 21 |
Parameter-Dynamic Adaptive Fusion and Calibration Network for RGBT Tracking |
提出PAFCNet以解决RGBT跟踪中的参数固定问题 |
multimodal |
|
|
| 22 |
IDraw: Artist Verification from Digital Drawing Images |
提出IDraw以解决数字绘画作者验证问题 |
multimodal |
|
|
| 23 |
Learning Where to Look and How to Judge: Resolution-agnostic Image Quality Assessment with Quality-aware Saliency |
提出ReLIQS以解决无参考图像质量评估中的分辨率依赖问题 |
multimodal |
|
|
| 24 |
UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization |
提出UniSim-SLAM以解决SLAM中的几何不一致性问题 |
foundation model |
✅ |
|