| 1 |
SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models |
提出SepPrune框架以高效修剪多模态大语言模型 |
large language model multimodal |
|
|
| 2 |
Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition |
提出PRISM-AH框架以解决视频级别的模糊性与犹豫性识别问题 |
large language model multimodal |
|
|
| 3 |
Food Image Segmentation with LLM-Derived Ingredient Labels and Multimodal Fusion |
提出多模态模块以提升食品图像分割性能 |
large language model multimodal |
|
|
| 4 |
Towards Reliable Stain Transfer: An Iterative Data-Model Co-Optimization Framework Based on Multimodal Expert-Guided Assessment |
提出DMCoStain框架以解决异质生物标志物的染色转移问题 |
multimodal instruction following |
✅ |
|
| 5 |
VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening |
提出VetClaw以解决兽医疾病筛查问题 |
multimodal |
|
|
| 6 |
LaP-Forensics: Latent-Pixel Consistency Guided Multimodal Reasoning for Deepfake Detection |
提出LaP-Forensics以解决深度伪造检测中的视觉伪影问题 |
multimodal |
|
|
| 7 |
Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models |
提出CRoMa以解决病理基础模型的稳健性问题 |
foundation model |
|
|
| 8 |
Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation |
提出代理智能以解决医学领域多步骤临床任务的挑战 |
large language model foundation model multimodal |
|
|
| 9 |
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition |
提出CLBench-V以解决多模态上下文学习评估问题 |
multimodal |
✅ |
|
| 10 |
FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding |
提出FORGE以解决长视频理解中的信息选择问题 |
large language model multimodal |
|
|
| 11 |
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities |
提出MODUS以解决多模态建模的局限性问题 |
multimodal |
|
|
| 12 |
Fine-Grained Food Image Understanding via Target-Aware Data Alignment |
提出目标感知数据对齐方法以解决细粒度食品图像理解问题 |
multimodal |
|
|
| 13 |
Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications |
提出基于指令的图像编辑方法以解决现有编辑系统的局限性 |
large language model |
|
|
| 14 |
Visual prompt engineering for video models |
提出视觉提示工程以提升视频模型的推理性能 |
foundation model |
✅ |
|
| 15 |
Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation |
提出Argus-Unified以解决多模态理解与生成的高成本问题 |
multimodal |
|
|
| 16 |
CD-RMOT-Bench: Benchmarking the Cross-Domain Referring Multi-Object Tracking |
提出CD-RMOT-Bench以解决跨域语言引导多目标跟踪问题 |
language conditioned |
|
|