| 1 |
MARS: What Retrieval Signals Are Hidden in Multimodal Large Language Models for Text-Video Retrieval? |
提出MARS框架以解决文本-视频检索中的信息压缩问题 |
large language model multimodal |
✅ |
|
| 2 |
LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory |
提出LookStep以解决视觉语言导航中的资源效率问题 |
VLN large language model multimodal |
✅ |
|
| 3 |
Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment |
提出TIC-Bench以解决多模态模型评估中的文本-图像深度交互问题 |
multimodal |
✅ |
|
| 4 |
Rendering-in-the-Loop: An Execution-Driven Agent for Interactive Web Development |
提出RILA以解决交互网页开发中的功能验证问题 |
large language model foundation model multimodal |
|
|
| 5 |
Morphology signal in whole slide image foundation models can automatically triage slides |
提出基于形态信号的模型以自动筛选全切片图像 |
foundation model |
|
|
| 6 |
ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding |
提出ShallowStream以解决流媒体视频理解中的计算开销问题 |
large language model multimodal |
✅ |
|
| 7 |
Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness |
研究表明多模态大语言模型系统性高估面部吸引力 |
large language model multimodal |
|
|
| 8 |
YesTrack: Referring Multi-Object Tracking via MLLM-based Yes/No Verification |
提出YesTrack以解决多目标跟踪中的语言表达匹配问题 |
large language model multimodal |
✅ |
|
| 9 |
Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image Understanding |
提出轻量适应方法以解决多光谱和SAR图像理解问题 |
foundation model instruction following |
|
|
| 10 |
Detecting Object Hallucinations in Large Vision-Language Models via Cross-Modal Attention Drifts and Mask-Based Verification |
提出CADMP框架以解决大规模视觉-语言模型中的对象幻觉问题 |
visual grounding |
|
|
| 11 |
Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework |
提出因果驱动评估框架以解析VLM生成过程 |
multimodal |
|
|