| 1 |
Just Noticeable Difference Modeling for Token Compression in Vision-Language-Action Models |
提出Action-JND以优化视觉语言行动模型中的令牌压缩 |
vision-language-action VLA foundation model |
|
|
| 2 |
A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving |
提出基于VLA的多模态协作交互以解决自主驾驶决策不可靠问题 |
vision-language-action VLA multimodal |
|
|
| 3 |
Recognition-Conditioned Reasoning: A Training-Free Multimodal-LLM Pipeline for Fine-Grained Micro-Action Understanding |
提出无训练的多模态LLM管道以解决微动作理解问题 |
large language model multimodal |
|
|
| 4 |
EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking |
提出EviRank以解决多模态图像重排序问题 |
multimodal chain-of-thought |
|
|
| 5 |
CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models |
提出CertVLA以解决视觉-语言-动作模型的物理攻击问题 |
vision-language-action VLA |
|
|
| 6 |
Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation |
提出Vis-Poison以解决多模态检索增强生成中的视觉知识中毒问题 |
large language model multimodal |
✅ |
|
| 7 |
Masking Is Not Enough: Generative Restoration for Multimodal De-Identification in Medical AI |
提出ClinX框架以解决医疗AI中的多模态去标识化问题 |
multimodal |
|
|
| 8 |
Latent Ordinal Evidence, Misaligned Outputs: Inference-Time Ordinal Lens Alignment for Multimodal LLMs |
提出Ordinal Lens Alignment以解决多模态LLMs的输出对齐问题 |
multimodal |
|
|
| 9 |
Kinematic Knowledge Maps for Pattern Alignment: Structured Latent Representational Learning in Multimodal Gait Analysis |
提出ScoliDetect以解决多模态步态分析中的对齐问题 |
multimodal |
|
|
| 10 |
SuppreSensing: Expert-Guided Feature Recalibration and Discrepancy Augmentation for Multimodal Object Detection |
提出SuppreSensing以解决多模态遥感目标检测中的语义异质性问题 |
multimodal |
|
|
| 11 |
EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue |
提出EmotionDialogCN以解决多模态对话数据集情感标注不足问题 |
multimodal |
|
|
| 12 |
MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model |
提出MV2GF以解决多视角行人检测中的几何泛化问题 |
foundation model |
|
|
| 13 |
Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds |
提出轻量级视觉支架以提升视觉语言模型空间推理能力 |
multimodal visual grounding |
|
|
| 14 |
MigrationNarrate: A Dataset for Detection of Migration Narratives in YouTube Videos |
提出MigrationNarrate数据集以解决移民叙事检测问题 |
large language model multimodal |
|
|
| 15 |
OccluRank: Controllable Occlusion-Aware Layout-to-Image Generation by Adding Just an Ordinal Rank |
提出OccluRank以解决布局到图像生成中的遮挡顺序问题 |
large language model multimodal |
|
|
| 16 |
AGIDefect-4K: A Richly Annotated Dataset for AI-Generated Image Defect Detection, Localization and Explanation |
提出AGIDefect-4K数据集以解决AI生成图像缺陷检测问题 |
large language model multimodal |
✅ |
|
| 17 |
OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs |
提出OmniAssistBench以解决交互式视频助手评估问题 |
large language model |
|
|
| 18 |
KoViDoRe: Korean Visual Document Retrieval |
提出KoViDoRe以解决韩文视觉文档检索问题 |
multimodal |
|
|
| 19 |
Identity-Preserving Text-to-Video Generation via Agentic Enhancement and Semantic Repair |
提出AESR框架以解决身份保留视频生成中的问题 |
instruction following |
✅ |
|