RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition
📰 ArXiv cs.AI
arXiv:2403.13805v2 Announce Type: replace-cross Abstract: CLIP (Contrastive Language-Image Pre-training) uses contrastive learning from noise image-text pairs to excel at recognizing a wide array of candidates, yet its focus on broad associations hinders the precision in distinguishing subtle differences among fine-grained items. Conversely, Multimodal Large Language Models (MLLMs) excel at classifying fine-grained categories, thanks to their substantial knowledge from pre-training on web-level
DeepCamp AI