Divide-and-Conquer Inference for Large-Scale Visual Recognition with Multimodal Large Language Models

📰 ArXiv cs.AI

arXiv:2605.24799v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong capabilities across a wide range of vision language tasks. However, when applied to large scale image classification, their performance degrades significantly as the label space expands a phenomenon we define as Performance Collapse in Long Sequence Recognition. Through an information theoretic analysis, we reveal that this collapse stems from a fundamental conflict between the esc

Published 26 May 2026

Full Article

Title: Divide-and-Conquer Inference for Large-Scale Visual Recognition with Multimodal Large Language Models

Abstract:
arXiv:2605.24799v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong capabilities across a wide range of vision language tasks. However, when applied to large scale image classification, their performance degrades significantly as the label space expands a phenomenon we define as Performance Collapse in Long Sequence Recognition. Through an information theoretic analysis, we reveal that this collapse stems from a fundamental conflict between the esc
Read full paper → ← Back to Reads