LensVLM: Selective Context Expansion for Compressed Visual Representation of Text

📰 ArXiv cs.AI

arXiv:2605.07019v1 Announce Type: cross Abstract: Vision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into long token sequences. Since VLM image encoders map fixed-size images to a fixed number of visual tokens, varying rendering resolution provides a fine-grained compression knob. However, accuracy deteriorates quickly as compression increases: characters shrink below the vision encoder's effective resolu

Published 11 May 2026

Full Article

Title: LensVLM: Selective Context Expansion for Compressed Visual Representation of Text

Abstract:
arXiv:2605.07019v1 Announce Type: cross Abstract: Vision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into long token sequences. Since VLM image encoders map fixed-size images to a fixed number of visual tokens, varying rendering resolution provides a fine-grained compression knob. However, accuracy deteriorates quickly as compression increases: characters shrink below the vision encoder's effective resolu
Read full paper → ← Back to Reads