Soft Label Pruning and Quantization for Large-Scale Dataset Distillation

📰 ArXiv cs.AI

arXiv:2604.18135v1 Announce Type: cross Abstract: Large-scale dataset distillation requires storing auxiliary soft labels that can be 30-40x larger on ImageNet-1K and 200x larger on ImageNet-21K than the condensed images, undermining the goal of dataset compression. We identify two fundamental issues necessitating such extensive labels: (1) insufficient image diversity, where high within-class similarity in synthetic images requires extensive augmentation, and (2) insufficient supervision divers

Published 21 Apr 2026
Read full paper → ← Back to Reads