CleanPatrick: A Benchmark for Image Data Cleaning
📰 ArXiv cs.AI
arXiv:2505.11034v2 Announce Type: replace-cross Abstract: Robust machine learning depends on clean data, yet current image data cleaning benchmarks rely on synthetic noise or narrow human studies, limiting comparison and real-world relevance. We introduce CleanPatrick, the first large-scale benchmark for data cleaning in the image domain, built upon the publicly available Fitzpatrick17k dermatology dataset. We collect 496,377 binary annotations from 933 medical crowd workers, identify off-topic
DeepCamp AI