Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
📰 ArXiv cs.AI
Learn to build cross-lingual quality classifiers for multilingual pretraining data selection to improve LLMs
Action Steps
- Investigate cross-lingual consistency of quality markers in embedding space
- Collect and preprocess multilingual datasets
- Train a quality classifier on a high-resource language
- Apply the trained classifier to low-resource languages
- Evaluate the performance of the cross-lingual classifier
Who Needs to Know This
NLP engineers and researchers can benefit from this technique to optimize their pretraining data and improve the performance of their LLMs
Key Insight
💡 Quality markers in embedding space can show cross-lingual consistency, enabling high-resource languages to subsidize filtering of low-resource languages
Share This
🚀 Improve LLMs with cross-lingual quality classifiers for multilingual pretraining data selection! 📊
Key Takeaways
Learn to build cross-lingual quality classifiers for multilingual pretraining data selection to improve LLMs
Full Article
Title: Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
Abstract:
arXiv:2604.20549v1 Announce Type: cross Abstract: As Large Language Models (LLMs) scale, data curation has shifted from maximizing volume to optimizing the signal-to-noise ratio by performing quality filtering. However, for many languages, native high quality data is insufficient to train robust quality classifiers. This work investigates the idea that quality markers in embedding space may show cross-lingual consistency, which would allow high-resource languages to subsidize the filtering of lo
Abstract:
arXiv:2604.20549v1 Announce Type: cross Abstract: As Large Language Models (LLMs) scale, data curation has shifted from maximizing volume to optimizing the signal-to-noise ratio by performing quality filtering. However, for many languages, native high quality data is insufficient to train robust quality classifiers. This work investigates the idea that quality markers in embedding space may show cross-lingual consistency, which would allow high-resource languages to subsidize the filtering of lo
DeepCamp AI