A Causal Language Modeling Detour Improves Encoder Continued Pretraining

📰 ArXiv cs.AI

arXiv:2605.12438v1 Announce Type: cross Abstract: When adapting an encoder to a new domain, the standard approach is to continue training with Masked Language Modeling (MLM). We show that temporarily switching to Causal Language Modeling (CLM) followed by a short MLM decay improves downstream performance. On biomedical texts with ModernBERT, this CLM detour outperforms MLM baselines trained on identical data and compute across 8 French and 11 English biomedical tasks, by +1.2-2.8pp and +0.3-0.8p

Published 13 May 2026
Read full paper → ← Back to Reads