A Causal Language Modeling Detour Improves Encoder Continued Pretraining
📰 ArXiv cs.AI
arXiv:2605.12438v1 Announce Type: cross Abstract: When adapting an encoder to a new domain, the standard approach is to continue training with Masked Language Modeling (MLM). We show that temporarily switching to Causal Language Modeling (CLM) followed by a short MLM decay improves downstream performance. On biomedical texts with ModernBERT, this CLM detour outperforms MLM baselines trained on identical data and compute across 8 French and 11 English biomedical tasks, by +1.2-2.8pp and +0.3-0.8p
DeepCamp AI