TY - GEN
T1 - When Does Language Transfer Help? Sequential Fine-Tuning for Cross-Lingual Euphemism Detection
AU - Sammartino, Julia
AU - Barak, Libby
AU - Peng, Jing
AU - Feldman, Anna
N1 - Publisher Copyright:
© 2025 Incoma Ltd. All rights reserved.
PY - 2025
Y1 - 2025
N2 - Euphemisms are culturally variable and often ambiguous, posing challenges for language models, especially in low-resource settings. This paper investigates how cross-lingual transfer via sequential fine-tuning affects euphemism detection across five languages: English, Spanish, Chinese, Turkish, and Yorùbá. We compare sequential fine-tuning with monolingual and simultaneous fine-tuning using XLM-R and mBERT, analyzing how performance is shaped by language pairings, typological features, and pretraining coverage. Results show that sequential fine-tuning with a high-resource L1 improves L2 performance, especially for low-resource languages like Yorùbá and Turkish. XLM-R achieves larger gains but is more sensitive to pretraining gaps and catastrophic forgetting, while mBERT yields more stable, though lower, results. These findings highlight sequential fine-tuning as a simple yet effective strategy for improving euphemism detection in multilingual models, particularly when low-resource languages are involved.
AB - Euphemisms are culturally variable and often ambiguous, posing challenges for language models, especially in low-resource settings. This paper investigates how cross-lingual transfer via sequential fine-tuning affects euphemism detection across five languages: English, Spanish, Chinese, Turkish, and Yorùbá. We compare sequential fine-tuning with monolingual and simultaneous fine-tuning using XLM-R and mBERT, analyzing how performance is shaped by language pairings, typological features, and pretraining coverage. Results show that sequential fine-tuning with a high-resource L1 improves L2 performance, especially for low-resource languages like Yorùbá and Turkish. XLM-R achieves larger gains but is more sensitive to pretraining gaps and catastrophic forgetting, while mBERT yields more stable, though lower, results. These findings highlight sequential fine-tuning as a simple yet effective strategy for improving euphemism detection in multilingual models, particularly when low-resource languages are involved.
UR - https://www.scopus.com/pages/publications/105034223034
U2 - 10.26615/978-954-452-098-4-122
DO - 10.26615/978-954-452-098-4-122
M3 - Conference contribution
AN - SCOPUS:105034223034
T3 - International Conference Recent Advances in Natural Language Processing, RANLP
SP - 1058
EP - 1065
BT - Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era, RANLP 2025
A2 - Angelova, Galia
A2 - Kunilovskaya, Maria
A2 - Escribe, Marie
A2 - Mitkov, Ruslan
PB - Incoma Ltd
T2 - 15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era, RANLP 2025
Y2 - 8 September 2025 through 10 September 2025
ER -