| Title |
Parameter-Efficient Fine-Tuning of Whisper with MixLoRA for Improved ASR Performance on ASD Speech |
| Authors |
박예슬(Yeseul Park) ; 이보원(Bowon Lee) |
| DOI |
https://doi.org/10.5573/ieie.2026.63.9.103 |
| Keywords |
Automatic speech recognition; Autism spectrum disorder; Parameter-efficient fine-tuning; MixLoRA |
| Abstract |
Speakers with autism spectrum disorder (ASD) exhibit atypical speech patterns with substantial inter-speaker variation, posing challenges to conventional automatic speech recognition (ASR) systems. To address this, we apply MixLoRA, a parameter-efficient fine-tuning method, originally designed for large language models in multi-task scenarios, to Whisper to enhance ASR performance for ASD speech. MixLoRA integrates low-rank adaptation (LoRA) with a mixture-of-experts framework, dynamically selecting experts per token for efficiency and effectiveness. Experiments on Whisper-base show that MixLoRA achieved 6.35% character error rate (CER) by fine-tuning only 12.89% of parameters, outperforming the baseline, full fine-tuning, and LoRA by 83.64%, 23.22%, and 15.56%, respectively. This approach consistently improves recognition performance across speakers and reduces speaker-level bias, suggesting its potential beyond ASD speech to broader low-resource ASR tasks with diverse speech patterns. |