ResearchGate (preprint) Preprint

Posted

How Much of Speech Recognition Must Be Learned? A Parameter-Free Analysis of Lexical Decoding

Po-Ting Lin 1

  1. 1 Independent Researcher
DOI
10.13140/RG.2.2.35863.33447
License
CC BY 4.0
Categories
Speech Recognition · Machine Learning

End-to-end recognisers learn one function from acoustics to text, which makes it impossible to ask where the difficulty sits. Separating the acoustic–phonetic stage from the lexical one, we find that given correct phonemes and no word boundaries a decoder with zero trainable parameters recovers words at 5.66% WER on LibriSpeech dev-clean – within 1.34 points of the floor that out-of-vocabulary words and homophones impose on any decoder over this lexicon, every residual error falling into one of three interpretable causes. Sweeping phoneme accuracy under controlled corruption yields a calibration curve, WER = 9.3% + 2.11 × PER (R² = 0.997), validated to about a point by seven trained CTC heads: the relationship is linear, not amplifying. A frozen language model reordering phoneme-licensed candidates improves accuracy up to a broad optimum with no collapse at high weight, provided both scores are expressed on a common scale.

  • speech recognition
  • error attribution
  • lexical decoding
  • parameter-free decoding
  • phoneme error rate
  • LibriSpeech
Loading PDF…
Cite as (BibTeX)
@misc{lin2026much,
  title = {How Much of Speech Recognition Must Be Learned? A Parameter-Free Analysis of Lexical Decoding},
  author = {Po-Ting Lin},
  year = {2026},
  howpublished = {ResearchGate},
  doi = {10.13140/RG.2.2.35863.33447}
}

← All publications

Curriculum Vitae

Choose a language

English PDF 中文 PDF