Digisensus Digisensus
Research

We share our research in ASR, LLMs and other AI technologies

Digisensus Research publishes studies on speech recognition, large language models and the other technologies behind our products.

Studies

Published studies

Study ltd26

Lithuanian speech recognition: effects of dialect training and transcript spelling

Does adding dialect recordings improve Lithuanian speech recognition, and should their transcripts use dialect or standard spelling? A controlled fine-tuning study on the LIEPA-3 corpus with 18 training runs, a pre-registered protocol and an independent customer-service test.

  • Dialect WER falls from 41.62% to 31.30% with 82 h of dialect speech; other speech of the same length reaches only 40.04%
  • The first 25 hours give about two thirds of the gain
  • Customer-service WER falls from 44.89% to 39.06% with standard-spelling dialect transcripts

Released: Code (MIT), three datasets and six fine-tuned models (CC BY 4.0), and the full report as PDF.

Build on it

The datasets, models and code from every study are on the Open source page, with their licenses.

Open source GitHub Hugging Face