# Lithuanian Dialect Speech Recognition Study | Digisensus

> Does dialect speech improve Lithuanian ASR, and should transcripts use dialect or standard spelling? Full report: 18 Parakeet-TDT runs, code, data and models.

[Research](/research/) Study ltd26 · 5 October 2026

# Lithuanian speech recognition: effects of dialect training and transcript spelling

This study is a controlled fine-tuning experiment by Digisensus that tests whether adding dialect recordings improves Lithuanian speech recognition, and whether dialect transcripts should use dialect or standard spelling. It covers 18 training runs on the LIEPA-3 corpus, a pre-registered protocol, an independent customer-service test, and releases the code, data and models.

Saulius Jarašiūnas · Study lead, Digisensus, Vilnius

[Full study, PDF (20 pages)](/assets/research/lithuanian-dialect-asr-study-2026-10-05.pdf) [Code on GitHub](https://github.com/Digisensus/lithuanian-dialect-asr) [🤗Data and models](https://huggingface.co/Digisensus)

Contents:

-   [Abstract](#abstract)
-   [Introduction](#introduction)
-   [Data](#data)
-   [Experimental design](#design)
-   [Results](#results)
-   [Summary of findings](#summary)
-   [Discussion](#discussion)
-   [Limitations](#limitations)
-   [Conclusion](#conclusion)
-   [Resources and links](#resources)
-   [References](#references)
-   [Complete results](#complete-results)

## Abstract

Enterprise conversation intelligence depends on accurate speech recognition, including when speakers use regional dialects. We test whether adding dialect speech improves Lithuanian speech recognition and whether dialect transcripts should use dialect or standard spelling.

We fine-tune Parakeet-TDT on 420.17 hours of Lithuanian telephone speech, adding up to 82.30 hours of dialect recordings. We compare this with adding other spontaneous speech and test dialect and standard transcript spelling. Training settings are fixed. We evaluate dialect recognition against both spellings and an "either-form" score that accepts either reference form. A separate 2.78-hour customer service speech test contains previously unseen speakers.

Adding dialect speech reduces word error rate (WER) against dialect references from 41.62% to 31.30%, averaged over two training runs. Adding other spontaneous speech lowers it only to 40.04%. In the single-run comparison, the first 25.43 hours provide about two thirds of the dialect-test improvement. Standard-spelling dialect transcripts give the lowest either-form WER among models also trained on telephone speech: 25.97%.

On customer service speech, adding dialect recordings lowers WER in each of two training runs, from an average of 44.89% to 39.43% with dialect spelling and to 39.06% with standard spelling; run-to-run variation on this test is larger than on the dialect test. WER on the LIEPA-3 telephone test changes little; that test may share speakers with training. These results support adding dialect recordings for regional speech and customer-service conversations. Transcript spelling affects the output learned and the accuracy measured.

**Main findings: word error rate (%), averaged over two training runs; lower is better. Every model also trains on 420.17 hours of standard telephone speech. The three dialect-test columns score the same recordings using different rules for accepted spelling.**

| Additional training data | Dialect test: dialect spelling | Dialect test: standard spelling | Dialect test: either form | LIEPA-3 telephone | Customer service speech |
| --- | --- | --- | --- | --- | --- |
| None (telephone speech only) | 41.62 | 35.82 | 34.22 | 10.61 | 44.89 |
| 82.36 h other spontaneous speech | 40.04 | 34.10 | 32.40 | 10.34 | 41.66 |
| 82.30 h dialect speech, dialect-spelling transcripts | 31.30 | 36.73 | 27.82 | 10.76 | 39.43 |
| 82.30 h dialect speech, standard-spelling transcripts | 37.48 | 26.94 | 25.97 | 10.56 | 39.06 |

## 1\. Introduction

Digisensus provides conversation intelligence to enterprise customers. This work depends on accurate transcription of telephone conversations, including speech from people who use regional dialects. Recording quality and regional pronunciation can both challenge speech recognition. This study focuses on regional speech: whether dialect recordings improve recognition and how their transcripts should be written. Callers may retain regional pronunciation features even when speaking mostly standard Lithuanian or living outside their region of origin. We therefore also test whether dialect training helps on independent customer-service recordings. We share these findings to help others developing speech recognition for languages with limited training resources.

Lithuanian provides a useful setting for this question. Its regional varieties differ from standard Lithuanian in pronunciation and inflection. A dialect recording can therefore be transcribed using forms that reflect the speaker's pronunciation or the corresponding standard Lithuanian forms. This choice affects both the model's training and its evaluation: a recogniser may identify the intended word but receive an error because its spelling differs from the reference transcript.

We investigate three practical questions. How much does recognition improve as more dialect speech is added to standard-speech training data? Does dialect speech help more than the same amount of other spontaneous speech? And should dialect training transcripts use dialect or standard spelling? We also measure changes on standard Lithuanian telephone speech, explore transfer to regions absent from dialect training, and compare held-out recordings from training speakers with recordings from unseen speakers.

To distinguish the effects of the two spellings, we evaluate the same model outputs against paired dialect and standard transcripts. We also use an "either-form" word error rate that accepts either reference spelling of each word. The example below illustrates why these scores can differ. Together, the three scores show how strongly the measured result depends on the spelling accepted during evaluation.

**Illustrative scoring example. The dialect and standard reference forms represent the same intended word.**

| Model output for the same spoken word | Dialect reference: giveno | Standard reference: gyveno | Either form accepted |
| --- | --- | --- | --- |
| giveno (dialect form of "lived") | Correct | Error | Correct |
| gyveno (standard form of "lived") | Error | Correct | Correct |
| gavo (a different word, "received") | Error | Error | Error |

### Related work

Previous research has examined speech recognition across accents and dialects, including English accents, Egyptian Arabic, and Swiss German. These studies address both the availability of representative speech recordings and the choice of written transcripts. For dialects without a single established spelling, the same spoken words may have several acceptable written forms.

Existing datasets take different approaches. The Arabic MGB-3 challenge transcribes Egyptian Arabic as spoken and allows multiple spellings. The Swiss Parliaments Corpus pairs Swiss German speech with Standard German text, treating recognition as translation between closely related language varieties. Our study compares dialect and standard spelling using the same Lithuanian recordings and training settings. Paired reference transcripts let us evaluate every model under both spelling conventions and with a score that accepts either form.

## 2\. Data

LIEPA-3 is an openly available corpus of 10,000 hours of annotated Lithuanian speech, developed with European Union funding through NextGenerationEU / Naujos kartos Lietuva for speech recognition and research. It contains read speech, spontaneous speech and 100 hours of dialect recordings, and is published by Vilnius University, Vytautas Magnus University and the Institute of the Lithuanian Language under CC BY 4.0.

We use three parts of LIEPA-3: standard telephone conversations, dialect recordings from four regions, and other spontaneous speech from radio and dictaphone recordings. Telephone speech provides the main training data. Dialect recordings let us test the value of adding regional speech, while other spontaneous recordings provide a comparison for adding a similar amount of data. We also evaluate the models on two public Lithuanian read-speech datasets, FLEURS and Common Voice. We additionally test on private customer service speech.

**Main training and test data. The dialect and control training sets have approximately equal durations. Telephone training and test recordings may share speakers.**

| Data | Hours | Purpose |
| --- | --- | --- |
| Telephone speech: training | 420.17 | Train the standard-speech baseline |
| Dialect speech: training | 82.30 | Test dialect data and spelling choices; 97 speakers |
| Other spontaneous speech: training | 82.36 | Compare with adding a similar amount of other speech |
| Dialect speech: main test | 5.01 | Test dialect recognition on 12 unseen speakers |
| Telephone speech: test | 4.36 | Measure changes on standard telephone speech |
| Customer service speech: test | 2.78 | Independent private test on previously unseen speakers |
| FLEURS: test | 2.97 | Additional evaluation on read speech |
| Common Voice 19.0: test | 7.15 | Additional evaluation on read speech |

### Dialect recordings and speaker separation

The dialect recordings cover Aukštaitija, Dzūkija, Suvalkija and Žemaitija. The main test contains three speakers from each region, selected to balance gender and age groups. None of these speakers appears in training.

Another 8 speakers provide 2.67 hours of development recordings used to monitor training. Their recordings are also excluded from training. All remaining recordings from the test and development speakers, 9.50 hours, are unused.

A separate 1.01-hour test contains recordings withheld from 12 speakers in the dialect training pool. This lets us compare new recordings from those speakers with recordings from unseen speakers. The withheld recordings themselves are never used for training. Models trained on smaller dialect subsets or with a region excluded may have seen only some of these speakers.

To study how improvement depends on data quantity, we build progressively larger dialect training sets. Each larger set includes the smaller one and adds speakers. The repository records the technical split identifiers and construction rules.

**Increasing amounts of dialect training speech. Each set includes all speakers and recordings from the smaller sets.**

| Dialect training set | Hours | Speakers |
| --- | --- | --- |
| Smallest subset | 12.83 | 16 |
| Second subset | 25.43 | 31 |
| Third subset | 50.47 | 61 |
| Full training pool | 82.30 | 97 |

### Two transcript versions for the same recordings

The original dialect transcripts represent the words as pronounced. We create a second version using the corresponding standard Lithuanian forms. The standardisation guidelines aim to preserve word order, grammatical meaning and word choice; dialect words without a standard counterpart remain unchanged.

The two versions are aligned word by word, allowing the same model output to be evaluated against either spelling. Standardisation uses language-model assistance, written guidelines and automatic alignment checks. The preparation and review procedure is documented in the repository. We also label the differences between paired words, for example changes to endings or pronunciation, to examine which word types remain difficult to recognise.

**Word classes used to analyse recognition errors.**

| Word class | Meaning | Example: dialect → standard |
| --- | --- | --- |
| Unchanged | Same word form in both spellings | – |
| Ending change | Only the ending differs | būdava → būdavo; seseris → seserys |
| Letters ą, ę, į, ų | Dialect form uses a, e, i, u in place of these letters | viska → viską; mūsu → mūsų |
| Other pronunciation change | Another sound change in the stem | giveno → gyveno; teip → taip |
| Dialect-only word | No standard counterpart; kept unchanged | rozu; bajino |

### Telephone speech and the comparison dataset

The telephone data follow the published training and test split of the dataset used in this study. Speaker identities are unavailable, so the two sets may contain recordings from the same callers. We remove training clips whose transcripts of six or more words repeat a test or development transcript. Telephone results therefore describe performance on this particular recording collection; they do not establish performance on entirely unseen callers.

The comparison dataset contains other spontaneous speech, mainly from radio and dictaphone recordings. Its duration, gender distribution, age groups and clip lengths are matched approximately to the dialect training data. This comparison helps assess whether dialect recordings offer more benefit than simply adding more speech.

### Customer service speech

Digisensus provides 2.78 hours of private customer service speech from speakers absent from the training, development and earlier test sets. We evaluate the existing models on these independent recordings to test whether dialect-training benefits extend beyond LIEPA-3.

### Public read-speech tests

FLEURS and Common Voice test recognition of read speech, which differs from spontaneous telephone conversations. We report these results separately; dataset details are recorded in the protocol.

## 3\. Experimental design

### What we compare

We compare a telephone-only baseline with models fine-tuned from the same pretrained checkpoint on different combinations of training data. We vary the amount and type of added speech, or the spelling of its transcripts. Each comparison answers a specific question.

**Questions and experimental comparisons.**

| Question | Comparison |
| --- | --- |
| How does adding more dialect speech affect recognition? | Add progressively larger dialect datasets: 12.83, 25.43, 50.47, 82.30 hours. |
| Does dialect speech help more than simply adding data? | Compare adding 82.30 hours of dialect speech with 82.36 hours of other spontaneous speech. |
| Which transcript spelling works better? | Train on the same dialect recordings with dialect-spelling or standard-spelling transcripts. |
| Does the benefit extend to customer-service conversations? | Evaluate all models on independent customer service speech from previously unseen speakers. |
| Does the benefit extend to an unrepresented region? | Compare models trained with and without recordings from the region being tested. |
| Are training speakers easier to recognise? | Compare withheld recordings from speakers in the training pool with recordings from unseen speakers. |
| Is standard telephone training necessary? | Compare combined training with training on dialect recordings alone. |

We also evaluate the original Parakeet-TDT and Whisper large-v3 models without additional training, providing reference points for the fine-tuned models.

### How we train the models

All fine-tuned models start from the same multilingual Parakeet-TDT checkpoint and use the same training settings. Each run trains for 10,000 updates. We evaluate the final checkpoint rather than selecting a checkpoint based on development or test results.

We repeat four main conditions twice: telephone speech alone, telephone speech plus other spontaneous speech, and telephone speech plus the full dialect dataset under each transcript spelling. The repeats use different random seeds, which change aspects such as training-data order. Other conditions have one run. Keeping the training procedure fixed makes the comparisons consistent, although different training datasets may need different amounts of training to reach their best performance. Full training settings are recorded in the repository.

### How we measure recognition

Word error rate (WER) counts substituted, missing and extra words, divided by the number of reference words, expressed as a percentage. Lower is better. We calculate it across all recordings in each test set. Before scoring, we normalise capitalisation, punctuation and spacing. For dialect speech, we report three scores separately.

**Three scoring rules applied to the same dialect recordings and model outputs. We do not average these scores.**

| Score | What counts as the correct word form? |
| --- | --- |
| Dialect-spelling WER | The form in the dialect transcript. |
| Standard-spelling WER | The form in the standard transcript. |
| Either-form WER | Either of the two aligned reference forms. |

These scores show how the measured result depends on the accepted spelling. The either-form score reduces penalties for choosing between the two reference forms; it does not establish that every remaining error is purely a recognition error. We describe changes using before-and-after WERs. When reporting a difference, we use percentage points: a fall from 40% to 30% is a reduction of 10 percentage points.

### How we assess uncertainty

We compare models on the same test recordings. We estimate 95% confidence intervals by repeatedly resampling speakers (10,000 resamples), or individual recordings where speaker identities are unavailable. For customer service speech, we resample whole calls. These intervals describe uncertainty from the test sample; they do not include variation between training runs or errors in the reference transcripts.

We report variation between repeated training runs separately. Regional comparisons contain only three test speakers per region and should be treated as exploratory. Likewise, the comparison between training speakers and unseen speakers uses different recordings and cannot isolate speaker familiarity from differences in recording difficulty.

### Protocol and changes

The original experimental protocol was frozen after the initial training runs and before their test evaluations. One unadapted model had already been evaluated to check the scoring pipeline. The Suvalkija exclusion experiment was added after the other regional results had been inspected, using the same training settings. The customer service speech evaluation used the already-trained models. Its analysis was exploratory: the plan was written after an earlier pilot evaluation on the same recordings.

After scoring, we moved the public read-speech benchmarks to secondary results because their scores varied substantially between training runs. We also did not perform the additional training runs required by the protocol for resolving small effects on telephone speech. We therefore leave the telephone-speech comparison between transcript spellings unresolved. The planned manual post-edit of the standard-spelling test references was not carried out; all scores use the automatically standardised version. The full protocol history and reference versions are recorded in the repository.

## 4\. Results

Adding dialect recordings improves dialect recognition more than adding a similar amount of other spontaneous speech. Much of the improvement comes from the first additions, and transcript spelling changes both the output and the measured error rate. Complete model scores are in the appendix at the end of this page.

### 4.1 How does adding more dialect speech affect recognition?

We test how dialect recognition changes as more dialect recordings are added to the standard telephone training data. We train models with no dialect speech and with 12.83, 25.43, 50.47, 82.30 additional hours. These models use dialect-spelling training transcripts and identical training settings.

Dialect recognition improves at every step. When evaluated against dialect-spelling references, word error rate falls from 42.11% without dialect training data to 31.28% with 82.30 hours. The first 25.43 hours provide about two thirds of this reduction. Further additions continue to help, but the improvements become smaller. The improvement also appears when either dialect or standard spelling is accepted: word error rate falls from 34.83% to 27.60%. This shows that the benefit is not limited to reproducing dialect spelling.

Customer service speech also improves as more dialect recordings are added. In the same single-run comparison, WER falls from 42.61% without dialect data to 37.01% with 82.30 hours. The improvement continues across the tested training amounts. This differs from the LIEPA-3 telephone test, where WER changes little: from 10.76% to 10.80%. These telephone values are low relative to the original models and likely benefit from shared callers and recording conditions between the telephone training and test sets. A second training run confirms the overall pattern: dialect-reference WER falls from 41.13% to 31.31%, while telephone WER changes from 10.46% to 10.72%.

Paired differences between the full-dialect model and the telephone-only model, with 95% intervals: dialect test −10.83 \[−14.02, −7.84\] pp in run 1 and −9.81 \[−12.83, −6.80\] pp in run 2; telephone test +0.04 \[−0.18, +0.27\] pp and +0.26 \[+0.06, +0.47\] pp. These results suggest that a relatively small dialect collection can provide a substantial improvement, with additional recordings giving diminishing returns. However, each larger dataset also includes more speakers, so this experiment does not separate the effect of recording hours from speaker diversity.

**One training run per condition (seed 1), using dialect-spelling transcripts. WER in %, lower is better. The abstract reports averages of two runs for the baseline and full-data models, which explains the slightly different values there.**

| Dialect speech added | Dialect-test WER: dialect spelling | Dialect-test WER: either spelling | LIEPA-3 telephone WER | Customer service speech WER |
| --- | --- | --- | --- | --- |
| None | 42.11 | 34.83 | 10.76 | 42.61 |
| 12.83 hours | 38.08 | 33.30 | 10.61 | 41.48 |
| 25.43 hours | 34.77 | 30.29 | 10.61 | 40.30 |
| 50.47 hours | 32.77 | 28.96 | 10.49 | 37.91 |
| 82.30 hours | 31.28 | 27.60 | 10.80 | 37.01 |

![Four line charts of WER against dialect training hours added (0, 13, 25, 50, 82): dialect test with dialect spelling falls from 42.11% to 31.28%; with either spelling accepted from 34.83% to 27.60%; LIEPA-3 telephone test stays near 10.5–10.8%; customer service speech falls from 42.61% to 37.01%.](/assets/research/ltd26-fig1_dose.png)

WER as dialect training hours increase, using dialect-spelling transcripts and seed 1 throughout. Top left: dialect references. Top right: either spelling accepted. Bottom left: LIEPA-3 telephone speech on an expanded vertical scale. Bottom right: customer service speech. All panels except the telephone panel share the same vertical scale.

### 4.2 Does dialect speech help more than simply adding more recordings?

We compare two additions to the same standard telephone training data: 82.30 hours of dialect speech and 82.36 hours of other spontaneous speech. The dialect recordings use dialect-spelling transcripts. This comparison tests whether the improvement comes simply from adding more training data.

Both additions improve dialect recognition, but dialect recordings provide a much larger benefit. Averaged over two training runs, WER against dialect references falls from 41.62% to 40.04% with other spontaneous speech, and to 31.30% with dialect speech. The advantage remains when either spelling is accepted: WER reaches 32.40% with other spontaneous speech and 27.82% with dialect speech. The result therefore does not depend solely on matching dialect spelling. On the LIEPA-3 telephone test, the other spontaneous recordings give the lower WER: 10.34%, compared with 10.76% after adding dialect recordings.

Paired differences, dialect speech minus other spontaneous speech, with 95% intervals: dialect test −8.62 \[−11.52, −5.84\] pp (run 1) and −8.86 \[−12.02, −5.78\] pp (run 2) against dialect references; −4.58 \[−6.75, −2.42\] pp and −4.57 \[−7.01, −2.16\] pp with either form accepted; telephone test +0.49 \[+0.28, +0.70\] pp and +0.34 \[+0.13, +0.55\] pp.

On the independent customer service speech test, average WER falls from 44.89% with telephone-only training to 41.66% with other spontaneous speech and 39.43% with dialect speech. Both dialect-training runs improve over their corresponding telephone-only run: from 42.61% to 37.01% and from 47.17% to 41.86%. Resampling calls gives a reduction of 5.46 percentage points with a 95% interval of 4.85 to 6.13, but this interval reflects only the test sample: the two telephone-only runs differ by more than its width, so run-to-run variation, not the test sample, is the main uncertainty on this test. Average WER is lower than with other spontaneous speech, but that advantage occurs in only one of the two training runs.

These results show that the type of additional speech matters. For recognising dialect speakers in this study, dialect recordings help substantially more than a similar amount of other spontaneous speech.

**Mean WER (%) of two training runs. All models also train on 420.17 hours of standard telephone speech. Lower WER is better.**

| Additional training speech | Dialect WER: dialect spelling | Dialect WER: either spelling | LIEPA-3 telephone WER | Customer service speech WER |
| --- | --- | --- | --- | --- |
| None | 41.62 | 34.22 | 10.61 | 44.89 |
| Other spontaneous speech | 40.04 | 32.40 | 10.34 | 41.66 |
| Dialect speech | 31.30 | 27.82 | 10.76 | 39.43 |

![Dot chart of customer service speech WER for each training run and their average: telephone only 42.61% and 47.17% (mean 44.89%); other spontaneous speech 42.42% and 40.90% (41.66%); dialect speech with dialect spelling 37.01% and 41.86% (39.43%); dialect speech with standard spelling 38.01% and 40.11% (39.06%).](/assets/research/ltd26-fig_customer.png)

Customer service speech: WER for each training run and their average. All conditions include the same telephone training data; the additions use approximately equal hours. Lower is better.

### 4.3 Which transcript spelling works better?

We compare two training conditions, each repeated twice, using the same standard telephone data and the same 82.30 hours of dialect recordings. Only the dialect transcripts differ: one condition uses dialect spelling, the other standard spelling. We evaluate both conditions against the same test recordings using all three scoring methods.

The reference spelling strongly affects the measured improvement. With dialect-spelling training transcripts, WER against dialect references falls from 41.62% to 31.30%. Yet against standard references, WER for the same training condition changes from 35.82% to 36.73%. Its apparent improvement therefore depends on which spelling the test accepts.

When either spelling is accepted, both training approaches improve on the model without dialect data. WER falls from 34.22% to 27.82% with dialect-spelling transcripts and to 25.97% with standard-spelling transcripts. Standard-spelling transcripts give the lower either-form WER in both repeated training runs: −1.54 \[−2.76, −0.34\] pp and −2.16 \[−3.20, −1.06\] pp relative to dialect spelling, and −8.78 \[−10.38, −7.15\] pp and −7.72 \[−9.20, −6.04\] pp relative to no dialect data.

Among models trained with the standard telephone data, standard-spelling dialect transcripts therefore give the best overall result under either-form scoring. The same pattern appears at the smaller training amounts of 25.43 hours and 50.47 hours: either-form WER of 27.91%, 26.39% and 26.06% with standard spelling in the single-run series. This experiment does not establish why standard spelling helps. One possibility is that the benefit is partly stylistic: the standard-spelling training transcripts and the standard-spelling test references were produced by the same standardisation procedure, so a model trained on that procedure's output is also scored against it. The dialect-spelling score and the dialect-only word class, where standard spelling does worse, are not affected by this.

On the LIEPA-3 telephone test, average WER is 10.76% with dialect-spelling transcripts and 10.56% with standard-spelling transcripts. Because this difference is small relative to the observed variation between training runs, we leave this comparison unresolved. On customer service speech, average WER is 39.43% with dialect-spelling transcripts and 39.06% with standard-spelling transcripts. Both improve over telephone-only training in each repeated run, but the small difference between spellings is inconclusive. The lowest observed customer-service WER is 36.14%, using 50.47 hours of dialect speech with standard spelling. That condition has one run and does not establish an optimal training amount.

There is also a trade-off for dialect-only vocabulary. In the single-run word-class analysis, the error rate attributed to words without a standard counterpart is 62.12% with dialect-spelling training transcripts and 73.84% with standard-spelling transcripts. This measure includes extra words assigned to that word class during scoring. The better overall score therefore does not mean that every type of word improves.

**Error rate (%) by reference word class on the dialect test, seed 1. Columns: stock Parakeet-TDT; telephone only (S); telephone + other spontaneous speech (S+X80); telephone + dialect with dialect spelling (S+D80.dial); with standard spelling (S+D80.std); dialect only with standard spelling (D80.std).**

| Reference word class | stock | S | S+X80 | S+D80.dial | S+D80.std | D80.std |
| --- | --- | --- | --- | --- | --- | --- |
| Unchanged | 63.33 | 27.31 | 24.61 | 22.78 | 20.60 | 17.44 |
| Other pronunciation change | 85.67 | 64.32 | 61.49 | 45.14 | 45.95 | 37.50 |
| Ending change | 82.05 | 57.78 | 55.87 | 39.42 | 38.60 | 30.32 |
| Letters ą, ę, į, ų | 77.50 | 37.70 | 35.10 | 37.90 | 32.10 | 28.10 |
| Dialect-only word | 96.16 | 76.06 | 72.93 | 62.12 | 73.84 | 69.80 |

![Bar chart of the same dialect test scored three ways. Dialect spelling accepted: 41.62% without dialect data, 31.30% with dialect-spelling transcripts, 37.48% with standard-spelling transcripts. Standard spelling accepted: 35.82%, 36.73%, 26.94%. Either spelling accepted: 34.22%, 27.82%, 25.97%.](/assets/research/ltd26-fig3_spelling.png)

The same dialect test scored three ways. Bars show mean WER over two training runs; dots show the individual runs. All models also train on standard telephone speech; the two dialect-trained models add the same 82.30 hours of recordings. Lower WER is better.

![Two line charts comparing dialect-spelling and standard-spelling training transcripts as dialect hours increase. Left: dialect test with either spelling accepted, where standard spelling stays lower at 25, 50 and 82 hours. Right: LIEPA-3 telephone test on an expanded scale, with both lines between 10.4% and 10.8%.](/assets/research/ltd26-fig_spelling_dose.png)

Transcript-spelling comparison as dialect training hours increase. Points show means where two runs are available and single runs otherwise; vertical segments show the range of the two runs. The telephone panel uses an expanded vertical scale. All models also train on standard telephone speech.

### 4.4 Does dialect training help a region excluded from training?

We test whether training on dialect speech from other regions helps recognise a region whose dialect recordings were excluded from training. The test covers all four regions: Aukštaitija, Dzūkija, Suvalkija and Žemaitija. It contains 12 test speakers in total, three from each region. None of these speakers appears in the training data.

For each region, we compare three training conditions: telephone speech alone; telephone speech plus dialect recordings from the other three regions; and telephone speech plus dialect recordings from all four regions. Each condition uses one training run. Dialect training transcripts use standard spelling, and scoring accepts either dialect or standard spelling. Within each region, all three conditions are evaluated on the same recordings.

Training on the other three regions lowers WER in every region compared with telephone-only training: from 34.52% to 26.55% in Aukštaitija, from 40.13% to 30.77% in Dzūkija, from 27.46% to 22.37% in Suvalkija, from 36.40% to 29.90% in Žemaitija. This supports a benefit from dialect training even when the tested region is absent from the dialect training data.

Training on all four regions gives the lowest measured WER in each comparison. However, the additional improvement for Žemaitija is uncertain: its 95% confidence interval includes no improvement. These are exploratory results based on three test speakers per region and one run per condition. Including the tested region also adds training hours, so we cannot separate the value of that region's speech from the benefit of more data.

**Either-form WER (%) on each region's three test speakers, one training run per condition, with paired differences in percentage points and 95% intervals. "Hours" is the dialect training data left after removing the region.**

| Region | Hours | Telephone only | Other three regions | All regions, 50 h | All regions, 82 h | Left out − all regions | Left out − telephone only |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Aukštaitija | 47.14 | 34.52 | 26.55 | 25.29 | 24.63 | +1.92 [+1.54, +2.12] | −7.97 [−9.44, −4.89] |
| Dzūkija | 68.01 | 40.13 | 30.77 | 28.33 | 28.59 | +2.19 [+1.31, +2.59] | −9.36 [−10.57, −8.13] |
| Suvalkija | 72.97 | 27.46 | 22.37 | 21.69 | 21.64 | +0.73 [+0.49, +1.01] | −5.09 [−8.09, −2.94] |
| Žemaitija | 58.78 | 36.40 | 29.90 | 29.73 | 28.85 | +1.04 [−0.16, +1.74] | −6.50 [−10.88, −4.72] |

![Grouped bar chart of either-form WER by region. Aukštaitija: 34.52% telephone only, 26.55% other three regions, 24.63% all regions. Dzūkija: 40.13%, 30.77%, 28.59%. Suvalkija: 27.46%, 22.37%, 21.64%. Žemaitija: 36.40%, 29.90%, 28.85%.](/assets/research/ltd26-fig4_regions.png)

Dialect recognition across four regions, evaluated on 12 test speakers in total (three per region). Each group compares telephone-only training, adding dialect speech from the other three regions, and adding dialect speech from all four regions. Bars show WER with either dialect or standard spelling accepted, using one training run per condition. Lower is better.

### 4.5 Are speakers represented in training easier to recognise?

We compare two sets of dialect recordings: withheld recordings from 12 speakers in the dialect training pool, and recordings from 12 speakers absent from training. None of the evaluated recordings was used for training. We expected models trained on the first group's other recordings to recognise those speakers more accurately.

The observed pattern is the opposite. Across all 19 fine-tuned training runs, WER is higher on the recordings from speakers in the training pool, by 0.77 to 2.41 percentage points. Both sets are scored by accepting either dialect or standard spelling. However, the same pattern appears in models trained without dialect speech, which had not encountered either group during fine-tuning. Models trained on smaller dialect subsets or with a region excluded had encountered only some speakers in the training pool. These results suggest that differences in the difficulty of the two recording sets affect the comparison.

We therefore cannot conclude that speaker familiarity provides no benefit. The two sets contain different speakers and recordings, so this experiment does not isolate its effect. It shows only that the withheld recordings from speakers in the training pool were harder for these models than the recordings from unseen speakers.

![Scatter plot of 19 fine-tuned training runs: either-form WER on unseen speakers (horizontal) against WER on withheld recordings from training-pool speakers (vertical). Every point lies above the equal-WER diagonal, by 0.77 to 2.41 percentage points.](/assets/research/ltd26-fig5_seen.png)

Each point represents one fine-tuned training run. Both axes show either-form WER (%). The dashed diagonal marks equal WER on both sets. Points above it indicate higher error on the withheld recordings from speakers in the dialect training pool. Speaker overlap depends on the model's training data; the two recording sets are the same across all runs.

### 4.6 What happens when we train on dialect speech alone?

We compare fine-tuning on dialect recordings alone with fine-tuning on standard telephone recordings alone or on both datasets together. "Dialect-only" refers to the data used for fine-tuning; all these models start from the same pretrained multilingual Parakeet-TDT model.

Using only 82.30 hours of dialect speech with standard-spelling transcripts gives the lowest either-form dialect WER observed in this study: 21.91%. However, its telephone-test WER is 19.63%. Combining the dialect recordings with standard telephone training data gives a higher dialect WER of 25.97%, but a much lower telephone WER of 10.56%. The size of this telephone gap partly reflects the possible speaker and channel overlap of the telephone test; on the independent customer service test the gap is smaller. Dialect-only training with dialect-spelling transcripts shows a similar trade-off: either-form dialect WER is 24.43%, while telephone WER is 28.20%.

The original Parakeet-TDT model, without fine-tuning on our study data, reaches 68.15% either-form dialect WER and 42.53% telephone WER; Whisper large-v3 reaches 53.15% and 27.43%, respectively. Fine-tuning Parakeet-TDT on telephone speech alone improves both scores relative to these original models, reaching 34.22% and 10.61%. These results show why the best training choice depends on the intended use. Dialect-only fine-tuning performs best on this dialect test, while combined training supports both dialect and telephone recognition. The dialect-only conditions each have one training run, so their apparent advantage needs confirmation with repeated runs.

**WER (%), lower is better. Telephone-only and combined-training values are means of two runs; dialect-only values come from one run. Original models are evaluated without additional training. Telephone and dialect training sets contain 420.17 and 82.30 hours, respectively.**

| Model and fine-tuning data | Dialect WER: either spelling | LIEPA-3 telephone WER | Customer service speech WER |
| --- | --- | --- | --- |
| Original Parakeet-TDT, no study fine-tuning | 68.15 | 42.53 | 59.51 |
| Original Whisper large-v3, no study fine-tuning | 53.15 | 27.43 | 54.67 |
| Parakeet-TDT: telephone speech only | 34.22 | 10.61 | 44.89 |
| Parakeet-TDT: dialect speech only, dialect spelling | 24.43 | 28.20 | 51.36 |
| Parakeet-TDT: dialect speech only, standard spelling | 21.91 | 19.63 | 42.59 |
| Parakeet-TDT: telephone + dialect speech, standard spelling | 25.97 | 10.56 | 39.06 |

### 4.7 Do the improvements carry over to public read-speech tests?

We also evaluate every model on FLEURS and Common Voice, two public Lithuanian datasets containing people reading aloud. These tests examine performance on a different type of speech from the spontaneous telephone conversations used for most training. Original Parakeet-TDT reaches 22.10% WER on FLEURS and 16.52% on Common Voice. After fine-tuning on telephone speech alone, the two training runs reach 28.62% and 28.65% on FLEURS, and 35.20% and 44.74% on Common Voice.

Adding the full dialect dataset improves both benchmarks relative to telephone-only training in both runs, under each transcript spelling. The size of the improvement varies between runs. Every fine-tuned model has higher WER than stock Parakeet on Common Voice, while some achieve lower WER on FLEURS. For example, with standard-spelling dialect transcripts, FLEURS WER is 26.87% in one run and 21.51% in the other. This variation makes small differences between training conditions difficult to interpret. Seed-to-seed standard deviation over the four repeated conditions is 3.14 pp on FLEURS and 4.38 pp on Common Voice, against 0.36 pp on the dialect test and 0.13 pp on the telephone test.

The recording conditions also differ. Most training speech has telephone bandwidth and is spontaneous, whereas the benchmarks contain wideband read speech. Dialect recordings are also wideband, so adding them changes both regional coverage and recording conditions. This experiment does not establish which differences cause the benchmark behaviour. We therefore treat these tests as secondary results.

**WER (%), run 1 / run 2; lower is better. Each fine-tuned condition has two training runs; the original model has no additional training. Combined models add 82.30 hours of dialect speech to 420.17 hours of telephone speech.**

| Fine-tuning data | FLEURS WER, run 1 / run 2 | Common Voice WER, run 1 / run 2 |
| --- | --- | --- |
| None: original Parakeet-TDT | 22.10 | 16.52 |
| Telephone speech only | 28.62 / 28.65 | 35.20 / 44.74 |
| Telephone + dialect speech, dialect spelling | 23.24 / 22.58 | 25.77 / 27.99 |
| Telephone + dialect speech, standard spelling | 26.87 / 21.51 | 27.75 / 23.85 |

![Two dot charts of WER on the FLEURS and Common Voice 19.0 Lithuanian test sets for every trained model, with the stock Parakeet-TDT and Whisper large-v3 as horizontal lines. Seed-to-seed differences are large on both benchmarks.](/assets/research/ltd26-figS1_benchmarks.png)

Public read-speech benchmarks, secondary results: FLEURS and Common Voice 19.0 Lithuanian test sets, every trained model (filled: seed 1, hollow: seed 2), with the stock models as horizontal lines. Seed-to-seed differences are large on both benchmarks. The training data are mostly telephone-band speech, while both benchmarks are wideband read speech.

## 5\. Summary of findings

Conclusions apply to the datasets, model and training settings tested in this study. The original protocol hypotheses and verdicts are preserved in the repository.

**Summary of findings.**

| Question | Finding |
| --- | --- |
| Does adding dialect speech help? | Yes. In the single-run series using dialect-spelling transcripts, dialect and customer service speech WER fall at every tested addition. Later additions give smaller dialect-test gains. LIEPA-3 telephone WER changes little. |
| Is the improvement simply from adding more data? | Dialect recordings improve dialect recognition substantially more than a similar amount of other spontaneous speech. |
| Which transcript spelling works better? | Among models also trained on telephone speech, standard-spelling dialect transcripts give the lowest either-form WER. Dialect-only vocabulary remains a weakness. |
| Does the reference spelling affect the result? | Yes. The same model can show a large improvement against dialect references and no improvement against standard references. |
| Does standard spelling improve LIEPA-3 telephone recognition? | Unresolved. The observed difference is small relative to variation between training runs. |
| Does dialect training help an excluded region? | Yes in all four regional comparisons, covering 12 test speakers in total. Each uses one run per condition. Regional coverage and training hours are mixed. |
| Are speakers represented in training easier to recognise? | The withheld recordings from those speakers have higher WER. Different speakers and recordings prevent a conclusion about the effect of familiarity. |
| Can we train on dialect recordings alone? | This gives the lowest dialect-test WER, but higher telephone-test WER than combined training. The dialect-only results each come from one run. |
| Does dialect training help customer-service speech? | Yes on this independent private test: both repeated dialect-training conditions improve over telephone-only training. |
| Do improvements extend to public read-speech tests? | Yes relative to telephone-only training in both repeated runs, but not consistently relative to stock Parakeet. |

## 6\. Discussion

For conversation-intelligence systems such as those developed by Digisensus, these results support collecting dialect recordings alongside standard telephone speech. Dialect data improve recognition of regional speakers more than a similar amount of other spontaneous speech. Much of the measured improvement comes from the first additions, suggesting that a modest, carefully selected dialect collection can be useful. However, this study increases recording hours and speaker diversity together, so it does not establish a minimum number of hours or speakers.

The independent customer service speech results strengthen the practical case: adding dialect recordings also improves recognition of previously unseen customer-service speakers. Retained regional pronunciation is one possible explanation, but the test does not label dialect features or callers' regional origins. The improvement alone does not establish its cause.

Transcript spelling should reflect the application's intended output. In this study, standard-spelling dialect transcripts give the lowest either-form error among models also trained on telephone speech. They are therefore a promising choice for applications that produce standard Lithuanian text. The gain is not uniform, however: on words that have no standard counterpart, standard-spelling training raised the error rate from 62.12% to 73.84% in the single-run analysis. Applications that need dialect forms, or that depend on such words, have a different requirement and should evaluate this word class separately.

Training-data choices also involve a trade-off between speech types. Dialect-only fine-tuning gives the lowest error on the dialect test, while combining dialect and telephone recordings gives much lower error on this telephone collection, whose test set may share callers with training, and lower error on customer service speech. On public read-speech tests, adding dialect speech improves both repeated runs relative to telephone-only training, but does not consistently beat stock Parakeet. Practical model selection should therefore use test recordings representative of the intended service.

Evaluation should make the accepted spelling explicit. Scoring the same output against dialect and standard references can produce different conclusions about improvement. Paired reference transcripts and either-form scoring help reveal this dependence. Reporting all three scores allows readers to distinguish accuracy under a required output convention from accuracy when either reference form is acceptable.

The findings motivate similar experiments in other languages with regional variation, but they do not establish that the same improvements will occur elsewhere. The training data come from one Lithuanian corpus; we study one pretrained model family and one training procedure, with supplementary evaluation on a private customer-service collection. They also do not directly measure improvements under noise, compression or poor telephone connections.

## 7\. Limitations

**Transcript accuracy.** The standard-spelling transcripts were prepared with language-model assistance and automatic checks. Errors in these transcripts can affect both training and evaluation. The team does not include a professional dialectologist, and the standardisation guidelines reflect practical decisions that others may judge differently. Because training and test standard-spelling transcripts share this procedure, part of the advantage of standard-spelling training may reflect consistency with the references rather than better recognition. The preparation procedure is recorded in the repository.

**Test coverage.** The main dialect test contains 12 speakers, with only three per region. Regional findings therefore remain exploratory. The LIEPA-3 telephone training and test sets may share speakers, so their results do not establish performance on entirely unseen callers. The comparison of familiar and unfamiliar speakers also uses different recordings and cannot isolate the effect of speaker familiarity.

**Experimental coverage.** We fine-tuned one pretrained model family using fixed training settings and 10,000 updates per run. Different datasets may need different training schedules to reach their best performance. Four main conditions have two training runs; the remaining conditions have one. More dialect hours also introduce more speakers, and excluding a region reduces the training data. These comparisons therefore combine several possible influences.

**Scope of the conclusions.** The small LIEPA-3 telephone-test difference between transcript spellings remains unresolved. Results on public read-speech datasets vary between training runs and differ from the telephone and dialect results. We did not isolate the effects of recording bandwidth, speaking style, noise or compression. The independent customer service speech test demonstrates improvement on one private collection, but does not establish the same benefit across other customer populations, languages or models. Its speakers have no dialect labels, so we cannot attribute the gain specifically to regional pronunciation.

## 8\. Conclusion

Adding dialect recordings improves Lithuanian dialect recognition more than adding a similar amount of other spontaneous speech. The benefit also extends to independent customer service speech: both repeated runs improve over telephone-only training under each transcript spelling. This supports combining telephone and dialect recordings for conversation-intelligence applications.

Standard-spelling dialect transcripts give the lowest either-form WER among models also trained on telephone speech. However, no training choice is best on every test or word class, and this study does not establish an optimal amount of dialect data. Model selection should use recordings representative of the intended service, with the accepted transcript spelling stated explicitly.

## Reproducibility and resources

Everything the study released is listed here with its link, version tag and license. The public analysis regenerates the reported results and tables from the released evaluation records. Customer service speech audio, transcripts and model outputs remain private; only aggregate results are reported. Other checkpoints are available on request from saulius@digisensus.com.

Report CC-BY-4.0

### [lithuanian-dialect-asr-study-2026-10-05.pdf](/assets/research/lithuanian-dialect-asr-study-2026-10-05.pdf)

The final report, 20 pages, with every table, figure and the complete results appendix.

[/assets/research/lithuanian-dialect-asr-study-2026-10-05.pdf](/assets/research/lithuanian-dialect-asr-study-2026-10-05.pdf)

Code MIT

### [Digisensus/lithuanian-dialect-asr](https://github.com/Digisensus/lithuanian-dialect-asr)

Study code, training recipes, frozen splits, the run registry, the pre-registered protocol and the analysis scripts that regenerate every number in the report.

[github.com/Digisensus/lithuanian-dialect-asr · v1.0](https://github.com/Digisensus/lithuanian-dialect-asr)

Dataset CC-BY-4.0

### [Digisensus/lithuanian-dialect-speech-liepa-3-100h-punctuated](https://huggingface.co/datasets/Digisensus/lithuanian-dialect-speech-liepa-3-100h-punctuated)

Dialect speech with speaker-disjoint splits, the four training subsets and the standard-spelling transcript layer with word alignment and change tags.

[huggingface.co/datasets/Digisensus/lithuanian-dialect-speech-liepa-3-100h-punctuated · v2.0](https://huggingface.co/datasets/Digisensus/lithuanian-dialect-speech-liepa-3-100h-punctuated)

Dataset CC-BY-4.0

### [Digisensus/lithuanian-phone-speech-liepa-3-429h-punctuated](https://huggingface.co/datasets/Digisensus/lithuanian-phone-speech-liepa-3-429h-punctuated)

Telephone speech with the study split lists.

[huggingface.co/datasets/Digisensus/lithuanian-phone-speech-liepa-3-429h-punctuated · v2.0](https://huggingface.co/datasets/Digisensus/lithuanian-phone-speech-liepa-3-429h-punctuated)

Dataset CC-BY-4.0

### [Digisensus/lithuanian-dialect-asr-eval](https://huggingface.co/datasets/Digisensus/lithuanian-dialect-asr-eval)

Evaluation records for the public test sets: references, the hypotheses of all 21 evaluated systems, scores and the control pool.

[huggingface.co/datasets/Digisensus/lithuanian-dialect-asr-eval · v1.0](https://huggingface.co/datasets/Digisensus/lithuanian-dialect-asr-eval)

Model CC-BY-4.0

### [Digisensus/parakeet-tdt-0.6b-lt-study-s](https://huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-s)

Telephone speech only. Run 1 of the condition, as a NeMo checkpoint with its training config.

[huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-s · v1.0](https://huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-s)

Model CC-BY-4.0

### [Digisensus/parakeet-tdt-0.6b-lt-study-s-d80-dial](https://huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-s-d80-dial)

Telephone + dialect speech, dialect spelling. Run 1 of the condition, as a NeMo checkpoint with its training config.

[huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-s-d80-dial · v1.0](https://huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-s-d80-dial)

Model CC-BY-4.0

### [Digisensus/parakeet-tdt-0.6b-lt-study-s-d80-std](https://huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-s-d80-std)

Telephone + dialect speech, standard spelling. Run 1 of the condition, as a NeMo checkpoint with its training config.

[huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-s-d80-std · v1.0](https://huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-s-d80-std)

Model CC-BY-4.0

### [Digisensus/parakeet-tdt-0.6b-lt-study-s-x80](https://huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-s-x80)

Telephone + other spontaneous speech (control). Run 1 of the condition, as a NeMo checkpoint with its training config.

[huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-s-x80 · v1.0](https://huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-s-x80)

Model CC-BY-4.0

### [Digisensus/parakeet-tdt-0.6b-lt-study-d80-dial](https://huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-d80-dial)

Dialect speech only, dialect spelling. Run 1 of the condition, as a NeMo checkpoint with its training config.

[huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-d80-dial · v1.0](https://huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-d80-dial)

Model CC-BY-4.0

### [Digisensus/parakeet-tdt-0.6b-lt-study-d80-std](https://huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-d80-std)

Dialect speech only, standard spelling. Run 1 of the condition, as a NeMo checkpoint with its training config.

[huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-d80-std · v1.0](https://huggingface.co/Digisensus/parakeet-tdt-0.6b-lt-study-d80-std)

**The six released models: run 1 of each headline condition. WER (%) on the dialect test (three reference views) and the LIEPA-3 telephone test; the report averages repeated runs where stated.**

| Model | Training data | Dialect ref. | Standard ref. | Either | Telephone |
| --- | --- | --- | --- | --- | --- |
| parakeet-tdt-0.6b-lt-study-s | Telephone only | 42.11 | 36.40 | 34.83 | 10.76 |
| parakeet-tdt-0.6b-lt-study-s-d80-dial | Telephone + dialect, dialect spelling | 31.28 | 36.29 | 27.60 | 10.80 |
| parakeet-tdt-0.6b-lt-study-s-d80-std | Telephone + dialect, standard spelling | 37.52 | 27.06 | 26.06 | 10.65 |
| parakeet-tdt-0.6b-lt-study-s-x80 | Telephone + other spontaneous (control) | 39.90 | 33.86 | 32.18 | 10.31 |
| parakeet-tdt-0.6b-lt-study-d80-dial | Dialect only, dialect spelling | 27.18 | 35.32 | 24.43 | 28.20 |
| parakeet-tdt-0.6b-lt-study-d80-std | Dialect only, standard spelling | 35.15 | 22.72 | 21.91 | 19.63 |

## References

1.  Vilnius University, Vytautas Magnus University and the Institute of the Lithuanian Language (2026). Large Lithuanian language speech corpus (LIEPA-3). CLARIN-LT repository. CC BY 4.0. [hdl.handle.net/20.500.11821/101](https://hdl.handle.net/20.500.11821/101)
2.  NVIDIA NeMo team (2025). parakeet-tdt-0.6b-v3: a multilingual FastConformer-TDT speech recognition model. [huggingface.co/nvidia/parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3)
3.  Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C. and Sutskever, I. (2023). Robust speech recognition via large-scale weak supervision. Proc. ICML.
4.  Conneau, A., Ma, M., Khanuja, S., Zhang, Y., Axelrod, V., Dalmia, S., Riesa, J., Rivera, C. and Bapna, A. (2022). FLEURS: Few-shot learning evaluation of universal representations of speech. Proc. IEEE SLT.
5.  Ardila, R., Branson, M., Davis, K., Henretty, M., Kohler, M., Meyer, J., Morais, R., Saunders, L., Tyers, F. M. and Weber, G. (2020). Common Voice: A massively-multilingual speech corpus. Proc. LREC.
6.  Bisani, M. and Ney, H. (2004). Bootstrap estimates for confidence intervals in ASR performance evaluation. Proc. ICASSP.
7.  Hinsvark, A. et al. (2021). Accented speech recognition: A survey. arXiv:2104.10747. [arxiv.org/abs/2104.10747](https://arxiv.org/abs/2104.10747)
8.  Ali, A., Vogel, S. and Renals, S. (2017). Speech recognition challenge in the wild: Arabic MGB-3. Proc. IEEE ASRU, 316–322. [doi.org/10.1109/ASRU.2017.8268952](https://doi.org/10.1109/ASRU.2017.8268952)
9.  Plüss, M., Neukom, L., Scheller, C. and Vogel, M. (2021). Swiss Parliaments Corpus, an automatically aligned Swiss German speech to Standard German text corpus. Proc. SwissText 2021. [ceur-ws.org/Vol-2957/paper3.pdf](https://ceur-ws.org/Vol-2957/paper3.pdf)
10.  Balode, L. and Holvoet, A. (2001). The Lithuanian language and its dialects. In Circum-Baltic Languages, Volume 1, 41–79. John Benjamins. [doi.org/10.1075/slcs.54.05bal](https://doi.org/10.1075/slcs.54.05bal)
11.  Bakšienė, R., Čepaitienė, A., Jaroslavienė, J. and Urbanavičienė, J. (2024). Standard Lithuanian. Journal of the International Phonetic Association, 54(1), 414–444. [doi.org/10.1017/S0025100323000105](https://doi.org/10.1017/S0025100323000105)
12.  Judžentytė-Šinkūnienė, G. and Nikartaitė, S. (2020). The concept of Samogitianness in the Northern Samogitian dialect. Valoda: nozīme un forma, 11, 39–62. [doi.org/10.22364/vnf.11.03](https://doi.org/10.22364/vnf.11.03)

Cite this study as: Jarašiūnas, S. (2026). Lithuanian speech recognition: effects of dialect training and transcript spelling. Digisensus, Vilnius. https://digisensus.com/lithuanian-dialect-speech-recognition/

## Appendix: complete results

WER (%) for every evaluated model on every test set; lower is better. Rows marked with ² are means of two training runs. The three dialect-test columns accept dialect spelling, standard spelling or either form. "Training speakers" means withheld recordings from speakers in the dialect training pool, scored with either form accepted; speaker exposure depends on the training subset. Model names: S = 420.17 hours of telephone speech; D12…D80 = dialect speech by hours; X80 = other spontaneous speech; .dial / .std = transcript spelling; -A, -D, -S, -Z = region left out.

**Word error rate (%) for every evaluated system. Stock models receive no study fine-tuning.**

| Model | Dialect test: dialect ref. | Dialect test: standard ref. | Dialect test: either | Training speakers: either | Telephone | FLEURS | Common Voice |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Original models, no study fine-tuning |  |  |  |  |  |  |  |
| parakeet-tdt-0.6b-v3 (stock) | 71.67 | 69.09 | 68.15 | 69.10 | 42.53 | 22.10 | 16.52 |
| Whisper large-v3 | 59.36 | 54.57 | 53.15 | 51.77 | 27.43 | 23.17 | 28.09 |
| Dialect-spelling transcripts, by hours |  |  |  |  |  |  |  |
| S² | 41.62 | 35.82 | 34.22 | 35.48 | 10.61 | 28.64 | 39.97 |
| S+D12.dial | 38.08 | 38.88 | 33.30 | 34.79 | 10.61 | 24.32 | 30.35 |
| S+D25.dial | 34.77 | 37.29 | 30.29 | 32.59 | 10.61 | 21.89 | 24.24 |
| S+D50.dial | 32.77 | 37.09 | 28.96 | 31.32 | 10.49 | 23.22 | 29.32 |
| S+D80.dial² | 31.30 | 36.73 | 27.82 | 29.88 | 10.76 | 22.91 | 26.88 |
| Control: other spontaneous speech |  |  |  |  |  |  |  |
| S+X80² | 40.04 | 34.10 | 32.40 | 33.61 | 10.34 | 28.48 | 37.34 |
| Standard-spelling transcripts |  |  |  |  |  |  |  |
| S+D25.std | 38.31 | 29.14 | 27.91 | 29.67 | 10.49 | 24.22 | 30.50 |
| S+D50.std | 37.67 | 27.47 | 26.39 | 27.94 | 10.46 | 20.35 | 18.99 |
| S+D80.std² | 37.48 | 26.94 | 25.97 | 27.40 | 10.56 | 24.19 | 25.80 |
| Dialect data only |  |  |  |  |  |  |  |
| D80.dial | 27.18 | 35.32 | 24.43 | 25.24 | 28.20 | 41.54 | 39.74 |
| D80.std | 35.15 | 22.72 | 21.91 | 22.68 | 19.63 | 30.34 | 26.71 |
| One region left out |  |  |  |  |  |  |  |
| S+D80-A.std | 38.00 | 27.95 | 26.92 | 28.69 | 10.53 | 21.38 | 20.33 |
| S+D80-D.std | 38.19 | 28.10 | 27.09 | 28.01 | 10.58 | 25.64 | 35.81 |
| S+D80-S.std | 38.13 | 27.58 | 26.67 | 28.09 | 10.61 | 27.37 | 34.35 |
| S+D80-Z.std | 37.44 | 27.37 | 26.20 | 28.14 | 10.64 | 24.43 | 32.43 |

Scorer: score.py 1.3 (2026-10-04), 10,000 bootstrap resamples. 21 systems evaluated in total.

## Read the full study

PDF, 20 pages, 5 October 2026. Questions: saulius@digisensus.com

[PDF](/assets/research/lithuanian-dialect-asr-study-2026-10-05.pdf) [All research](/research/)

---

- Canonical: https://digisensus.com/lithuanian-dialect-speech-recognition/
- Lithuanian version: https://digisensus.com/lt/lietuviu-tarmiu-snekos-atpazinimas/
- Book a demo: https://calendar.app.google/MM4rtjEa8ctThvoZ8
- Contact: saulius@digisensus.com · +370 620 69969
