Digisensus Digisensus
Open source

Open-source apps, models and datasets

We share what we build along the way: a free call recorder for Mac, the speech and language models we tune, and the Lithuanian speech data we prepare to train them. Everything here is free to use under its license.

{ } OPEN SOURCE APPS Digisensus Recorder LLM MODELS Qwen3.8 27B ASR MODELS Parakeet · Lithuanian DATASETS 429 h + 100 h

Apps

1 item

Software you can install and use today, with the full source code.

Digisensus Recorder

Digisensus/digisensus-recorder
GPL-3.0

A free call recorder and AI meeting note taker for Mac. It records Zoom, Google Meet, Microsoft Teams, FaceTime, WhatsApp and phone calls with no bot joining the call, then gives you transcripts that show who said what.

  • Recordings stay on your Mac
  • macOS 15 or later, Apple silicon

LLM models

1 item

Large language models we tune and quantize for our own GPU cluster, shared with the setup that serves them.

Qwen3.8 27B Digisensus NVFP4

Digisensus/Qwen3.8-27B-Digisensus-NVFP4-RTX5090-256K
Apache-2.0

Qwen3.8-27B on a single RTX 5090 with the full 256K-token context. Each part of the network gets the precision it needs, so it stays closer to the original model at the same memory.

  • About 115 tokens per second for one user
  • 24.5% closer to the original than another NVFP4 quant tuned for the RTX 5090
  • Made for tool calling, JSON output and reasoning

ASR models

1 item

Speech recognition (speech-to-text) models fine-tuned for Lithuanian, with honest results on held-out and real-world test sets.

Parakeet Lithuanian Phone 0.6B

Digisensus/parakeet-tdt-0.6b-lt-phone
CC-BY-4.0

Lithuanian speech recognition for telephone speech, fine-tuned from NVIDIA Parakeet TDT 0.6B v3 on 426 hours of calls. It writes readable text: punctuation, capitals, and numbers, dates and amounts as digits.

  • Word error rate 12.4% on held-out LIEPA-3 phone speech, down from 42.7%
  • 39.3% on independent real calls, down from 59.5%: a research release, not yet production grade

Datasets

2 items

Lithuanian speech datasets built from the openly licensed LIEPA-3 corpus, with transcripts in written form so models learn to output readable text.

Lithuanian Phone Speech 429 h

Digisensus/lithuanian-phone-speech-liepa-3-429h-punctuated
CC-BY-4.0

429 hours and 417,000 utterances of spontaneous Lithuanian telephone speech. To our knowledge, the largest openly licensed Lithuanian phone speech dataset with written-form transcripts.

  • Written-form and normalized transcripts side by side
  • 16 kHz audio in Parquet, with train, validation and test splits

Lithuanian Dialect Speech 100 h

Digisensus/lithuanian-dialect-speech-liepa-3-100h-punctuated
CC-BY-4.0

100.5 hours and 111,970 utterances of spontaneous dialect speech from 117 speakers in Aukštaitija, Žemaitija, Dzūkija and Suvalkija, with dialect word forms kept as spoken.

  • Three transcripts per clip: written form, normalized, and phonetic with stress marks
  • Speaker IDs and regions for every clip

Why we publish

Speech technology for smaller languages moves faster when data and models are shared. We publish what we can, including results that show where a model still falls short, so others can build on it and check our work.

github.com/Digisensus huggingface.co/Digisensus