FPSC (Faroese Parliament Speech Corpus) is a large Faroese speech dataset created from publicly available recordings from the Faroese Parliament. The corpus contains around 1,600 hours of speech consisting of more than 89,000 parliamentary speeches, 368 parliamentary sessions, and 75 different speakers.
Each speech segment includes an audio recording together with an automatically generated transcription. The dataset also contains metadata such as speaker, age, gender, place of origin, dialect, political party, session, date, topic, and type of speech.
The transcriptions were generated automatically using several Faroese-adapted speech recognition models. The outputs from the models were combined using ROVER voting to select the transcription considered most reliable. The transcriptions have not been manually verified, so the corpus should be regarded as weakly supervised data rather than as a fully corrected speech corpus.
FPSC is intended for research and development in areas such as Faroese automatic speech recognition (ASR), speech technology for low-resource languages, dialect research, sociolinguistics, continued training of speech models, and multilingual transfer learning.
The recordings and original metadata come from the publicly available material of the Faroese Parliament.
Release: 2026
Authors: Dávid í Lág, Barbara Scalvini, Carlos Mena and Jón Guðnason
Paper: FPSC: A Sustainable Pipeline for Building a Faroese Parliamentary Speech Corpus, LREC 2026.
Contact: davidl@setur.fo




