Language Model Evaluation
Á talvuni niðanfyri sæst ein sjálvvirkandi meting av, hvussu væl ymiskir málmyndlar duga føroyskt. Tað er EuroEval, sum hevur kannað málmyndlarnar og gjørt yvirlitið. Úrslitini byggja á eina røð av máluppgávum, har málmyndlarnir verða royndir í eitt nú málfatan, spurningum, tekstflokking og øðrum uppgávum. Endamálið er at geva eina lætta og yvirskipaða mynd av, hvussu málmyndlarnir klára seg á føroyskum samanborið við hvønn annan.
The results are indicative and should be viewed as a comparison of how well the language models perform on the tasks included in EuroEval (you can read more about these tasks here). Máltøknidepilin hevur verið við til at gera summar av hesum uppgávunum. Úrslitini siga ikki einsamøll alt um, hvussu góður ein málmyndil er til allar hugsandi uppgávur, men tey kunnu geva eina góða ábending um førleikarnar hjá málmyndlinum á føroyskum.
What Do the Headings Mean?
Generative Leaderboard
This compares generative language models. These are language models that can both understand and write text. The result is based on several different tasks for Faroese.
NLU Leaderboard
NLU stands for Natural Language Understanding. This focuses especially on how well the language models understand language, for example text classification, questions, and other language tasks. This list may also include language models that are not designed to generate text (generative language models).
Generative Scatter Plot / NLU Scatter Plot
This is a visual representation of the same results. Each point represents a language model. The horizontal axis shows the size of the language model (number of parameters), while the vertical axis shows the Rank score. This makes it possible to see whether larger language models also achieve better results.
What do the numbers mean?
Rank score is an overall measure of how well a language model performs compared with the other language models in the comparison. A lower number is better, and a score close to 1.00 means that the language model is among the best-performing models.
For example, 1.42 ± 0.07 means that the Rank score is approximately 1.42. The number after ± indicates the uncertainty in the measurement. If two language models have very similar results, the difference may therefore be too small to say with certainty that one is better than the other.
Rank: 1st, 2nd, 3rd …
This shows where the language model is placed on the list. Several language models can share the same rank if their results are so close that there is no clear difference between them.
Other columns
Parameters shows the size of the language model. A number such as 31B means approximately 31 billion parameters.
Open-weight shows whether the weights of the language model are publicly available.
Commercial shows whether the language model may be used for commercial purposes.
Trained from scratch shows whether the language model was trained from scratch rather than built on top of another language model.

