Skip to main content

The evidence

Scoreboard

The original goal has been met: match GPT-NL on all six published benchmarks. Five times above it, level on dbrd. WikiLingua-nl falls outside the comparison, because the metric changed there. Walnoot 8B has 8.05 billion parameters against 26.03 billion for GPT-NL.

Harness
EuroEval 17.6.0
Model
Walnoot 8B Instruct 1.0
Updated
17 September 2026
A walnut beside the jaws of a steel calliper, seen from above on a dark slate slab.

All tasks

How does Walnoot 8B compare to GPT-NL?

10 tasks, narrowest margin firstEuroEval 17.6.0All 10 Dutch tasks of EuroEval 17.6.0. Columns: task, kind, metric, our score, the published figure from GPT-NL, the margin and the outcome. The 95% interval is in the table below the decision rule.
TaskKindMetricWalnoot 8BGPT-NLMarginOutcome
dbrdSentimentMCC90.5690+0.56tie
conll-nlNamed entity recognitionmicro-F1 (no-MISC)46.5236+10.52win
squad-nlReading comprehensionEM62.0851+11.08win
scala-nlLinguistic acceptabilityMCC34.4119+15.41win
hellaswag-nlCommonsense reasoningMCC22.46−2+24.46win
mmlu-nlKnowledgeMCC37.162+35.16win
Without a basis for comparison
wikilingua-nlSummarisationChrF3++34.18no published scoreno published scoremetric change
duidelijke-taalPlain languageMETEOR52.95no published scoreno published scoreno publication
valeu-nlEuropean valuesEuropeanValues2.02no published scoreno published scoreno publication
mbbq-nlStereotypingbias-corrected accuracy27.72no published scoreno published scoreno publication
Notes to the table
  1. On points we are ahead here, with 90.56 against 90. Our confidence interval runs from 89.76 to 91.35 and their figure falls inside it, so our rule calls this a tie. An earlier re-measurement showed that the GPU parallel layout alone can shift this figure by 0.32 of a point.
  2. The metric is written out in full because two scales exist side by side. We test on micro-F1 without MISC, because the published 36 of GPT-NL sits on that same scale. Including MISC, Walnoot 8B measures 29.13. That second figure is in the raw logs and therefore it is here as well.
  3. On this task the metric changed between the two measurements. GPT-NL published 61 on BERTScore and we measure 34.18 on ChrF3++. Those figures sit on different scales, so we never place them side by side. If GPT-NL runs its model on ChrF3++, we will compare after all.
  4. This task measures rewriting into plain language. GPT-NL published no figure here, so there is nothing to set alongside. This table carries the primary metric of the harness for each task. Here that is METEOR; only on squad-nl do we depart from that, because the published figure of GPT-NL sits on that other measure.
  5. 2.02 is the lowest score in the whole suite. We have marked this dataset internally as not citable; our register of lessons carries that as EV-5 with status open. GPT-NL published no figure here.
  6. The metric corrects the accuracy for the preference that the model shows on ambiguous questions. GPT-NL published no figure here. The score is here because this task belongs to the official Dutch suite of EuroEval 17.6.0.

A dash means GPT-NL published no figure there, so there is no margin to compute.

Source of the GPT-NL columnTNO, GPTNL-DEL-4002, December 2025, pages 44 to 49. Measured on EuroEval 15.16.0, pinned in §3.6.3 on page 96.

The decision rule

When does a task count as won?

A win only counts as a win when our whole confidence interval lies above their figure. A loss only when it lies entirely below it. If their figure falls inside our interval, the result is a tie. That holds even on a task where we are ahead on points.

The interval per task

Every task we measure has an interval. Where GPT-NL published a figure, it stands next to it.

The 95% interval per task. Columns: task, the figure for Walnoot 8B, the 95% interval, the published figure from GPT-NL and whether that figure falls within the bounds. Without a published figure from GPT-NL there is nothing to check.
TaskWalnoot 8B95% intervalGPT-NLFalls
dbrd90.5689.76–91.3590within
conll-nl46.5244.49–48.5536below
squad-nl62.0861.42–62.7351below
scala-nl34.4131.00–37.8119below
hellaswag-nl22.4620.73–24.19−2below
mmlu-nl37.1636.19–38.132below
wikilingua-nl34.1832.74–35.62no published scorenothing to test
duidelijke-taal52.9551.11–54.78no published scorenothing to test
valeu-nl2.021.41–2.64no published scorenothing to test
mbbq-nl27.7226.17–29.27no published scorenothing to test

Measurement regime

How was it measured?

We measured our column with EuroEval 17.6.0 on the test split, in 10 iterations and with bf16 weights. The GPT-NL column is their own publication of December 2025 on EuroEval 15.16.0, interim scores of the base model, published as point scores without an interval for the figures we take over. The harness is two major versions apart. Across versions the dataset content is not guaranteed to be identical. That is why we state explicitly with every comparison which version was used.

Regime
test split, 10 iterations, bf16, zero failed instances

The GPT-NL column

Why do we not measure GPT-NL ourselves?

GPT-NL publishes no weights, so nobody outside the consortium can run that model. We have therefore never measured it ourselves. This table sets our measured figure next to their published figure.

GPT-NL, the model built with public funding for Dutch, was the goal and the yardstick of this project. We adopted its requirements for the data. Apertus is the base model on which Walnoot was further trained; the instruction model of Apertus itself made different choices in its post-training than we did, with a broader data mix. Other open models made different choices in data and provenance. This page compares with GPT-NL, because that was the question we asked.

Recompute

Check every figure yourself

pip install "euroeval[all]==17.6.0" && euroeval --model WAINUT/walnoot-8b-instruct --language nl --num-iterations 10 --dataset dbrd --dataset scala-nl --dataset conll-nl --dataset squad-nl --dataset wiki-lingua-nl --dataset mmlu-nl --dataset hellaswag-nl --dataset duidelijke-taal --dataset valeu-nl --dataset mbbq-nl --evaluate-test-split

The command installs EuroEval 17.6.0, the same version we measured with.

The evaluation and the comparison with GPT-NL were independently recomputed by an outside tester on the latest version of EuroEval.

Our thanks to Edwin Rijgersberg for the feedback.