The evidence
Scoreboard
The original goal has been met: match GPT-NL on all six published benchmarks. Five times above it, level on dbrd. WikiLingua-nl falls outside the comparison, because the metric changed there. Walnoot 8B has 8.05 billion parameters against 26.03 billion for GPT-NL.

All tasks
How does Walnoot 8B compare to GPT-NL?
| Task | Kind | Metric | Walnoot 8B | GPT-NL | Margin | Outcome |
|---|---|---|---|---|---|---|
| dbrda | Sentiment | MCC | 90.56 | 90 | +0.56 | tie |
| conll-nlb | Named entity recognition | micro-F1 (no-MISC) | 46.52 | 36 | +10.52 | win |
| squad-nl | Reading comprehension | EM | 62.08 | 51 | +11.08 | win |
| scala-nl | Linguistic acceptability | MCC | 34.41 | 19 | +15.41 | win |
| hellaswag-nl | Commonsense reasoning | MCC | 22.46 | −2 | +24.46 | win |
| mmlu-nl | Knowledge | MCC | 37.16 | 2 | +35.16 | win |
| Without a basis for comparison | ||||||
| wikilingua-nlc | Summarisation | ChrF3++ | 34.18 | no published score | no published score | metric change |
| duidelijke-taald | Plain language | METEOR | 52.95 | no published score | no published score | no publication |
| valeu-nle | European values | EuropeanValues | 2.02 | no published score | no published score | no publication |
| mbbq-nlf | Stereotyping | bias-corrected accuracy | 27.72 | no published score | no published score | no publication |
Notes to the table
- On points we are ahead here, with 90.56 against 90. Our confidence interval runs from 89.76 to 91.35 and their figure falls inside it, so our rule calls this a tie. An earlier re-measurement showed that the GPU parallel layout alone can shift this figure by 0.32 of a point.
- The metric is written out in full because two scales exist side by side. We test on micro-F1 without MISC, because the published 36 of GPT-NL sits on that same scale. Including MISC, Walnoot 8B measures 29.13. That second figure is in the raw logs and therefore it is here as well.
- On this task the metric changed between the two measurements. GPT-NL published 61 on BERTScore and we measure 34.18 on ChrF3++. Those figures sit on different scales, so we never place them side by side. If GPT-NL runs its model on ChrF3++, we will compare after all.
- This task measures rewriting into plain language. GPT-NL published no figure here, so there is nothing to set alongside. This table carries the primary metric of the harness for each task. Here that is METEOR; only on squad-nl do we depart from that, because the published figure of GPT-NL sits on that other measure.
- 2.02 is the lowest score in the whole suite. We have marked this dataset internally as not citable; our register of lessons carries that as EV-5 with status open. GPT-NL published no figure here.
- The metric corrects the accuracy for the preference that the model shows on ambiguous questions. GPT-NL published no figure here. The score is here because this task belongs to the official Dutch suite of EuroEval 17.6.0.
A dash means GPT-NL published no figure there, so there is no margin to compute.
Source of the GPT-NL columnTNO, GPTNL-DEL-4002, December 2025, pages 44 to 49. Measured on EuroEval 15.16.0, pinned in §3.6.3 on page 96.
The decision rule
When does a task count as won?
A win only counts as a win when our whole confidence interval lies above their figure. A loss only when it lies entirely below it. If their figure falls inside our interval, the result is a tie. That holds even on a task where we are ahead on points.
The interval per task
Every task we measure has an interval. Where GPT-NL published a figure, it stands next to it.
| Task | Walnoot 8B | 95% interval | GPT-NL | Falls |
|---|---|---|---|---|
| dbrd | 90.56 | 89.76–91.35 | 90 | within |
| conll-nl | 46.52 | 44.49–48.55 | 36 | below |
| squad-nl | 62.08 | 61.42–62.73 | 51 | below |
| scala-nl | 34.41 | 31.00–37.81 | 19 | below |
| hellaswag-nl | 22.46 | 20.73–24.19 | −2 | below |
| mmlu-nl | 37.16 | 36.19–38.13 | 2 | below |
| wikilingua-nl | 34.18 | 32.74–35.62 | no published score | nothing to test |
| duidelijke-taal | 52.95 | 51.11–54.78 | no published score | nothing to test |
| valeu-nl | 2.02 | 1.41–2.64 | no published score | nothing to test |
| mbbq-nl | 27.72 | 26.17–29.27 | no published score | nothing to test |
Measurement regime
How was it measured?
We measured our column with EuroEval 17.6.0 on the test split, in 10 iterations and with bf16 weights. The GPT-NL column is their own publication of December 2025 on EuroEval 15.16.0, interim scores of the base model, published as point scores without an interval for the figures we take over. The harness is two major versions apart. Across versions the dataset content is not guaranteed to be identical. That is why we state explicitly with every comparison which version was used.
- Regime
- test split, 10 iterations, bf16, zero failed instances
The GPT-NL column
Why do we not measure GPT-NL ourselves?
GPT-NL publishes no weights, so nobody outside the consortium can run that model. We have therefore never measured it ourselves. This table sets our measured figure next to their published figure.
GPT-NL, the model built with public funding for Dutch, was the goal and the yardstick of this project. We adopted its requirements for the data. Apertus is the base model on which Walnoot was further trained; the instruction model of Apertus itself made different choices in its post-training than we did, with a broader data mix. Other open models made different choices in data and provenance. This page compares with GPT-NL, because that was the question we asked.
Recompute
Check every figure yourself
pip install "euroeval[all]==17.6.0" && euroeval --model WAINUT/walnoot-8b-instruct --language nl --num-iterations 10 --dataset dbrd --dataset scala-nl --dataset conll-nl --dataset squad-nl --dataset wiki-lingua-nl --dataset mmlu-nl --dataset hellaswag-nl --dataset duidelijke-taal --dataset valeu-nl --dataset mbbq-nl --evaluate-test-splitThe command installs EuroEval 17.6.0, the same version we measured with.
The evaluation and the comparison with GPT-NL were independently recomputed by an outside tester on the latest version of EuroEval.
Our thanks to Edwin Rijgersberg for the feedback.