Do not believe us. Run it yourself.
Dutch language model that you run in your own environment
Crack it open.See what is inside.
The model
Walnoot 8B is a Dutch language model from WAINUT. You download the model and run it in your own environment, including for applications where data may not leave that environment. Everything we measure and publish, you can check yourself.
Scoreboard
The original goal has been met: match GPT-NL on all six published benchmarks. Five times above it, level on dbrd.
- Model
- Walnoot 8B Instruct
- Version
- 1.0
- Licence
- Apache 2.0
- Evaluation
- EuroEval 17.6.0
WAINUT
Why we built this.

Photo: Yasin Celik
Every day we sit down with organisations that handle sensitive data. We kept running into the same problem there. A letter, case file or document comes in and you want to do something with it, but the information isn't allowed to leave your own environment. In that situation a frontier model, like GPT or Claude, simply isn't always the right tool. So that kind of application gets shelved. And we see that at very different organisations, in more than one sector.
At the same time, especially at government bodies, we kept hearing the same name: GPT-NL. The need for a Dutch language model that organisations can run and check themselves turned out to be big there.
And then an American provider made its newest model temporarily available only in the United States. That made it clear once again how dependent we in Europe are on a small number of foreign companies.
We have an opinion about that. And then you can keep talking about it, or you can try to build something yourself.
So we set one concrete goal up front: at least match GPT-NL on six comparable tasks. From there we started building a compact Dutch language model for bounded tasks.
Why Walnoot
What it gives you.

Dutch first
Walnoot has had continued pretraining on Dutch material and was measured on ten Dutch tasks, each with its confidence interval alongside.
Stay in control
Sensitive material does not have to leave your own environment, because you download the weights and run them on your own machine or server.
Know what you are using
The provenance per source, the recipe, the evaluation logs and the limits are open, so you do not have to take our word for it.
Open under Apache 2.0
The weights are under Apache 2.0, including for commercial use. Because Walnoot is open-weight, what you download stays yours, even if you move to another supplier later.
Grows with you
The licence and the recipe let you use Walnoot as it is and adapt it later if your organisation needs more.
Built on
We did not start from scratch.

European base
Apertus-8B-2509 from the Swiss AI Initiative, published under Apache 2.0.
What WAINUT added
Nothing. This is the work of its makers.
Continued pretraining on Dutch
WAINUT trained that model further on Dutch-language material, with a cooling-down phase at the end. The mix was 70 per cent Dutch, 20 per cent English and 10 per cent code.
What WAINUT added
A Dutch base of our own. We do not publish it separately.
Instruction layer
On top of that the model learned to follow Dutch instructions: answering questions about a document, extracting information from a text and classifying texts.
What WAINUT added
An instruction layer of our own, with a chat template and a system prompt.
Post-training with feedback
Finally the model was steered on its own answers that a judging step approved.
What WAINUT added
The last step of the recipe. The technical term is rejection fine-tuning.
The instruction model of Apertus itself chose a different data mix and a different post-training. We chose data that you can trace back by source and a recipe that you can check. The training run took less than a week and the first model was running within about a month.

Scoreboard
What we measured.
The original goal has been met: match GPT-NL on all six. Five times above it, level on dbrd.
squad-nl
Reading comprehension
- Our score
- 62.08
- GPT-NL
- 51
win
conll-nl
Named entity recognition
- Our score
- 46.52
- GPT-NL
- 36
win
scala-nl
Linguistic acceptability
- Our score
- 34.41
- GPT-NL
- 19
win
mmlu-nl
Knowledge
- Our score
- 37.16
- GPT-NL
- 2
win
hellaswag-nl
Commonsense reasoning
- Our score
- 22.46
- GPT-NL
- −2
win
dbrd
Sentiment
- Our score
- 90.56
- GPT-NL
- 90
tie
More than three times smaller than GPT-NL and still higher on five of the six.
Walnoot 8B has 8.05 billion parameters against 26.03 billion for GPT-NL.
The evaluation and the comparison with GPT-NL were independently recomputed by an outside tester on the latest version of EuroEval.
Walnoot 8B is built for work that can be bounded.
Read the limits on the model cardView all ten tasks with their interval
Why not simply
Three questions we asked ourselves too.
Why not simply a frontier API?
Often that is simply the best choice. For complex analyses, broad reasoning work and tasks with a lot of context, frontier models such as GPT and Claude are stronger.
Walnoot is meant for something else: bounded tasks where data may not leave your own environment. Think of document question answering, information extraction or automatic classification of documents.
The data stay where you want them and you decide for yourself where the model runs.
Why not simply Apertus?
Apertus is the European base on which Walnoot is built and we are simply transparent about that. WAINUT built on it in three steps: additional Dutch pretraining, an instruction layer of our own in Dutch and post-training on approved answers (rejection fine-tuning).
On top of that we evaluated Walnoot on ten Dutch tasks. We also publish the provenance of the sources we used and the recipe with which the model was made.
The instruct model of Apertus was made with a different data mix and different post-training. We made our own choices there: Dutch data whose provenance can be traced per source and a recipe in which you can see what happens at every step.
Is GPT-NL not a low bar?
That is a fair question. GPT-NL was the yardstick we fixed for ourselves beforehand. Our goal was clear: to perform at least as well on the comparable tasks. We did not move that bar during the build when results disappointed.
On the scoreboard we show all ten tasks, including the four for which no direct comparison with GPT-NL is possible.
The evidence layer
Everything above can be checked.

Where the data comes from.
Our Dutch training step runs on the GPT-NL Public Corpus under CC BY 4.0, with 28 subsets at a pinned revision and per source the legal basis and the processing steps included. What is not in it is stated too, such as the newspaper archives we left out on quality grounds.
What is actually open.
| Artefact | Licence | Access |
|---|---|---|
| Weights (8B) | Apache 2.0 | Public, without registration |
| Evaluation logs | Apache 2.0 | Public, raw jsonl |
| Training recipe | Apache 2.0 | Public |
| Evaluation configuration | Apache 2.0 | Public, pinned to 17.6.0 |
| Our Dutch corpus | Documented per source | Provenance public, source data not published by us |
| Training data of the base model | Terms per source | Documented by the Swiss AI Initiative, not ours to publish |
Reproduce
How you run it yourself.
One file and a guide per program. If you want to recompute the figures, the full command with the pinned evaluation version is there alongside.
The command
euroeval --model WAINUT/walnoot-8b-instruct --num-iterations 10 --evaluate-test-split --dataset …Shortened view. The full command sits behind the copy button.
Show the full command
pip install "euroeval[all]==17.6.0" && euroeval --model WAINUT/walnoot-8b-instruct --language nl --num-iterations 10 --dataset dbrd --dataset scala-nl --dataset conll-nl --dataset squad-nl --dataset wiki-lingua-nl --dataset mmlu-nl --dataset hellaswag-nl --dataset duidelijke-taal --dataset valeu-nl --dataset mbbq-nl --evaluate-test-splitFurther in the dossier