📊 Full opportunity report: Minerva. The opposite path. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Italy’s Minerva-3B, a sovereign LLM trained from scratch on 2.5 trillion tokens, shows limited academic performance despite large-scale investment. This questions assumptions about language-specific model development.
Italy’s Minerva-3B, a large-scale sovereign language model trained entirely from scratch on 2.5 trillion tokens with approximately 50% Italian content, scored just 4.9% on the INVALSI Italian school-exam benchmark, raising questions about the effectiveness of large-scale native-language investment in AI development.
Developed by Sapienza University of Rome and Italy’s national research institutions, Minerva-3B was designed to be a fully open, native-Italian model, with weights, data, and code released publicly. It was trained on a dataset of 2.5 trillion tokens, with half in Italian, and involved 15 researchers and technical support from NVIDIA and CINECA, utilizing Italy’s CINECA supercomputing infrastructure.
Despite the large investment and the model outperforming comparable multilingual models on Italian benchmarks, Minerva-3B’s performance on the INVALSI academic test was notably poor, with a score near chance. Researchers concluded that dataset size and parameter count are more critical for complex language tasks than the proportion of native-language data alone.
This empirical result suggests that even significant native-language training at large scales may not produce the expected depth of country-specific knowledge, challenging assumptions about the direct correlation between investment scale and model performance in specialized tasks.
Minerva.
The opposite
path.
Italy spent years building a European sovereign LLM from scratch. Then Minerva-3B scored 4.9% on the INVALSI Italian school exam.
Where AMÁLIA layered Portuguese specialization onto a multilingual foundation, Minerva trained from scratch on 2.5 trillion tokens with approximately 50% Italian content. Where AMÁLIA’s weights are not yet public, Minerva published weights, training data, and code as truly-open from day one. By every institutional measure, the Italian approach worked. But the empirical results contain a finding the press coverage has been quiet about — and it has implications that extend well beyond Italy.
Same problem. Opposite path.
European sovereign-LLM development has two primary architectural approaches. Italy chose from scratch with substantial native-language foundation. Portugal chose continuation pre-training of a multilingual model. The structural comparison surfaces what each commitment actually requires operationally.
The comparison is not “Italy did it better than Portugal.” Both projects respond to the same structural problem with different architectural strategies under different institutional and economic constraints. Italy’s national-AI investment is structurally larger by an order of magnitude — and Minerva is the visible artifact of that scale.

Accelerate Everything with Tensor Cores: A Developer’s Guide to High-Performance AI, Efficient Training, and Scalable Models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
4.9% on INVALSI. The bitter lesson surfaces.
In June 2024, researchers evaluated Minerva-3B on the Italian school-exam benchmark. The result was unambiguous. This is not a critique of Minerva — it is a critique of the public discourse around what Minerva’s empirical results actually demonstrate.

NVIDIA HPE Tesla P40 24GB Computational Accelerator (Renewed)
This Certified Refurbished product is tested and certified to work and look like new by a specialized third-party…
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
350M to 7B. Four parameter scales, one architecture.
The Minerva model family covers four parameter tiers, each with specific training corpora. Each scale level reveals what the from-scratch path actually requires at different operating points.
Italian + English
100B English
~50% English
+ 200B code

Data Analysis with Open Source Tools: A Hands-On Guide for Programmers and Data Scientists
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three answers. Same question.
Minerva, AMÁLIA, and OpenEuroLLM represent the three operational answers to the European sovereign-LLM question. Each makes different architectural and institutional bets. The strategic discourse benefits from treating all three as data points in the same empirical experiment.
AI research dataset storage
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three standards the movement should adopt.
The structural critique generalizes beyond Minerva. The European sovereign-LLM movement benefits from internalizing these lessons across every subsequent national project. Italy modeled the openness standard; the movement should adopt it as norm.
Minerva is one valid answer to the European sovereign-LLM question. AMÁLIA is another. OpenEuroLLM is potentially a third. The strategic discourse benefits from treating all three as data points in the same empirical experiment rather than as competing national-prestige projects. More analysis like this is needed. Not less.
Implications for European Sovereign AI Strategies
The results from Minerva-3B highlight a fundamental challenge for European countries pursuing sovereign AI models: scaling native-language data and parameters may not suffice to achieve deep, country-specific knowledge. This raises questions about the cost-effectiveness of large-scale native-language training and suggests that European AI initiatives may need to reconsider their investment strategies to meet complex language and knowledge demands.
While Italy’s approach demonstrated impressive institutional coordination and technical achievement, the limited performance on academic benchmarks indicates that the current scale may still be insufficient. This could influence future policy and funding decisions across Europe, emphasizing the importance of aligning model scale with targeted knowledge outcomes.
European Sovereign LLM Development Approaches and Challenges
Italy’s Minerva project represents a contrasting approach to the European sovereign-LLM debate, which includes Portugal’s AMÁLIA model. While AMÁLIA layered Portuguese onto a multilingual foundation with relatively modest native-language data, Minerva was built from scratch, emphasizing native-Italian data and open weights.
Prior to Minerva’s release, European efforts focused on balancing multilingual capabilities with native-language specialization. The structural lesson emerging from Minerva’s performance is that large-scale native-language training alone may not guarantee the desired depth of knowledge, especially at current parameter scales. This challenges the assumption that more native data automatically translates into better country-specific language understanding.
Italy’s investment involved significant institutional coordination, including funding through Italy’s PNRR, and the use of the CINECA supercomputer. Despite these resources, the empirical results underscore the complexity of scaling language models for academic and nuanced language tasks.
Unresolved Questions About Model Scaling and Performance
It remains unclear whether increasing model size or dataset volume beyond current levels will significantly improve Minerva’s academic and complex language task performance. The relationship between native-language data proportion and model depth at various scales is still under investigation, and further iterations of Minerva are ongoing.
Next Steps for Minerva and European Sovereign AI Projects
Researchers plan to continue refining Minerva, including experimenting with larger models and different training methodologies. European policymakers and AI developers will likely reassess investment strategies, considering the empirical limitations revealed by Minerva’s performance. Further benchmarking and comparative studies are expected to clarify the path toward effective country-specific AI models.
Key Questions
Why did Minerva-3B perform poorly on the INVALSI tests?
Despite large-scale native-language training, Minerva-3B’s limited performance suggests that dataset size and model scale are more important for complex tasks than native-language proportion alone.
Does this mean native-language models are not worth the investment?
Not necessarily. The results highlight the challenge of scaling models effectively; strategic investment and larger models may still be needed to achieve country-specific knowledge depth.
How does Minerva compare to other European sovereign models like AMÁLIA?
Minerva was trained from scratch with a larger native-language dataset, yet it underperformed on academic benchmarks compared to expectations, contrasting with AMÁLIA’s approach of layered multilingual training.
What are the implications for European AI policy?
The findings suggest a need to reevaluate funding and development strategies, emphasizing larger scale investments and possibly new methodologies to meet complex language and knowledge demands.
Will further iterations improve Minerva’s performance?
Researchers are continuing to refine Minerva, including increasing scale and experimenting with different training techniques, which may improve future results.
Source: ThorstenMeyerAI.com