While the advent of large language models (LLMs) has marked a transformative phase in AI, existing models often fall short of the specific needs of the public sector and other users in Europe.
Proprietary and 3rd-party LLMs offer powerful capabilities, but have limitations around language diversity, particularly for low-resource languages (those with fewer speakers and less linguistic data available).
The exact data sources are rarely reported in detail, but a widely used resource is Common Crawl, where most EU languages are highly underrepresented. For example, Latvian accounts for only 0.09% of the total dataset, Irish 0.07% and Maltese 0.03%. The least-represented half of all EU official languages add up to merely 2.4%. Other challenges include data quality, copyright safety, transparency and freedom from bias.
Contributing to the European ecosystem of LLMs
The European Commission is working to address these limitations by using the high-quality multilingual data generated by the EU institutions to contribute to the European ecosystem of LLMs – which are better suited to the EU’s multilingual landscape.
This work is part of DG Translation’s partnership with the Directorate-General for Communications Networks, Content and Technology (DG CONNECT) for AI-based multilingual services under the Digital Europe programme.
AI for a multilingual Europe
Models that cover only a limited number of languages and underperform on low-resource languages are major obstacles for multilingual organisations and societies. This is particularly relevant for a multilingual Europe, especially when it comes to European AI projects that require a broad range of EU languages.
The EU institutional LLM – enhanced with formal texts from the EU institutions – is a powerful addition to European sovereign AI, complementing other efforts and contributing to a diverse landscape of European solutions. Its enhanced EU language capabilities and EU knowledge are better tailored for use by EU public administrations, small businesses, academia and non-governmental organisations.
The project is a stepping stone in the EU’s ambitions to become a major player in AI innovation and strategic technologies, while being able to rely on its own digital systems and tools.
Creating a high-quality EU institutional LLM
To realise this vision, DG Translation’s expert engineers are building on existing European open-source LLMs, with the goal of improving their multilingual capabilities. This work draws on 2 key European assets:
- the supercomputers provided by the European High Performance Computing Joint Undertaking (EuroHPC JU)
- the datasets of the European Advanced Multilingual Information System (Euramis), a unique and voluminous corpus of multilingual text from all the EU institutions.
Better coverage of EU languages
In addition, DG Translation’s proximity to language professionals and their direct feedback gives us a major advantage when it comes to preparing data and evaluating models. The resulting models demonstrate better coverage of all EU official languages and an improved ability to handle EU topics.
Early results confirm this potential. In a benchmark based on EU institutional texts, the EU institutional LLM consistently outperformed the original model across all EU languages tested. The gains were most significant for languages typically underrepresented in global AI training data:
- Irish nearly quadrupled its score
- Estonian improved by almost 80%
- Greek nearly doubled its result
- Latvian and Lithuanian also recorded gains of around 70 to 75%.
In various stages of the EU institutional LLM project, the European Commission has been able to access supercomputing infrastructure through 3 projects with the EuroHPC JU. First, the goal was to develop a cutting-edge skillset and demonstrate the ability to train large AI models with the MeluXina supercomputer in Luxembourg. Then followed more advanced and intense training of LLMs with the Leonardo supercomputer in Bologna. Currently, we are testing our algorithm's efficiency on the MareNostrum 5 supercomputer in Barcelona to optimise future training.
The EU institutional LLM has been built with European technology at its core, which is why an existing European open LLM was selected for the project. Models created by Mistral AI (Mixtral 8x7B and 8x22B) are enhanced with our Euramis language data.
The EU institutional LLM in practice
The EU institutional LLM already powers eSummary, one of our AI-based multilingual tools. Future iterations will power more of our tools.
The model is available for download from the European Language Data Space by any EU-based legal entity:
- EU institutional LLM v1 base – continually pretrained multilingual base model
- EU institutional LLM v1 instruct – instruction-tuned version of the base model
What are the main features of DG Translation’s project?
- Comprehensive inclusion of all 24 EU official languages
Each low-resource language is represented with at least 1 billion tokens (units of text). - Bottom-up approach
The Euramis data used is meticulously curated in line with stringent quality standards. It is aligned with core European values and devoid of any copyright infringements. - Pre and post-training
The project focuses on continuing the basic training of a state-of-the-art LLM with EU institutional data, fine-tuning it for use cases in EU-public administrations and aligning it to EU values and EU-specific preferences.
The current model is the first step in a longer-term programme with future iterations planned.
This page was last updated on 16 July 2026
