ON
← Back to feed
More security in the office: How the Bundesdruckerei project is taming Möve's public sector AI
Germany🏛️ PoliticsCenteryesterday

More security in the office: How the Bundesdruckerei project is taming Möve's public sector AI

The article discusses the Bundesdruckerei's initiative called 'Möve,' which aims to evaluate AI language models specifically tailored for use by German public authorities. It highlights the limitations of international rankings that focus primarily on English comprehension and mathematical skills, emphasizing the need for models that accurately handle the German language, legal reliability, security, and sustainability. The project has evaluated over 50 large language models (LLMs), including GPT, Gemma, Llama, and Mistral, and provides a comparative framework based on performance, security, transparency, and alignment with democratic values. Gemma4 and GPT-40 currently lead the rankings, while Mistral Small 3.1 places within the top ten. The initiative includes test data directly sourced from administrative practices and evaluates seven core criteria, including factual accuracy and adherence to EU values.

The German federal government has launched a comprehensive initiative aimed at evaluating large language models (LLMs) for use in public administration, addressing concerns over reliability, security, and alignment with democratic values. The project, named Möve, short for "Modelle für die öffentliche Verwaltung evaluieren", was initiated by the state-owned Bundesdruckerei, which operates under full federal ownership. The initiative began in 2025 and compares more than 50 LLMs, including commercial offerings such as GPT, Gemma, Llama, Claude, DeepSeek, and Mistral, alongside open-source alternatives. Möve seeks to provide a systematic framework for assessing how well these AI systems perform in areas critical to public administration, such as legal accuracy, linguistic precision, security, sustainability, transparency, and compliance with democratic principles. According to the Bundesdruckerei, international rankings often prioritize factors like general knowledge and mathematical ability, which may not translate effectively into real-world administrative tasks. Instead, the project emphasizes the need for models that can handle complex bureaucratic texts, official documents, and legal terminology accurately in the German context. Frederik Blachetta, CEO of the Bundesdrackerei Group, emphasized that while global benchmarks offer some guidance, they fall short of capturing the nuanced requirements of public sector operations. He noted that the ultimate test lies in whether a model behaves reliably and ethically within Germany’s specific legal and cultural environment. This focus on practical application sets Möve apart from other evaluations, which tend to emphasize theoretical performance metrics rather than real-world utility. To ensure the evaluation reflects actual administrative challenges, the project employs nine specially developed German-language datasets drawn directly from public administration practices. These include real legal texts, official documentation, and publications from federal ministries. Based on this data, researchers have established seven core criteria for assessment, ranging from the ability to summarize lengthy documents accurately to the capacity for multi-document thematic analysis. Additional considerations include adherence to European Union fundamental values and the frequency of factual errors generated by the models. One key component of the evaluation involves analyzing the tendency of AI systems to produce hallucinations, plausible yet factually incorrect information. The project also examines the efficiency and resource consumption of each model, along with the extent to which providers comply with transparency obligations outlined in the EU's AI Act. Results from these assessments are compiled into an interactive online comparison tool, allowing public agencies to filter and compare models based on their specific needs. Collaboration with external experts plays a central role in expanding the scope of the project, particularly in the area of IT security. The Bundesdruckerei works closely with the Federal Office for Information Security (BSI) and the Fraunhofer Institute for Applied and Integrated Security (AISEC). These partnerships aim to address potential risks associated with deploying powerful AI systems, especially given their dual-use nature, capable of both enhancing cybersecurity defenses and being exploited for automated attacks. The ongoing expansion of Möve underscores the growing recognition among German authorities of the need for rigorous oversight of AI technologies in sensitive sectors. By combining technical evaluation with ethical and regulatory scrutiny, the project aims to equip public institutions with the tools necessary to make informed decisions about adopting AI solutions. The findings will continue to evolve as new models emerge and as the landscape of AI governance becomes clearer.

How each side covered it

The same event, grouped by the political lean of the outlets covering it.

How each side covered it

Support independent, bias-aware news and unlock the social pulse, community voting, and your personalized For You feed.

Become a Supporter

Covered around the world

The same event as reported in other countries.

Covered around the world

Support independent, bias-aware news and unlock the social pulse, community voting, and your personalized For You feed.

Become a Supporter

Claims check

Key factual claims, and how many sources assert vs dispute each.

Claims check

Support independent, bias-aware news and unlock the social pulse, community voting, and your personalized For You feed.

Become a Supporter

Go to the primary sources (2)

The official sources this coverage is built on. Read them directly to bypass framing.

1 reports

heise online logoheise onlineIndependentCenterFactual 75Objective 70yesterday
More security in the office: How the Bundesdruckerei project is taming Möve's public sector AI

The article discusses the Bundesdruckerei's initiative called 'Möve,' which aims to evaluate AI language models specifically tailored for use by German public authorities. It highlights the limitations of international rankings that focus primarily on English comprehension and mathematical skills, emphasizing the need for models that accurately handle the German language, legal reliability, security, and sustainability. The project has evaluated over 50 large language models (LLMs), including GPT, Gemma, Llama, and Mistral, and provides a comparative framework based on performance, security, transparency, and alignment with democratic values. Gemma4 and GPT-40 currently lead the rankings, while Mistral Small 3.1 places within the top ten. The initiative includes test data directly sourced from administrative practices and evaluates seven core criteria, including factual accuracy and adherence to EU values.

Bias read (Center): The article presents a balanced overview of the technical and regulatory challenges faced by public administration in adopting AI models. It does not take a clear ideological stance but rather focuses on the practical implications and evaluation criteria set by the Bundesdruckerei. While the subject

Why factuality (75): The article accurately describes MÖVE as a benchmarking framework for evaluating LLMs in the German public sector. It mentions the inclusion of over 50 models and highlights specific models like Gemma4 and GPT-40. However, it incorrectly states the project has been running since 2025 when the primar

Why objectivity (70): The article presents information in a generally neutral manner but shows some bias by highlighting certain models (Gemma4 and GPT-40) as leading performers without providing detailed evidence. The tone is informative but slightly promotional of the Bundesdruckerei initiative.

Keep the news honest.

ObjectiveNews is reader-funded and ad-free — we show you the bias instead of hiding it. Support independent journalism for €5/month.

Become a Supporter

Related stories