Veća sigurnost u uredu: Kako projekt "Federal Printing House" obuzdava umjetnu inteligenciju u državnim službama
U članku se raspravlja o inicijativi Bundesdruckerei pod nazivom "Möve", čiji je cilj ocjenjivanje modela jezika umjetne inteligencije posebno prilagođenih za uporabu njemačkim javnim tijelima. Naglašava ograničenja međunarodnih rangiranja koja se prvenstveno fokusiraju na razumijevanje engleskog jezika i matematičke vještine, naglašavajući potrebu za modelima koji točno obrađuju njemački jezik, pravnu pouzdanost, sigurnost i održivost. Projekt je procijenio više od 50 velikih jezičnih modela (LLM), uključujući GPT, Gemma, Llama i Mistral, te pruža usporedni okvir temeljen na performansama, sigurnosti, transparentnosti i usklađenosti s demokratskim vrijednostima. Gemma4 i GPT-40 trenutno vode u rangiranju, dok se Mistral Small 3.1 nalazi među prvih deset. Inicijativa testiranja uključuje podatke izravno iz administrativnih praksi i ocjenjuje sedam ključnih kriterija, uključujući činjeničnu točnost i pridržavanje vrijednostima EU-a.
The German federal government has launched a comprehensive initiative aimed at evaluating large language models (LLMs) for use in public administration, addressing concerns over reliability, security, and alignment with democratic values. The project, named Möve, short for "Modelle für die öffentliche Verwaltung evaluieren", was initiated by the state-owned Bundesdruckerei, which operates under full federal ownership. The initiative began in 2025 and compares more than 50 LLMs, including commercial offerings such as GPT, Gemma, Llama, Claude, DeepSeek, and Mistral, alongside open-source alternatives. Möve seeks to provide a systematic framework for assessing how well these AI systems perform in areas critical to public administration, such as legal accuracy, linguistic precision, security, sustainability, transparency, and compliance with democratic principles. According to the Bundesdruckerei, international rankings often prioritize factors like general knowledge and mathematical ability, which may not translate effectively into real-world administrative tasks. Instead, the project emphasizes the need for models that can handle complex bureaucratic texts, official documents, and legal terminology accurately in the German context. Frederik Blachetta, CEO of the Bundesdrackerei Group, emphasized that while global benchmarks offer some guidance, they fall short of capturing the nuanced requirements of public sector operations. He noted that the ultimate test lies in whether a model behaves reliably and ethically within Germany’s specific legal and cultural environment. This focus on practical application sets Möve apart from other evaluations, which tend to emphasize theoretical performance metrics rather than real-world utility. To ensure the evaluation reflects actual administrative challenges, the project employs nine specially developed German-language datasets drawn directly from public administration practices. These include real legal texts, official documentation, and publications from federal ministries. Based on this data, researchers have established seven core criteria for assessment, ranging from the ability to summarize lengthy documents accurately to the capacity for multi-document thematic analysis. Additional considerations include adherence to European Union fundamental values and the frequency of factual errors generated by the models. One key component of the evaluation involves analyzing the tendency of AI systems to produce hallucinations, plausible yet factually incorrect information. The project also examines the efficiency and resource consumption of each model, along with the extent to which providers comply with transparency obligations outlined in the EU's AI Act. Results from these assessments are compiled into an interactive online comparison tool, allowing public agencies to filter and compare models based on their specific needs. Collaboration with external experts plays a central role in expanding the scope of the project, particularly in the area of IT security. The Bundesdruckerei works closely with the Federal Office for Information Security (BSI) and the Fraunhofer Institute for Applied and Integrated Security (AISEC). These partnerships aim to address potential risks associated with deploying powerful AI systems, especially given their dual-use nature, capable of both enhancing cybersecurity defenses and being exploited for automated attacks. The ongoing expansion of Möve underscores the growing recognition among German authorities of the need for rigorous oversight of AI technologies in sensitive sectors. By combining technical evaluation with ethical and regulatory scrutiny, the project aims to equip public institutions with the tools necessary to make informed decisions about adopting AI solutions. The findings will continue to evolve as new models emerge and as the landscape of AI governance becomes clearer.
Kako je izvijestila svaka strana
Isti događaj, grupiran prema političkom nagibu medija koji su o njemu izvještavali.
progresivno
sredina
konzervativno
★
Kako je izvijestila svaka strana
Podržite neovisne vijesti svjesne pristranosti i otključajte društveni puls, glasovanje zajednice i svoj personalizirani feed Za tebe.
U članku se raspravlja o inicijativi Bundesdruckerei pod nazivom "Möve", čiji je cilj ocjenjivanje modela jezika umjetne inteligencije posebno prilagođenih za uporabu njemačkim javnim tijelima. Naglašava ograničenja međunarodnih rangiranja koja se prvenstveno fokusiraju na razumijevanje engleskog jezika i matematičke vještine, naglašavajući potrebu za modelima koji točno obrađuju njemački jezik, pravnu pouzdanost, sigurnost i održivost. Projekt je procijenio više od 50 velikih jezičnih modela (LLM), uključujući GPT, Gemma, Llama i Mistral, te pruža usporedni okvir temeljen na performansama, sigurnosti, transparentnosti i usklađenosti s demokratskim vrijednostima. Gemma4 i GPT-40 trenutno vode u rangiranju, dok se Mistral Small 3.1 nalazi među prvih deset. Inicijativa testiranja uključuje podatke izravno iz administrativnih praksi i ocjenjuje sedam ključnih kriterija, uključujući činjeničnu točnost i pridržavanje vrijednostima EU-a.
Procjena pristranosti (Sredina): U članku se iznosi uravnotežen pregled tehničkih i regulatornih izazova s kojima se javna uprava suočava pri usvajanju modela umjetne inteligencije.
Zašto činjenice (75): The article accurately describes MÖVE as a benchmarking framework for evaluating LLMs in the German public sector. It mentions the inclusion of over 50 models and highlights specific models like Gemma4 and GPT-40. However, it incorrectly states the project has been running since 2025 when the primar
Zašto objektivnost (70): The article presents information in a generally neutral manner but shows some bias by highlighting certain models (Gemma4 and GPT-40) as leading performers without providing detailed evidence. The tone is informative but slightly promotional of the Bundesdruckerei initiative.
★
Neka vijesti ostanu poštene.
ObjectiveNews financiraju čitatelji i bez oglasa je – pristranost vam pokazujemo, ne skrivamo. Podržite neovisno novinarstvo za 5 €/mjesec.