HEGEMON: Holistic Evaluation of Generative Foundation Models in the Security Context

Funding: Agentur für Innovation in der Cybersicherheit GmbH (Cyberagentur) through pre-commercial procurement
Project duration: 36 months

The HEGEMON project investigates how generative foundation models can be evaluated and adapted for safety-critical applications. Its goal is to develop holistic, domain-specific benchmarks for assessing whether these models are not only technically capable, but also robust, trustworthy, efficient, and strategically suitable for use in internal and external security.

Within HEGEMON, the UniBw AIML group leads the research and conceptual development of the overarching benchmarking framework. The group will analyse existing evaluation approaches, identify research gaps, and define evaluation criteria and metrics across five main dimensions: a) technical performance and adaptability, b) security and robustness against adversarial manipulation, c) trustworthiness, including explainability, transparency, fairness, usability, and compliance, d) monetary and non-monetary costs, such as inference time, memory, and energy consumption, and e) strategic considerations, including technological sovereignty and resilience.

A central contribution of the UniBw AIML group is the operationalisation of this framework for safety-critical geoinformation applications. This includes defining quantitative and qualitative metrics, establishing suitable baselines and normalised scoring methods, and developing procedures for criteria that require expert or end-user assessment. The group will also contribute to translating the general framework into use-case-specific benchmarks for natural-language geospatial dossiers, vector-data generation and cartographic generalisation, and multimodal question answering based on maps and remote-sensing data.

Through this work, the UniBw AIML group aims to create a reproducible and application-oriented foundation for comparing generative AI models in security-relevant settings. The results will include a general benchmarking concept, use-case-specific evaluation methods, scientific publications, and practical guidance for assessing the safe and responsible deployment of foundation models.

Official HEGEMON programme information