Proceedings of the
European Safety and Reliability Conference (ESREL2026)
14 – 19 June 2026, Braga, Portugal

BAAS: Benchmarking Adversarial Agent Strategies - A Comparative Study of Gradient-based, Reinforcement, and Evolutionary Paradigms for Safety-Critical Scenario Generation

Carmen Mei-Ling Frischknecht-Gruber

School of Engineering & Department of Informatics, Zurich University of Applied Sciences & University of Fribourg, Switzerland.

frsh@zhaw.ch

Prof. Dr. Monika Reif

School of Engineering, Zurich University of Applied Sciences, Switzerland.

reif@zhaw.ch

Prof. Dr. Andreas Fischer

Department of Informatics, University of Fribourg, Switzerland.

andreas.fischer@unifr.ch

ABSTRACT

The verification and validation of autonomous driving systems demand systematic exposure to safety-critical interactions that reveal weaknesses in control policies. Traditional verification methods rely on parameter sweeps or fixed test suites, which offer limited coverage of rare but high-risk situations. Recent adversarial approaches increase scenario criticality, i.e., generate more challenging and risky conditions, but often prioritize collision occurrence over behavioral plausibility or diversity. This work investigates the generation of adversarial agents as a reproducible and scalable approach to create diverse yet coherent safety-critical driving interactions. Five paradigms are compared: (1) parameter sweeps as baseline, (2) gradient-based optimization, (3) reinforcementlearning (RL) adversaries, (4) quality-diversity (QD) search via MAP-Elites and QD-RL. A pre-trained driving policy serves as the ego under test in a highway-driving environment. Each method is evaluated on three dimensions: effectiveness, diversity, and behavioral plausibility and feasibility. Additionally, we analyze the impact of singleadversary setups and discuss extensions toward multi-adversary configurations and hybrid curricula combining MAP-Elites seeding with adversarial RL. The study quantifies trade-offs between adaptiveness, diversity and feasibility in the generation of safety-critical scenarios and identifies mechanisms for exposing weaknesses in robust control policies. The framework provides a structured comparison of learning-based, evolutionary and gradientbased paradigms, supporting systematic robustness evaluation of autonomous driving functions and the creation of reusable adversarial scenario databases for future studies.

Keywords: Scenario generation, autonomous systems, testing, reinforcement learning, machine learning.



Download PDF