Proceedings of the
European Safety and Reliability Conference (ESREL2026)
14 – 19 June 2026, Braga, Portugal
AI Agent for Assessing the Quality of FMEA Reports
Laboratoire de Génie Industriel, CentraleSupélec, Université Paris-Saclay, France.
Laboratoire de Génie Industriel, CentraleSupélec, Université Paris-Saclay, France.
Laboratoire de Génie Industriel, CentraleSupélec, Université Paris-Saclay, France.
ABSTRACT
Anticipating and managing risks in the early development of industrial systems is essential to reduce costs and hazards through design improvements. Widely accepted methods such as Failure Modes and Effects Analysis (FMEA) were designed precisely to support risk prevention during system design. However, conducting an FMEA is costly, and because FMEAs are qualitative analyses that rely on natural language, using Large Language Models (LLMs) to automate parts of this work is increasingly attractive. Yet, ensuring the quality of such AI-generated analyses remains challenging. Currently, FMEAs are mainly crafted and assessed by human experts, while too few automated methods exist to evaluate their quality. Meanwhile, the "LLM-as-a-judge" paradigm has emerged, where LLM agents evaluate task outputs. In this work, we: (1) propose a method to automatically generate FMEA reports, (2) introduce a LLM-as-a-judge framework to automatically produce review reports that assess FMEA quality, and (3) compare AI-generated and human-generated FMEAs using both our LLM-as-a-judge framework and human evaluations.
Keywords: Failure Modes and Effects Analysis (FMEA), Large Language Model (LLM), LLM-as-a-judge, Evaluation, FMEA generation.

