Proceedings of the
European Safety and Reliability Conference (ESREL2026)
14 – 19 June 2026, Braga, Portugal
Knowledge and LLM-Based Approach to Develop Data-Driven Causal Model for Extending Human Reliability Analysis
Department of Computer Sciences, UCLA, USA.
The B. John Garrick Institute for the Risk Sciences, UCLA, USA.
ABSTRACT
Human performance critically determines safety in high risk domains where organizational failures underlie major accidents. While numerous organizational factors (OFs) influencing human reliability have been identified, their causal interrelations remain poorly understood. A key challenge is that Human Reliability Analysis (HRA) data exists primarily as unstructured text (accident reports, interviews, and regulatory documents) which conventional data mining techniques cannot effectively analyze for implicit, context dependent causal relationships. This study employs a Large Language Model (LLM) based approach to extract causal relationships among OFs from 266 validated academic and technical documents. Following comparative evaluation of three LLMs (GPT5, GPT-5-mini, Grok-4), GPT-5-mini was selected based on cost-effectiveness and identification performance. Through structured prompt engineering defining 19 organizational factors and 38 target causal relationships, the system extracted 1,538 causal relationship instances with supporting textual evidence and justifications. Results provide empirical frequency distributions quantifying the prevalence of specific OF interactions. This framework demonstrates a scalable, reproducible method for transforming unstructured linguistic data into analyzable causal knowledge, enhancing the data-driven foundation of HRA. Future work will implement causality-specific prompting and LLM-as-a-judge validation to support development of probabilistic Bayesian network models for quantitative risk assessment.
Keywords: Human Reliability Analysis, Organizational Factors, Large Language Models, Causal Relationships, Information Extraction, Prompt Engineering, Natural Language Processing, Bayesian Networks, High-Risk Industries, Safety Analysis.

