Proceedings of the
European Safety and Reliability Conference (ESREL2026)
14 – 19 June 2026, Braga, Portugal

Overcoming the ``Glass Ceiling'' of LLMS for Automated Risk Analysis in Occupational Safety

Gregory Forte

PREVENTEO, France.

gregory.forte@preventeo.com

ABSTRACT

Large language models (LLMs) are increasingly explored for automating expert reasoning tasks, including risk analysis in occupational health and safety (OHS). This study evaluates their reliability and reproducibility in this safety-critical context. Preventeo conducted 600 systematic tests across 25 representative scenarios, comparing the outputs of GPT-4o and Gemini Pro with expert-validated reference grids assessing hazard identification and risk evaluation accuracy. Results show that both models achieve high precision (  ∼ 0.95 ), confirming their ability to produce relevant and contextually coherent assessments. However, their recall remains limited ( ∼ 0.75), leading to an F1-score below industrial acceptance thresholds. Moreover, significant reproducibility issues were observed, arising from the stochastic nature of LLMs and unsignaled updates of proprietary models. These factors create a "glass ceiling" that prevents direct, autonomous deployment of LLMs in domains where reliability, traceability, and compliance are essential. This work provides two main contributions. To our knowledge, this work provides one of the first large-scale quantitative evaluations of LLM-based OHS risk-grid generation against expert-validated reference grids, highlighting the gap between linguistic fluency and domain reliability in safety-critical contexts. Second, it introduces a reproducible evaluation framework combining quantitative metrics (precision, recall, F1score) and expert validation. Based on these insights, the study advocates for hybrid modular architectures integrating LLMs with knowledge graphs, retrieval-augmented generation (RAG), and confidence scoring to improve robustness, explainability, and compliance with industrial standards.

Keywords: AI, LLM, Risk assessment.



Download PDF