Upcoming talk: Safety First! Modelling requirements from the GPT-5 system card (MoDRE’26)
Safety First! Modelling Requirements from GPT-5 System Card using Lightweight Safety Models will be presented on August 17th, 2026 at the 16th International Model-Driven Requirements Engineering workshop (MoDRE), co-located with the 34th IEEE International Requirements Engineering conference (RE 2026) in Montreal.
Where the ICSE-NIER paper argued the principle, this one runs the experiment at full size. The paper takes the GPT-5 model and system card, pulls out 74 claims, and traces them back to 22 requirements nobody ever wrote down. The requirements that turn out to be missing, and the claims left with no evidence attached, are the actual result: current documentation practice hides both.
The paper is part of the FATES-MLOps project, funded by the ANR in France and NSERC in Canada, and it comes out of the long-running collaboration between McSCert and the French teams in Toulouse and Nice that the project is built on. Two of its four authors, Mireille Blay-Fornarino and Nicolas Lacroix, sit on the Nice side of that consortium.
Abstract
Requirements engineering for Machine Learning (ML) models is gaining renewed importance as organizations increasingly publish model cards and system cards to document the behaviour, capabilities, and limitations of their models. These artifacts have become a de facto standard for communicating model properties to downstream developers and auditors. However, the claims made in such cards are rarely linked to explicit requirements, and they are often expressed in vague or imprecise language, particularly with respect to safety. This makes it difficult to determine which requirements a card actually addresses, or whether the supporting evidence is adequate. We propose an approach based on argumentation modelling, drawing on Toulmin’s model in the form of justification models, to systematically capture the claims made in model and system cards and trace them back to upfront requirements. The resulting diagrams make each claim’s grounds and warrants explicit and establish traceability links between documented assertions and the requirements they are intended to satisfy. We evaluate the approach on the model and system card for GPT-5, identifying 74 claims and tracing them to 22 implicit requirements. The analysis surfaces unspecified requirements and gaps in evidence, showing that argumentation-based traceability can expose weaknesses obscured by current documentation practices.
Reference
Kalvin Thuan-Phong Khuu, Nicolas Lacroix, Mireille Blay-Fornarino, and Sébastien Mosser. Safety First! Modelling Requirements from GPT-5 System Card using Lightweight Safety Models. In 16th International Model-Driven Requirements Engineering (MoDRE), co-located with IEEE RE 2026. Montreal, Canada. August 2026. [HAL]