Abstract
AI systems in safety-critical settings such as healthcare and robotics must not only perform well on average, but also detect unsafe behaviour and recover at run time. We present SelfHeal, a lightweight deployment-time assurance framework that wraps existing models with a two-layer runtime controller: (i) a safety shield (R0) that blocks unsafe actions before they reach the environment, and (ii) a recovery layer (R1) that searches for safe alternatives when blocking alone would cause deadlock or loss of service. Explainable AI analyses are used as calibration signals to quantify weak sensitivity to safety-critical inputs and to set safety rules, margins, and risk triggers. We instantiate SelfHeal in two domains. In a Synthea-based treatment recommendation task, rule-based shielding removes all injected drug-allergy conflicts, reducing the unsafe recommendation rate from 14.8% to 0%. With recovery enabled, the system restores performance to 95.1% accuracy (compared with 95.96% for the baseline) with only a modest increase in decision time. In the Safety-Gymnasium SafetyPointGoal1-v0 benchmark, shielding reduces constraint violations from 186 to 45 cost events per 1,000 steps and improves success but can introduce navigation stalls. Adding recovery further reduces violations to 11 per 1,000 steps, increases success from 63.8% to 92.6%, and cuts the stall rate by more than half. Across both domains, the combination of shielding, recovery, and incident logging reduces unsafe behaviour while keeping performance close to the baseline, supporting practical self-healing operation in safety-critical AI systems.
| Original language | English |
|---|---|
| Publisher | TechRxiv |
| Number of pages | 21 |
| DOIs | |
| Publication status | Published - 29 Dec 2025 |
Fingerprint
Dive into the research topics of 'SelfHeal: A Self-Healing AI Control Framework for Autonomous Recovery from Unsafe Behaviour'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver