Artificial Intelligence–Based Incident Detection, Response, and Recovery in Cloud-Native Computing Environments: An Intelligent Framework for Autonomous Cyber Resilience
DOI:
https://doi.org/10.64235/xd3wma19Keywords:
Artificial intelligence; AIOps; cloud-native computing; Kubernetes; autonomous cyber resilience; incident response; root-cause analysis.Abstract
Cloud-native computing has transformed the deployment and operation of digital services through microservices, containers, Kubernetes orchestration, distributed applications, and dynamically scalable infrastructure. Although these technologies improve flexibility, availability, and deployment speed, they also increase operational complexity. Failures can emerge across interconnected services, containers, networks, databases, configurations, and security controls, making incident detection, diagnosis, response, and recovery more difficult. Conventional monitoring approaches often rely on static thresholds, fragmented dashboards, and manually coordinated response processes. These approaches can generate excessive alerts, provide limited contextual understanding of incidents, and delay the restoration of affected services.
This study proposes an artificial intelligence-based framework for autonomous cyber resilience in cloud-native computing environments. The framework integrates telemetry collection and analysis, anomaly detection, incident classification, distributed root-cause analysis, risk-aware response orchestration, recovery validation, and continuous learning. It combines operational data from logs, metrics, traces, Kubernetes events, and security alerts to identify abnormal behaviour and determine probable causes across distributed microservice environments. The response layer applies predefined policies and confidence thresholds to recommend or execute remediation actions, including pod restart, workload scaling, traffic rerouting, deployment rollback, service isolation, and security escalation. High-risk actions remain subject to human approval and governance controls.
The study adopts a design-science and experimental evaluation methodology. The proposed framework will be assessed in a Kubernetes-based cloud-native testbed using simulated operational and cybersecurity incident scenarios, including resource exhaustion, service failures, network latency, configuration errors, database disruptions, and suspicious access events. Performance will be evaluated using detection accuracy, precision, recall, F1-score, root-cause localisation accuracy, mean time to detect, mean time to respond, mean time to recover, recovery success rate, and service availability. The study contributes an integrated and policy-governed approach to intelligent incident management that aims to improve detection speed, reduce recovery time, strengthen service continuity, and support safer automated remediation in complex cloud-native environments.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

