AI and Observability: Securing the Cloud in the Digital Age
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
The Importance of Observability in Cloud Security
In the age of artificial intelligence, observability has become a central element of cloud security. It enables the unification of visibility, detection, and risk analysis, providing a deep understanding of incidents. Security teams rely on observability data to analyze system behaviors during outages or security incidents. Long associated with performance monitoring, this visibility is now essential for security analyses. To reconstruct the sequence of an attack, it is crucial to understand how cloud identities, services, and resources interact. Observability data, including metrics, events, logs, and traces (MELT), provide this indispensable level of context.
With the rise of applications and services based on language models (LLM) in increasingly distributed environments, the costs and visibility issues related to the fragmentation of tools and infrastructures are growing. Vibe Coding significantly increases the number of application vulnerabilities, while new attacks leveraging AI can cause unpredictable behaviors in certain systems. However, AI can also be a major asset for security teams, accelerating investigations through its ability to correlate vast volumes of signals. Thus, observability data becomes an essential component of the security framework, at the intersection of application performance and security.
Leveraging Observability Data for Security
To understand cloud incidents, unified and centralized visibility is essential. However, security signals alone can often prove insufficient. Each incident generates a multitude of streams associated with authentications, applications, and infrastructure data, capturing each contextual element at different moments in time. Without the ability to correlate security and observability contexts, teams find themselves piecing together the puzzle afterward.
The Cloudflare incident at the end of 2023 concretely illustrates the importance of cross-referencing operational data and security signals during an incident. By exploiting compromised credentials and an access token through a third-party provider, attackers managed to penetrate internal systems, including the self-hosted Atlassian environment. Faced with a series of weak signals, such as unusual authentication flows, reconnaissance activities, and attempts to access other systems, no single element allowed for an understanding of the attack. It was the correlation of observability data that ultimately enabled teams to reconstruct the sequence of events and establish an accurate timeline.
This example shows that teams can rely on their existing data to answer key questions in a security investigation: what has changed in the system? Who or what is responsible for this change? What other elements have been impacted? Which endpoints have been called? In the case of services integrating AI, an additional question arises: which prompt or tool call triggered the action?
This need for shared context drives organizations to bring their SRE and security teams closer together, or even merge them. By combining a deep understanding of architecture with security expertise, they make their systems more resilient to failures. Cloud security is then no longer seen as a separate layer but as a natural extension of the observability of environments. Security signals thus become more faithful to the actual behavior of systems, making them more relevant and actionable in the event of an incident.
AI and Threat Analysis
When threat analysis relies on existing observability data, teams can reuse the same context as performance monitoring to link a system's behavior to an attack path or an exploited vulnerability. This data foundation also allows AI to become more relevant, facilitating the generation of investigation summaries and remediation recommendations.
We see immediate value in linking observability data to security signals through AI-based analysis, particularly to generate clear and actionable event mappings during incidents. This contribution is especially visible during the triage and investigation phases, which are often lengthy and iterative, where teams must verify the relevance and priority of various signals.
For example, imagine a SIEM signal indicating the addition of an AdministratorAccess policy to a service account, associated with a source IP address identified as a suspicious residential proxy. In a traditional investigation, the analysis would generally follow several steps: verifying whether the IP address corresponds to a legitimate administrator and if the session aligns with their login habits; reconstructing events that occurred at the same time, such as policy changes, access key creations, authentication failures, or unusual API calls; identifying impacted services and resources to assess the extent of possible access; and analyzing associated network behaviors, including unusual locations and spikes in outgoing traffic.
AI-assisted analysis, leveraging observability data, can condense this complex process into a single assessment. Teams can thus reduce investigation time from several hours to a few minutes, allowing them to focus on corrective actions, such as disabling compromised credentials or strengthening least privilege principles.
Prioritizing Risks Through Observability
Teams must be able to rely on observability data at every stage of the software development lifecycle (SDLC). With the acceleration of code production driven by AI-assisted development workflows, this data becomes essential for effectively prioritizing risks. In practice, only a fraction of critical vulnerabilities warrants immediate attention. By linking code analysis results to production behaviors and impacted services, it becomes possible to identify genuinely significant risks earlier. When leveraging this data, AI-based approaches also help reduce false positives and avoid unnecessary remediation efforts.
Integrating this data into code reviews, through LLM-based analysis applied to pull requests (PR), allows for the identification of risks before production deployment. AI proves particularly useful for analyzing large PRs, where malicious code can hide among seemingly innocuous changes.
To distinguish malicious code from legitimate modifications, models need context on actual attacks and common development practices. This is why datasets must be continuously enriched, both with observed or simulated attacks and with standard PRs, to reduce false positives.
Towards Resilient Cloud Security
Cloud environments will continue to evolve with the adoption of new technologies, while the integration of AI adds an additional layer of complexity. With each new layer, new blind spots emerge, and the attack surface expands. In this context, to interpret, investigate, and correct accurately, observability data must be considered the foundation of cloud security, especially in AI-dependent environments. By adopting this approach, organizations can better anticipate threats, enhance the resilience of their systems, and ensure optimal security.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.