
Security operations are entering an important transition. For years, the SIEM has been the center of the SOC. It collects telemetry, correlates events, generates alerts, and gives analysts a place to investigate. But the assumptions behind that model are breaking down.
The problem is not simply that organizations are generating more logs. They are generating orders of magnitude more data, from more sources, with more context, and needing to retain it for longer to find and mitigate the root cause of incidents before damage is done. Cloud infrastructure, identity systems, endpoints, applications, APIs, network telemetry, DNS, SaaS platforms, and increasingly AI systems all contribute to the ever-expanding security data footprint. At large enterprises, that can mean tens of terabytes of telemetry every day and ultimately petabytes of historical data.
Traditional SIEM architectures were not designed for this reality.
For years, the answer to rising data volumes has been predictable. Companies ingest less, retain less, or move older data somewhere cheaper, such as cold or frozen storage, where it can take hours or days to retrieve. That approach may make the economics of the SIEM manageable, but it creates a security problem because you cannot investigate an incident with only partial evidence. Not to mention, if it takes hours or days to retrieve that data, investigations are stalled, and the damage may have already happened.
And now the industry is adding AI agents, a new consumer of security data. While this may seem like another enhancement of the automation capabilities offered by a SIEM, there is real urgency behind this shift. Attackers are already using AI to move faster, automating all aspects of attacks and compressing the time between initial compromise and impact from days to hours. Defenders who investigate at human speed against adversaries operating at machine speed are fighting a losing race. Agentic security operations are not a luxury; they are the only way to keep pace.
Agentic security operations promise to change how SOCs work. Instead of asking an analyst to manually pivot among dozens of tools, an investigation agent can take an alert, gather relevant evidence, correlate activity across identities, endpoints, cloud infrastructure and applications, and develop a hypothesis about what happened.
That is potentially transformative, but for it to work, the agent needs access to all of the evidence.
For example, an agent investigating a suspicious login might need to determine whether the credentials were compromised, where they were subsequently used, whether the user accessed an unusual application, whether a privilege changed, whether an endpoint exhibited related activity, and whether the same identity or infrastructure was involved in earlier events.
The relevant evidence may span weeks or months and cross many different data sources.
If the SIEM contains only a fraction of that telemetry or if older data has been moved into a tier that requires a restore before it can be searched, the agent cannot reconstruct the incident. It can only reason about the subset of reality that happens to be available to it.
That creates a distinction between AI that can reason and AI that can investigate.
An agent may be extraordinarily capable at reasoning over evidence. But if the evidence is missing, fragmented, or too slow to retrieve, better reasoning does not solve the problem.
This becomes even more important when we consider how agents perform investigations.
A good investigation is iterative. An agent starts with an alert, asks a question, examines the answer, develops another hypothesis, and queries the data again. Each answer determines the next question.
Suppose an agent sees an anomalous authentication event. It might ask:
- Has this identity behaved this way before?
- What device was involved?
- What other resources did the identity access?
- Did the account’s privileges change?
- What network connections followed?
- Did another identity exhibit the same behavior?
- When did this activity first begin?
The agent cannot know all of those questions in advance. Investigation is a process of discovery.
That means the underlying data layer has to support fast, exploratory queries across enormous historical datasets. If every investigation requires an analyst or an agent to identify a time window, request that data be restored, wait for it to become searchable, and then begin the next query, the investigation becomes serial and slow.
This is especially problematic for attacks whose root cause is separated from the initial alert by days, weeks, or months. Credential compromise, lateral movement, and long-running persistence rarely respect the retention window of a SIEM.
The irony is that we’re building agents specifically to investigate incidents at machine speed while giving those agents a data architecture that may require human-speed retrieval.
There is a temptation to think that the answer is simply to put more AI into the SOC.
But adding an investigation agent on top of an incomplete data set does not create visibility. It creates an intelligent system operating with incomplete evidence. That can be worse than a slow investigation because an agent can produce a confident conclusion from an incomplete picture.
The issue is not that agents need to ingest every byte of telemetry directly into an LLM. They don’t. Sending petabytes of raw logs into a model would be economically and technically absurd. To get agents the speed and full-fidelity access to historical data they need, some companies, especially MSSPs, are looking into adding a high-performance security data layer to their stack. Sitting beneath the SIEM, the data layer is a place where complete, high-fidelity telemetry can be retained for years, economically, and queried in seconds when an investigation requires it.
The architecture separates the problems of storing and accessing security data from the problems of detecting threats and orchestrating response. That means organizations can continue using their existing SIEM for the functions it does well—detection, alerting, correlation and analyst workflows—while maintaining a broader historical security record that remains immediately searchable.
The old assumption was that keeping everything hot was inherently too expensive. At petabyte scale, that assumption no longer has to hold. Modern data architectures can use highly efficient compression, indexing, and columnar storage to retain enormous quantities of telemetry while keeping it available for rapid search. The objective is not simply to store more data. It is to make the data operationally useful.
A petabyte sitting in an archive is technically retained, but it is not necessarily available to an investigator at the moment of need. A petabyte that can be searched immediately becomes part of the organization’s active security context.
For agentic SOCs, that difference is critical. The most valuable security data is not necessarily the data associated with the alert that just fired. It may be the event that happened 45 days earlier and explains how the attacker got there.
Reasoning is a model problem. Investigation is a data problem.
The next generation of security operations will not be defined simply by whether an organization has a SIEM or whether it has deployed AI agents. It will be defined by whether its architecture gives those agents and its human analysts access to the data they need, when they need it.
Agentic security operations change the economics of investigation. An analyst might investigate one incident at a time. An agent can potentially investigate thousands of questions continuously, including delivering on the long-unfulfilled promise of proactive threat hunting — constantly sweeping recent and historical telemetry for signs of compromise that no alert ever fired on. That creates enormous leverage, but it also multiplies the consequences of a weak data foundation. If an agent has to wait for data, its advantage disappears. If an agent cannot search historical telemetry, its investigation has blind spots. If high-value sources were never ingested because they were too expensive to put into the SIEM, the agent cannot reason over them. And if the data is scattered across disconnected systems with different schemas and retention policies, the agent spends its time finding and preparing evidence rather than finding the root cause.
The answer is not to abandon the SIEM. It is to stop asking the SIEM to be everything.
Attackers have already made their architectural decisions. The question is whether defenders’ data can keep up with their agents.
______
About the Author:
Matteo Rebeschini is a Field CISO at Hydrolix, where he leads the company’s expansion into the cybersecurity and security analytics market. With deep expertise spanning SIEM, threat detection, large-scale log management, and cloud-native security architectures, Matteo helps enterprise organizations design and deploy modern security data platforms built for the realities of petabyte-scale environments.
Join our LinkedIn group Information Security Community!











