Aqua Blog

Supply Chain Security Has the Operating Model Backwards

Supply Chain Security Has the Operating Model Backwards

TL;DR Supply chain security usually treats runtime as the final stage, the place where teams find out whether earlier controls worked. Unfortunately, this operating model is backwards. This is not an argument to scan less or to wait until production to secure software. Runtime shows what is actually running, which code is reachable and how software behaves under production conditions, and that evidence should decide what gets secured first upstream. Runtime can also stop malicious behavior at the moment it executes, and each of those catches can strengthen the controls that come before it.

Software supply chain security has spent years getting better at finding risk before software reaches production. We scan packages, dependencies and images. We generate software bills of materials (SBOMs). We evaluate provenance and try to stop malicious artifacts before they execute.

That work is essential. But finding something and knowing how much it matters are different security decisions.

Runtime provides the context for the second decision. It tells us what is actually running, which vulnerable code is reachable and what software is doing under production conditions. That intelligence should not sit at the end of the lifecycle waiting to catch what earlier controls missed. It should flow back through the lifecycle and determine what security teams prioritize upstream.

What did the Graphalgo research find?

On September 22, Aikido published research on the Graphalgo campaign after identifying Go malware in at least two Terraform providers and at least two Go modules. Aikido said this was the first time its researchers had observed malware distributed through Terraform providers. The campaign was linked to earlier Graphalgo activity in npm through shared infrastructure and cryptographic material.

The activation mechanism matters more than the package names. In the Terraform providers, hidden code ran only when the SHA-256 hash of two Terraform variables, containerName and networkID, matched a hardcoded value. Until then, the provider behaved normally. Once the condition was met, it decrypted embedded content and launched it with a detached go run . command. One of the Go modules used a similar pattern, activating only when it processed an object with a specific price value.

The malicious logic was present in the artifact from the start. The conditions that activated it were execution conditions.

Why should runtime be where prioritization starts?

Most supply chain security workflows move in one direction. They inspect source code, dependencies, build systems, images and deployment artifacts. Findings move into remediation workflows. Production sits near the end, where security teams discover whether earlier controls were sufficient.

That architecture leaves some of the best prioritization evidence at the wrong end of the process.

Consider two vulnerabilities with the same severity. One exists in an image, but its vulnerable code is never loaded by the running application. The other sits on a reachable code path in a production workload. A finding-first model puts both into the same queue. Runtime context changes the decision because it shows which finding represents real production exposure. The gap is not small. A widely cited 2023 industry report found that 87% of container images running in production had high or critical vulnerabilities, yet only 15% of the high and critical vulnerabilities with an available fix were in packages actually loaded at runtime. We are wasting 85% of our time.

Scanning tells teams what could matter. Runtime context tells them what matters in the applications they are actually defending. The point of connecting the two is not to produce another finding. It is to make the existing findings more useful.

Would runtime-first prioritization have caught Graphalgo?

Would runtime prioritization have caught Graphalgo? Not until it ran. Code that waits for a specific hash or a specific data value sits on a path that normal operation never reaches. A reachability-based model would see that path as unexercised and rank it low. By its own logic, it would be right to. Dormant malicious code is designed to look like theoretical risk until the moment it stops being theoretical.

If prioritization were the only thing runtime did, that would be a real gap. It is not the only thing runtime does. Two other functions matter here, and they explain why runtime belongs at the center of the program rather than at the end.

Runtime enforcement acts at the moment of execution. When the trigger fires, the software starts doing things it has never done before. In the Terraform variant, it decrypts hidden content and launches new code that was not part of the artifact as it was built. Those are behaviors, and behaviors can be evaluated at the point when and where they happen. A workload with a defined policy and behavioral detection can deny a process it has never run, whatever triggered it. It can also deny processes that behave like malware, even if they have never seen it before. The malicious logic is still present in the dependency, but it never gets to act. That is prevention, not after-the-fact detection. It also does not depend on anyone having found the package first.

The runtime catch moves left. A denied execution is evidence. It identifies the dependency, the behavior it attempted and the workload it attempted it in. That evidence can shift left and become a build policy, an admission rule and a remediation priority for every other application carrying the same component. One execution attempt in one production workload becomes prevention across the pipeline. Pre-deployment controls did not catch the package, but they catch it from then on, because runtime told them what to look for.

There is a limit worth naming. As documented, this campaign reached developers through fake job offers, and Terraform providers execute wherever an engineer or pipeline runs terraform init. Developer workstations and CI runners need their own controls, and production runtime security does not replace them. But the technique itself is not specific to Terraform or to developer machines. The Go module variant activated on a specific data value, which is the kind of input a production service processes every day. A dependency built on that pattern and compiled into a production workload is exactly the case this runtime first model is designed for.

Does runtime-first mean moving security to the right?

No. That would simply reverse the same linear model.

Runtime-first is not a lifecycle sequence. It is a decision model. Security teams should still inspect source, packages and images before deployment, verify provenance, control build infrastructure and block known malicious artifacts. What changes is the direction in which intelligence moves.

When the feedback loop works, runtime makes earlier security more precise. When the loop does not exist, organizations maintain one system that produces way too many irrelevant findings and another that eventually discovers which findings mattered.

How do the three supply chain security models compare?

Decision dimension Inventory-first  Risk-prioritized Runtime-first
Primary question What risks exist in our software? Which findings appear most dangerous? Which risks matter in applications actually running?
Main evidence Packages, CVEs, SBOMs and signatures Severity, exploit intelligence, asset and reachability context Production reachability, execution, behavior and threat intelligence
Prioritization Finding based Context enriched Production informed
Feedback upstream Findings drive remediation Enriched findings drive remediation Runtime evidence continuously refines remediation and policy
Response model Find, ticket and remediate Rank, ticket and remediate Prioritize upstream, then detect, enforce and contain at runtime
Security outcome Better inventory Better queues Security effort aligned to actual application risk

Runtime-first security does not discard what came before it. Inventory still matters. Vulnerability and exploit intelligence still matter. Runtime makes those inputs more useful by grounding them in what the application is actually doing, and it adds a control for the cases those inputs miss.

What should security leaders change?

Start with a look at the remediation queue. If production evidence shows that one vulnerable component is reachable and another is not, does that evidence change what gets fixed first? If the answer is no, runtime intelligence and vulnerability management are still operating as separate programs.

Then look at the feedback path. When a runtime control denies unexpected behavior from a dependency, how long does it take for that evidence to become a build policy or an admission rule? If the answer is that it never does, production is catching the same problem repeatedly instead of teaching the pipeline once.

Finally, ask what happens when a dangerous component is already running. Prioritization can be correct and remediation can still be incomplete. At that point, the security problem changes from deciding what matters to controlling what is allowed to execute.

The complete operating model should form a loop: pre-deployment controls reduce what reaches production, runtime tells teams which remaining risks matter most, enforcement stops unsafe behavior at execution, and every runtime catch strengthens the controls that come before it.

How can Aqua help make supply chain security runtime-first?

The Aqua Platform connects application context from code and images with live workload behavior in production. Before deployment, Aqua evaluates images as part of software supply chain security and vulnerability management, because known risk should be addressed before production whenever possible.

Once applications are running, Aqua adds the production context that changes prioritization. Aqua correlates vulnerability intelligence with runtime context and can use factors including reachability, Exploit Prediction Scoring System (EPSS) data and evidence of active exploitation to help teams separate exploitable risk from theoretical findings. Aqua also observes process activity, network connections, file changes and privilege usage under actual operating conditions.

Runtime also provides control when prioritization is not enough, including for dormant code that no prioritization model would have flagged. Aqua can deny unauthorized actions at the point of execution, block malicious execution, prevent privilege escalation and contain attacker activity in running workloads. Those controls do not replace fixing the software supply chain. They address the period when software is already running and the organization needs a control that does not depend on having found the problem first.

FAQ
Does runtime-first supply chain security replace scanning?

No. Scanning, provenance, SBOMs and pre-deployment policy remain necessary because known risks should be addressed before production when possible. Runtime-first means using production evidence to decide which upstream findings deserve attention first, and providing control when unsafe behavior executes.

Would runtime prioritization have flagged the Graphalgo malware before it activated?

No. Code that stays inert until a specific input arrives sits on a path normal operation does not reach, so a reachability model would rank it low. Runtime enforcement solves that case differently: it evaluates behavior at the moment of execution and can deny it, and the resulting evidence can inform build and admission policy upstream.

Could static analysis have detected the Graphalgo malware?

Aikido’s research does not establish that the malicious packages were impossible to identify statically. It establishes that the malicious behavior was conditionally activated by specific inputs, which makes execution conditions relevant to understanding how the malware operated.

What does runtime enforcement add after prioritization?

Prioritization determines what deserves attention first. Enforcement addresses what happens when unsafe or malicious behavior attempts to execute, including behavior from components no one had identified as risky. Runtime controls can deny specific actions and contain malicious activity while the underlying issue is addressed.

Matthew Richards
Matt is the Chief Operating Officer at Aqua Security. Prior to Aqua, he was the CMO of Datto where he helped grow the company from late-stage startup through a successful IPO in October 2020. Before Datto he served as the VP of Products and Markets at ownCloud from 2012 to 2016. He previously held management positions at CA Technologies, Novell, and IBM. Richards earned bachelor’s degrees in mechanical engineering and engineering sciences from Dartmouth College and earned his MBA from the MIT Sloan School of Management.
Need to secure enterprise workloads?

Aqua Cloud Native Application Protection Platform (CNAPP)

Go cloud native with the experts!