400 Dependabot alerts. Where do you start?
An example with 400 alerts: assess exploitation, exposure and impact to find what needs urgent action and give the rest a reasoned plan.
The day the alarm got taped over
Imagine a team turning on dependency scanning. The numbers here are an example, not a measurement from a project. It starts with good craftsmanship: visibility at last!
Then the report arrives. 400 findings. Red, orange and yellow in every direction.
For the first few weeks somebody tries to work through them. Then sprint pressure arrives, and the list grows. After three months the alerts are background noise: a channel everyone has muted, a number nobody mentions.
This is alert fatigue. The scan still finds weaknesses, but without someone to assess and follow up on the findings, it offers little help in deciding what the team should do first.
400 is not one number. It’s four.
The mistake is treating the list as a single queue. Start by identifying what needs urgent action. Then divide the rest according to how it should be handled:
1. Findings that need prompt action. A known exploited vulnerability that an attacker can reach in your system belongs high in the queue. So may a severe weakness with a plausible attack path even if exploitation has not yet been observed. That path may go through an internal service or a build tool; internet exposure is not a requirement.
Assign an owner and a short deadline for a fix or mitigation. If the attack path is unclear, give the investigation a deadline too. In this example, we assume twelve of the original 400 findings land here. Your number may be different.
2. Findings assessed as less exposed in your usage. The library is vulnerable, but you have established that the affected function is unused, or that a specific control limits the attack path. The upgrade can then be planned as ordinary maintenance. Record why it was deprioritized, who will follow up, and when the assessment should be revisited. The assumption “we don’t use that function” needs to be rechecked when code, configuration or knowledge about the vulnerability changes.
3. Findings that can be upgraded in batches. The remaining findings assessed as lower risk can often be handled in batches through routine upgrades. But “transitive” only means that a package comes in through another dependency. And a build tool may have access to source code, keys and release artifacts even if it never ships to production. GitHub describes how a compromised runner can expose secrets and the repository. Assess where the tool runs and what it can access before putting the finding in the maintenance queue.
4. Zombie dependencies. Packages nobody remembers the reason for. Check whether they are still used, including in builds and tests. If they are redundant, remove them and check that the system still works. That removes both the vulnerability and the need for future upgrades of the package.
So the sorting criterion is not the CVSS score alone. Ask whether exploitation is known or likely, whether an attacker can reach the vulnerable code, and what the impact would be. These questions need to be considered together, not treated as a formula for multiplying three numbers. Severity needs context to become an order of work.
Two free sources help with the question of exploitation:
- EPSS (Exploit Prediction Scoring System, from FIRST) estimates the probability of observed exploitation of a CVE in the wild over the next 30 days. The score is updated daily and freely available via API. Dependabot lets you filter on it. FIRST reports that its partners observe exploitation activity for around 2.5–3 percent of published CVEs in a 30-day window. That describes observed activity across systems, not the probability that your particular system will be affected. A low score does not exempt you from assessing the impact in your environment.
- CISA KEV (Known Exploited Vulnerabilities Catalog) lists vulnerabilities with confirmed exploitation in the wild. It had 1,733 entries on October 2, 2026. A match is a strong signal to investigate and act quickly, even when EPSS is low. Check affected versions, attack paths and controls in your environment; the catalog does not know your configuration.
A concrete example of this approach is CISA BOD 26-04, issued on June 10, 2026, for US federal civilian agencies. Its deadlines depend on exposure, KEV status, whether an attack can be automated, and how much control the attacker gains. They range from three to sixty days, while some combinations can wait until the next major upgrade or rebuild.
Some cases also require an investigation into whether the system is already compromised. This is their framework, not a deadline table Norwegian teams are required to follow. It shows how context can determine deadlines.
To investigate whether your application calls the vulnerable code, you can use reachability analysis. Snyk and Semgrep offer it. GitHub also uses detected vulnerable function calls in prioritization. Check which languages, packages and integrations the tool supports. Semgrep Supply Chain is free for organizations with up to ten contributors under its counting rules; above that limit, a paid license is required.
The analysis helps you prioritize, but it does not clear the rest of the list. Snyk stresses that failing to find a call path does not prove the code is unreachable. Dynamic calls and incomplete information can create blind spots. Use the result to support a reasoned assessment, and keep a plan for findings that are not urgent.
Make the queue impossible to ignore by making it small
Three mechanisms keep the system healthy after the sorting:
- Stop new urgent findings before introducing them. A PR check can block new dependencies with risk the team cannot accept. GitHub’s dependency review examines dependency changes; it does not resolve the existing backlog by itself. Give existing critical findings an owner and a deadline, and make sure the gate lets the fix itself through. A necessary exception needs a reason, an owner with authority to accept the risk, a mitigation and an expiry date.
- The rest gets batched on a rhythm. A fixed, small dose of upgrades every sprint: boring, predictable, never a heroic cleanup sprint. Same logic as all improvement that actually happens (in Norwegian): it lives in the day-to-day.
- Measure how long critical risk remains. Track the time from discovering a critical finding until the fix or mitigation is verified in the affected environment. A merged PR is not enough. Also look at which critical findings remain open and how old they are; a falling total can hide one important finding that is never addressed.
Out of the background noise
400 alerts is not visibility. It’s fog with color codes.
In the example, twelve findings received prompt attention. The other 388 were not declared safe; they got a reasoned plan for maintenance or removal.
Start with the oldest finding you assess as critical. Agree on who will follow up, when the risk should be reduced, and how you will verify it.