Dependency Scanning Automation That Runs Itself
Dependency Scanning Automation That Runs Itself
TL;DR: I was spending 45 minutes every Monday manually running dependency audits across three projects, reviewing the output, and deciding what to escalate. I automated the entire workflow — scan, triage, ticket creation, and summary — into a pipeline that runs unattended and only pages me when something is genuinely actionable. Here’s the exact pattern and why it’s reusable across any recurring security task.
Every enterprise security team has a version of this problem. There’s a manual task that has to happen regularly — weekly, daily, sometimes multiple times a day — and it’s too important to skip but too repetitive to stay sharp on. You do it on autopilot, which is exactly when you miss things.
For me, the task was dependency scanning. Three active repositories, each with its own package.json or requirements.txt, each capable of picking up a new CVE between one Monday and the next. The right answer is always “automate this.” The harder question is how to automate it in a way that’s actually trustworthy — not just a cron job that silently fails or floods a Slack channel with noise nobody reads.
In 25 years of enterprise security work, I’ve seen more “automated” processes that created alert fatigue than ones that genuinely removed human toil. The difference is usually in how the output is designed, not the scan itself.
This post covers the full pattern: what I built, the decisions I made, and how to adapt it to your own recurring manual tasks.
The Problem With “Just Schedule the Script”
The naive solution is to take whatever you’re running manually and wrap it in a cron job. Scan at 8am Monday, output goes somewhere. Done.
This fails in predictable ways:
- No triage logic. You get 40 findings, three of which matter. The human who used to run this manually knew which three to care about. The cron job doesn’t.
- No failure visibility. The script errors out, the scan doesn’t run, nobody notices. The false sense of coverage is worse than not automating at all.
- Output designed for humans, consumed by no one. Dependency scanner output is verbose and context-free. It tells you a CVE exists. It doesn’t tell you whether you’re actually using the affected code path, whether a patch is available, or whether you already accepted this risk last quarter.
For automation to actually remove toil — rather than redistribute it — you need to build three things: the scan, the triage, and the escalation gate.
The Dependency Scanning Automation Pattern
Here’s what I built. The components are simple. The design decisions are what matter.
Component 1: The Scan (the easy part)
For Python projects I use pip-audit. For Node I use npm audit. Both produce structured JSON output, which is the critical requirement — you cannot reliably parse human-readable output downstream.
pip-audit --output json --format json -o audit-results.json
Or for Node:
npm audit --json > audit-results.json
The scan runs on a schedule via GitHub Actions. I chose GitHub Actions over a local cron job for one reason: it’s observable. You can see the run history, the logs, the exit codes. A cron job that fails at 3am fails silently. A GitHub Actions workflow that fails sends a notification and leaves a red badge.
Component 2: The Triage (the part that makes it useful)
This is where most automation stops short. Raw scanner output goes into a bucket, and you’ve just moved the manual work from running the scan to reviewing the output.
I wrote a small Python triage script that does four things:
- Parses the JSON output
- Filters out findings below a configurable severity threshold (currently CVSS 7.0 or higher)
- Checks against a local
accepted-risks.jsonfile — if a finding has been reviewed and explicitly accepted, it doesn’t re-escalate - Checks whether a fixed version is available
import json
import sys
SEVERITY_THRESHOLD = 7.0
ACCEPTED_RISKS_FILE = "accepted-risks.json"
def load_accepted_risks():
try:
with open(ACCEPTED_RISKS_FILE) as f:
return json.load(f)
except FileNotFoundError:
return []
def triage(audit_file):
with open(audit_file) as f:
findings = json.load(f)
accepted = load_accepted_risks()
actionable = []
for vuln in findings.get("vulnerabilities", []):
cvss = vuln.get("cvss", 0)
vuln_id = vuln.get("id")
if cvss < SEVERITY_THRESHOLD:
continue
if vuln_id in accepted:
continue
actionable.append(vuln)
return actionable
if __name__ == "__main__":
results = triage(sys.argv[1])
print(json.dumps(results, indent=2))
sys.exit(1 if results else 0)
The exit code is intentional. Zero means nothing actionable. Non-zero means something needs attention. GitHub Actions can branch on exit codes, which is how the escalation gate works.
Component 3: The Escalation Gate
If the triage script exits non-zero, the workflow creates a GitHub Issue with a structured summary. If it exits zero, the workflow logs “clean scan” and terminates.
- name: Run triage
id: triage
run: |
python triage.py audit-results.json > actionable.json
continue-on-error: true
- name: Create issue if actionable findings exist
if: steps.triage.outcome == 'failure'
uses: actions/github-script@v7
with:
script: |
const fs = require('fs');
const findings = JSON.parse(fs.readFileSync('actionable.json', 'utf8'));
const body = findings.map(f =>
`**${f.id}** — CVSS ${f.cvss}\n${f.description}\nFixed in: ${f.fix_version || 'no patch available'}`
).join('\n\n---\n\n');
await github.rest.issues.create({
owner: context.repo.owner,
repo: context.repo.repo,
title: `[SECURITY] Dependency audit findings — ${new Date().toISOString().split('T')[0]}`,
body,
labels: ['security', 'dependencies']
});
The issue contains exactly what’s actionable: the CVE IDs, severity scores, affected packages, and whether a fix is available. No noise. No low-severity informational findings. No re-escalation of risks already accepted.
The Accepted Risks Pattern (and Why It’s Not a Shortcut)
The accepted-risks.json file deserves its own discussion, because it looks like a way to make findings disappear and it isn’t.
Every entry in that file requires a comment explaining the acceptance rationale and an expiry date. After 90 days, the finding re-escalates regardless. This isn’t optional — it’s enforced by the triage script checking the expiry field.
[
{
"id": "GHSA-xxxx-xxxx-xxxx",
"accepted_by": "jared",
"reason": "Only affects Python 2.x codepath we don't use. Upstream has no patch.",
"expires": "2025-10-01"
}
]
In enterprise IAM work, this is the same principle as access review cycles. You don’t grant permanent exceptions — you grant time-bounded exceptions that require re-evaluation. Risk acceptance without an expiry is just risk deferral with extra steps.
The Pattern Is Reusable
I’ve now applied the same three-component structure — scan, triage, escalate — to two other recurring tasks:
- SSL certificate expiry checks: Script queries each domain, filters certs expiring within 30 days, issues only if action is needed
- IAM permission drift detection: Compares current IAM policy JSON against a committed baseline, escalates on any new wildcard or cross-account permission
The implementation differs, but the pattern doesn’t. Structured output, severity filter, accepted-risk exclusion list with expiry, exit-code-driven escalation gate.
If you have a recurring security task that a human is doing on a schedule, ask: does this task produce output that can be evaluated programmatically? If yes, this pattern applies.
What I Got Wrong the First Time
The first version had no accepted-risks logic. Every scan, the same three legacy dependency findings escalated. Within two weeks I stopped reading the issues. Alert fatigue from your own automation is a real failure mode.
The second version had accepted risks but no expiry enforcement. I accepted a finding in February, forgot about it, and it was still silently suppressed in August when the upstream project actually patched it and a newer related CVE appeared under a different ID. I only caught it because I was auditing the automation itself, not because the system caught it.
The third version — the one described above — has held up. Expiry enforcement is the thing most people skip and the thing that makes the pattern trustworthy long-term.
Key Takeaways
- Automation that doesn’t triage creates a different kind of toil. Moving from manual scanning to automated noise is not progress.
- Exit codes are the correct escalation mechanism. Build your triage scripts to exit non-zero on actionable findings and let your orchestration layer handle the branching.
- Accepted-risk lists need expiry dates. Risk acceptance without re-evaluation is just deferral. Enforce the expiry in code, not in policy documents nobody reads.
- GitHub Actions over cron for observability. Silent failures in security automation are worse than no automation.
- The three-component pattern — scan, triage, escalate — is portable. Every recurring security check I’ve automated follows this structure. Build it once, apply it everywhere.
The 45 minutes I was spending every Monday now happens unattended. I see a GitHub issue when there’s something that requires a decision. I see nothing when everything is clean. That’s what useful automation actually looks like.
Comments