← All articles
July 24, 2026·7 min read

Building Threat Intel Tooling for Organizations Without a SOC: Design Decisions and Tradeoffs

ArchitectureThreat IntelEngineeringOSINT

Article one made the case for why small organizations, the businesses, nonprofits, and public sector offices without a dedicated security team, represent an underexamined gap in national cyber resilience. This one is about the harder part: what it actually takes to build something useful for them, and the specific tradeoffs involved in doing that well.

The core design constraint

Most threat intelligence tooling is built on an assumption that doesn't hold here: that the person consuming the output already has security expertise. Enterprise platforms surface raw indicators, confidence scores, and technical detail because their users are analysts who know what to do with that. The moment you remove that assumption, the entire design has to change. The hard problem isn't ingesting threat data, that part is comparatively easy. The hard problem is translating that data into something a nontechnical decision maker can act on without a security background standing between them and the signal.

That constraint shaped every decision below.

Pipeline structure

The tool runs as a five stage pipeline: ingestion, organization profiling, filtering, triage, and digest formatting. Each stage has exactly one job, and each one's output is the next stage's input, nothing more. That modularity wasn't an aesthetic choice, it's what makes each piece independently testable and replaceable. The filtering logic, for instance, can be rewritten entirely without touching how data gets ingested or how the final digest gets formatted. For a v1 built solo, that separation matters more than it would in a larger team context, since there's no code review catching a change that accidentally breaks three unrelated things at once.

Why CISA's KEV catalog, specifically, for v1

There were several plausible starting feeds: AbuseIPDB, URLhaus, AlienVault OTX, Shodan. KEV won for a reason that matters more than convenience. Every entry in CISA's Known Exploited Vulnerabilities catalog has already been confirmed as under active exploitation, not theoretical, not "could be exploited," but observed happening. For an audience with zero security background, that distinction is everything. A tool that says "this is being actively exploited right now" earns trust and prompts action. A tool that says "this could theoretically be a problem" gets ignored, correctly, since most nontechnical users have no way to evaluate that kind of uncertainty. Starting with the highest confidence, lowest ambiguity feed was a deliberate choice to establish credibility before adding noisier sources later.

The matching problem: exact versus substring

This is the piece that looked simple and wasn't. The tool needs to match an organization's declared software (from their profile) against each KEV entry's vendor and product fields. Exact string matching alone misses obvious variants, "Microsoft 365" in a profile won't match "Microsoft Office" in a KEV entry, even though they clearly overlap. But substring matching alone creates the opposite problem: a profile listing "Office" as a product will substring match against unrelated entries that happen to contain that word anywhere in a longer product name, reintroducing exactly the noise problem this tool exists to eliminate.

The resolution was a tiered approach: attempt exact match first, since it's the highest confidence signal, and fall back to substring matching only for phrases above a minimum length threshold, filtering out short, generic words that would otherwise cause false positives. Each match gets tagged with which method found it, exact or substring, so that confidence information isn't lost, it propagates downstream into the triage stage, where an exact match can be weighted more heavily than a substring match. This is a small detail with an outsized effect on whether the final digest feels trustworthy or noisy.

Why rule based triage, not machine learning

Given how much attention ML based scoring gets in security tooling generally, it's worth explaining why this project deliberately avoided it, at least for v1. Rule based triage (a KEV entry is Critical if it's an exact match and recently added, Watch if it's an exact match but older or a substring match, Informational otherwise) is fully explainable. Every priority assignment can be traced to a specific, statable reason. For an audience without the background to sanity check a model's output, explainability isn't a nice to have, it's the entire basis for trust. A nonprofit director who gets told "this is Critical" needs to be able to ask why and get a real answer, not a probability score they have no way to interpret. There's a version of this project down the line where more sophisticated scoring earns its place, but that's only worth doing once there's real usage data justifying the added complexity, not before.

The "nothing critical" decision

One easy trap in building an alerting tool is assuming silence is fine when there's nothing urgent to report. It isn't, for this specific audience. An organization with no dedicated IT security staff has no independent way to verify that a tool is still running and still checking. If the digest simply doesn't arrive on a quiet day, the most likely interpretation isn't "everything's fine," it's "the tool broke, and nobody noticed." So the digest formatter explicitly handles the empty case, sending a short, plain message confirming the check happened and nothing urgent came up. It's a small design decision, but it's the difference between a tool that gets trusted over time and one that quietly gets forgotten the first time it goes a few days without an alert.

What's next

The pipeline is fully functional end to end against live CISA data, tested and confirmed with real email delivery. The next real test isn't technical, it's operational: deploying this against an actual organization's real software profile rather than a synthetic test one, and finding out what breaks, what's missing, and what a real nontechnical user actually needs that a solo build inevitably misses on the first pass. That's the next piece of this project worth writing about.

Follow the build

This project is open source and public from day one. Track progress, read the design docs, or contribute at github.com/jdeveloping-ux/oss-threat-intel ↗