NexGuard.

← Field Notes · Detection

Signatures Still Work: The Case for Exact Matching

Hash-based detection is treated as the industry's oldest, least glamorous tool. It is also its most reliable.

Ask a security researcher what signature-based detection cannot do, and the list comes quickly: it cannot catch malware it has never seen, it cannot catch a file that has been recompiled or even slightly altered, and it does nothing at all against an attacker who writes custom code for a single target. These are real limitations, and they are the reason signature matching has spent two decades being described in industry writing as legacy technology, a floor to build on rather than a solution in itself.

What that framing tends to leave out is what signature matching actually buys, which is certainty. A cryptographic hash, most commonly SHA-256 in modern tools, reduces a file to a fixed-length fingerprint that changes completely if even a single bit of the file changes. When a scanner checks a file's hash against a list of hashes known to belong to malicious samples, a match is not a guess, a score, or a probability. It is confirmation that the file on the machine is byte-for-byte identical to a sample that was previously identified, analyzed, and cataloged as malicious. There is no false-positive rate to discuss, because there is no inference involved.

That certainty has a cost, and the cost is coverage. A hash list only catches what has already been collected, analyzed, and added to it, and it catches nothing that has been modified since. Attackers know this, and automated tools that recompile or repack malware to produce a new hash for every deployment are common and cheap to run. This is the honest reason no serious vendor, including NexGuard, claims that hash matching alone is a complete antivirus strategy. It is a floor, not a ceiling.

But a floor matters more than its critics usually credit. Large public malware-sharing efforts, the kind that feed threat-intelligence feeds used across the industry, catalog enormous numbers of confirmed-malicious samples precisely because so much real-world malware is reused, repackaged, and redeployed with only minor changes across many victims. A commodity infostealer sold on a criminal forum and installed on thousands of machines by dozens of unrelated buyers often does not change at all between installations. Hash matching catches that traffic with zero false positives and negligible computational cost, freeing more expensive and error-prone detection methods, heuristics, behavioral monitoring, machine-learning classifiers, to focus on what actually requires judgment.

The industry's more sophisticated products layer these approaches rather than choosing between them, and that layering is the right architecture. But it is worth saying plainly, in an industry that markets almost exclusively on the sophistication of its newest detection method, that the oldest technique in the field still does the one thing none of the newer ones can promise: when it flags a file, it is not guessing.