Research · Measurement

How security gateways distort phishing simulation results.

A click is an event, not yet a person.

What causes false positives?

Email security gateways and related cloud services may follow links, download remote images, submit URLs for analysis, or revisit content from a different network. Prefetch features in mail clients and browsers can produce similar requests. These events are useful evidence, but they are not automatically employee intent.

A simulation platform normally sees an HTTP request before it knows why the request exists. The request may come from the employee, a corporate proxy acting for the employee, a link scanner running before delivery, a sandbox detonating the message, or a later analyst review. The difficult part is not collecting the event. It is assigning the right interpretation.

The safest operating rule is to preserve the raw event and calculate the human-facing metric from a classified view. That keeps the evidence auditable while preventing automated traffic from silently inflating click, open, or submission rates.

Collect context before classifying.

No single header, IP range, or timing rule can identify all security traffic. A defensible classifier combines several weak signals and records why an event received a label.

SignalWhat it can suggestWhy it is not enough alone
TimingRequests immediately after send may be automated inspection.A fast employee action is still possible.
Network ownershipCloud, security-vendor, or data-center networks may indicate a scanner.Employees can use cloud browsers, VPNs, or shared egress.
User agent and TLS fingerprintRepeated machine-like clients can reveal a scanning fleet.Fingerprints change and can be shared by legitimate clients.
Request sequenceAsset-fetch patterns and link traversal can look unlike human navigation.Privacy and prefetch features can mimic partial browsing.
Cross-event consistencyA later browser event from the employee environment can confirm likely human action.Some environments intentionally minimize identifiable context.

Environmental signals can be privacy-sensitive. Collect only what the authorized exercise and applicable policy permit, define retention, and avoid turning a training program into covert employee monitoring.

Use a staged, explainable method.

  1. Preserve the raw event. Keep the original timestamp and the evidence needed for later review.
  2. Apply deterministic labels first. Known internal scanners, allowlisted test systems, and documented gateway behavior should be handled before probabilistic scoring.
  3. Combine multiple contextual signals. Timing, network, client, sequence, and campaign history are more useful together than separately.
  4. Represent uncertainty. “Likely scanner” and “unclassified” are more honest than forcing every event into human or machine.
  5. Review important exceptions. A disputed high-risk event should be explainable to the security team and, where appropriate, the employee.
  6. Recalculate the metric from the labeled view. Do not delete evidence to make the number cleaner.
TaiGong 1.3.4 boundary

Core campaign events and basic analytics are available in the current product. Advanced environmental analysis, threat intelligence, and reporting are licensed capabilities. The exact current boundary is published in the support matrix.

Report a range of evidence, not one heroic percentage.

A useful report separates delivery, observed requests, classified human actions, excluded security traffic, uncertain events, employee reports, and follow-up. It also records the classification policy used for that exercise. This makes results comparable over time and prevents a gateway policy change from being mistaken for a sudden change in employee behavior.

When two departments use different mail gateways, compare like with like. When a gateway configuration changes between quarters, annotate the series. The quality of the measurement model matters as much as the campaign design.

A classifier improves judgment; it does not create certainty.

Security infrastructure and legitimate clients keep evolving. Some automated systems imitate browsers, while privacy protections intentionally remove identifying signals. The goal is not perfect attribution. It is a transparent method that is less wrong, reviewable, and appropriate for an authorized awareness program.

See the human-risk operating model →