Explore RIFT.

22 pages
Illustrative event trace moving through telemetry and alerting panels.
Detection validation

An event is only the beginning.

A detection rule is one part of a longer chain. The event must be observable, the relevant data must reach the right place and the resulting alert must give someone a useful starting point. Validation follows that chain under defined conditions and records where the expected outcome differs from what was actually observed.

LONG-FORM / 10 MIN READDEFINED SCOPE · USEFUL EVIDENCE
Keep the output useful

What the conversation should produce.

  • A telemetry coverage matrix listing each source, the fields it provides, and the confirmed delivery path to the collection layer.
  • A rule inventory entry for each detection with owner, last validation date, known gaps, and the next scheduled re-validation condition.
  • A response log recording alert receipt, routing destination, assignment, and first investigative action for every validation test run.
BRIEFING / 01

Start with the behaviour or event the detection should identify

Detection validation begins by stating the behaviour you expect to see, not the alert you hope to fire. Write the activity as an observable sequence: who acted, what tool was used, where the data moved, and which system recorded it. A vague target such as "unauthorised access" produces vague tests. A precise target such as "a service account used to query an external storage endpoint after hours" produces a testable hypothesis. The difference determines whether your validation reveals gaps or merely confirms assumptions.

State the environment and data sources on which the detection depends. A rule designed around particular event fields may need different collection or normalisation in another environment, even when the underlying behaviour is relevant to both. Cross-environment coverage should be demonstrated rather than assumed from a rule name. Describe the identities, systems and data relationships involved, then compare that description with the available telemetry and actual logic. This makes the intended boundary clear enough for engineering and business owners to discuss the same validation question.

BRIEFING / 02

Check that required telemetry exists and reaches the right system

Check whether the required telemetry is available and whether the selected detection can use it as expected. Inspect representative records from the relevant sources, including the fields, timing and transformations that matter to the rule. Missing or changed fields, delayed delivery and filtering can all affect the outcome. The effect depends on the actual detection logic, including any fallback behaviour, so avoid declaring the cause from a schema difference alone. Record the observation and the next check needed to establish its significance.

Telemetry validation also requires checking routing and retention. Logs may reach the collector but be discarded by an intermediary filter, or stored in a tier that expires before investigation requires them. Confirm the pipeline path end-to-end and the retention window against your investigation timeline. Document which sources are confirmed, which are partial, and which are absent. This record becomes the baseline for every future validation and prevents the illusion that a detection exists simply because a rule sits in a repository.

BRIEFING / 03

Define expected alert content, routing and investigation context

An alert is not a conclusion; it is a handoff. Define what the alert must contain: the triggering event, the affected identity, the relevant systems, and the confidence level assigned by the detection logic. Vague alerts force analysts to reconstruct context from scratch, which delays response and increases fatigue. Specify the routing destination: which team receives the alert, which ticketing system captures it, and which escalation path activates when severity thresholds are crossed. Routing decisions should be written, not inherited from tribal knowledge.

Define the context an analyst needs to begin the agreed investigation. That may include the relevant identity, system, event source, rule reference and a way to obtain supporting records. Not every useful piece of context has to be embedded automatically in the alert; an accessible and understood investigation path may also serve the purpose. Validate the handover against the actual workflow, including permissions and availability of supporting data. An alert that technically exists but cannot be interpreted usefully leaves a different problem from a rule that never matched.

BRIEFING / 04

Run a bounded validation and record missed and partial outcomes

Run the agreed validation and record full, partial and absent outcomes without assigning their cause too early. A rule that matches one part of a scenario may be behaving as designed, or it may reveal a coverage gap. Missing fields in an alert may originate in collection, transformation, enrichment or presentation. Compare the observation with the stated expectation and trace the relevant evidence through the chain. Preserve enough context to support a targeted follow-up rather than turning every missed expectation into a single generic failure.

Bounded validation requires explicit boundaries. State which systems were tested, which were excluded, and why. State the time window, the identities used, and the data volume. Without boundaries, results are unrepeatable and conclusions are untrustworthy. Record the test conditions alongside the outcomes so that a future reviewer can reproduce the exercise and compare results. Where exact repetition is unavailable or inappropriate, state the limitation and preserve the context needed to interpret the observation.

BRIEFING / 05

Distinguish absent events from absent rules and absent response

Separate the stage where an expected outcome was absent from the cause of that absence. The source may not have recorded the event, the data may not have arrived, the rule may not have matched, or the alert may not have reached an investigation workflow. A rule can exist without covering the selected conditions, and a missing response can have several contributing causes. Use evidence from the relevant owners to narrow the explanation before deciding which change is appropriate.

The distinction matters because it determines where investment yields results. Throwing more detection rules at a telemetry gap wastes engineering effort. Throwing more alerts at a response gap wastes analyst time. The disciplined approach is to classify every missed outcome before deciding on a remedy. This classification becomes the evidence base for prioritisation, and it prevents the organisation from treating symptoms while the underlying failure persists. Absence is informative only when you know which kind of absence you are looking at.

BRIEFING / 06

Document rule changes, ownership and a future validation condition

Detection logic changes constantly: rules are tuned, retired, replaced, and re-tuned. Without documentation, every change erases the evidence that justified it. Record what changed, why it changed, who approved it, and what validation condition triggered the change. A rule that was relaxed because it produced excessive noise is not the same as a rule that was relaxed because the noise was misclassified. The distinction determines whether the relaxation was correct and whether it should be reversed when conditions shift. Ownership is equally critical: every rule must have a named owner who is accountable for its validation, its tuning, and its retirement.

Define when the detection should be examined again and what question that check will answer. A relevant system change, altered telemetry, a rule revision or an agreed review point may be a useful trigger. Keep the earlier expectation and observation available so the team can compare meaningfully. Revalidation is a way to examine continued performance under stated conditions; the existence of a scheduled test alone does not establish effectiveness, and a single successful result should not erase known limits.

Illustrative scenario / Not a client case study

Illustrative scenario: service account querying external storage after hours

A security team receives an alert that a service account accessed an external storage endpoint at 02:17. The alert fires from a rule monitoring outbound API calls. The team must determine whether this represents a genuine compromise, a misconfigured automation job, or a telemetry artefact. The scenario is illustrative and used to demonstrate how detection validation distinguishes evidence from assumption. It does not describe any real incident, any real organisation, or any real detection rule. The purpose is to show how the six validation questions apply to a concrete situation.

  1. Confirm the expected use of the service identity with its owner and compare that context with the observed event. An unfamiliar time or destination is a reason to investigate, not sufficient evidence by itself to classify the activity.
  2. Check whether the outbound call was logged by the network device or the application, since each source carries different fidelity and different blind spots.
  3. Compare the alert with the context requirements agreed for this detection, including the relevant identity, destination and supporting event reference. Record which missing information affected the actual investigation rather than requiring a universal set of fields.

This scenario establishes that an alert is a starting point, not a conclusion. It shows how validation separates telemetry gaps from rule gaps from response gaps, and why each requires a different remedy. It does not prove the incident is benign, nor does it prove the detection works; it proves that the team can articulate what they know, what they do not, and what they must test next.

Locate the next useful validation question

QuestionWhat to establishUseful output
Did the behaviour occur or did the telemetry miss it?Whether the event was recorded at the source and reached the collector intactA telemetry coverage matrix listing each source, field, and confirmed delivery path
Which rule fired and which rule should have fired?The exact rule ID, its tuning parameters, and the scope boundaries it was designed to coverA rule inventory entry with owner, last validation date, and known gaps
Was the alert routed and investigated, or did it sit unattended?The routing destination, the ticket created, the analyst assigned, and the time from fire to triageA response log showing alert receipt, assignment, and first investigative action

Scroll the table horizontally on smaller screens.

Useful questions.

How do we know a validation test was thorough enough?

Judge adequacy against the agreed question, relevant variations and evidence requirements, not only whether a test can be repeated. Record the conditions, exclusions and observed outcomes, then ask whether they support the intended conclusion. Repeatability helps, but a repeatable test can still be too narrow. Unanswered variations should remain visible for the next planning decision.

What should we do when a rule fires but produces no actionable alert?

Compare the observed alert with the agreed investigation needs, then trace the missing context through collection, enrichment, rule output and routing. Do not assume which component is responsible simply because a rule matched. Assign a follow-up to the relevant owner and verify the specific improvement. A useful diagnosis explains the gap and its effect without exaggerating what the observation proves.

Further reading: MITRE ATT&CK: adversary behaviour knowledge base. This independent resource provides background; no affiliation, certification or endorsement is implied.

The next useful question

What do you need
the evidence to tell you?

Start with the decision, the environment and the constraints. Create a scope brief you can download, review with your team and refine before any engagement is considered.

Build your scope brief ↗