Skip to content

Change Detection


About This Document

This document is the second step in the detection rule configuration flow. After completing the configuration, return to the main document to continue to step 3: Event Notification.

Change Detection identifies anomalies by comparing the absolute change or relative percentage change of the same metric across two different time periods. This method is commonly used to track metric peaks or fluctuations. When an anomaly is detected, more precise event records can be generated for subsequent analysis and processing.

It is suitable for monitoring the relative change or rate of change of data in a short period against longer-term data. For example, configure the MySQL connection count metric so that the percentage difference between the most recent 15 minutes and the average over the past 1 day is greater than 500%. This means that if the average connection count in the most recent 15 minutes exceeds 5 times the average over the past 1 day, the system will trigger an alert.

It is recommended to use statistical functions such as average (avg), maximum (max), and minimum (min) to calculate these metrics, rather than the last value (last) function, to reduce the impact of anomalous data and improve monitoring accuracy.

Detection Metric

Define the detection data source and aggregation method based on DQL (❗️ Avoid selecting high-cardinality fields as detection dimensions. Improper configuration may make trigger conditions too permissive and cause frequent alerts. The current query returns a maximum of 100,000 records).

Detect change anomalies by comparing metric data between two time periods:

Result = [difference/difference percentage] of the detection metric between [Time Period A] and [Time Period B]

Configuration Elements

Configuration Item Description
Workspace Defaults to the current workspace; you can switch to other authorized workspaces

After authorization, you can use detection metrics from other workspaces under the current account to create monitors. Once the rule is created, cross-workspace alert configuration is enabled. Note that when you select another workspace, the detection metric dropdown only displays data types authorized for use in the current workspace
Data Source Type Metrics, logs, traces, RUM data, etc.
Query Method Simple query, expression query
Filter Conditions Filters detection metric data based on metric tags to limit the data range under detection; supports adding one or more tag filters; supports fuzzy match and fuzzy non-match filter conditions
Aggregation Algorithm Avg by (average value), Min by (minimum value), Max by (maximum value), Sum by (sum), Last (last value), First by (first value), Count by (number of data points), Count_distinct by (number of distinct data points), p50 (median value), p75 (value at the 75th percentile), p90 (value at the 90th percentile), p99 (value at the 99th percentile)
Detection Dimension Any string (keyword) field in the configured data can be selected as a detection dimension. Up to three fields are supported. By combining multiple detection dimension fields, you can identify a specific detection object. The system checks whether the statistical metric of a detection object meets the trigger threshold; if so, an event is generated.

(For example, if the detection dimensions host and host_ip are selected, the detection object can be {host: host1, host_ip: 127.0.0.1}.)
Alias Custom name for the detection metric

Time Period Configuration

Configuration Item Description
Time Period A Recent data period used as the baseline
Time Period B Historical data period used for comparison
Comparison Method Difference: absolute difference between the two periods (A - B)
Difference Percentage: relative percentage change ((A - B) / B × 100%)
  • Available time periods:
Time Period Type Options
Historical Period Last month, Last week, Yesterday, 1 hour ago, Previous period, Last 15 minutes, Last 30 minutes, Last 1 hour, Last 4 hours, Last 12 hours, Last 1 day
Recent Period Last 1 minute, Last 5 minutes, Last 15 minutes, Last 30 minutes, Last 1 hour, Last 4 hours, Last 12 hours, Last 1 day
Note

For the detection windows "Yesterday" and "1 hour ago", the difference or difference percentage of the detection metric is compared within the same time range; for other detection windows, the difference or difference percentage of the detection metric is compared between two different time periods.

Click to view Query Method Details.

Detection Frequency

The execution frequency of the detection rule is automatically matched to the detection window with the larger time range among the two selected windows.

  • Selected by default: 5 minutes

Trigger Conditions

Configure trigger conditions and severity levels. When the query returns multiple values, an event is generated if any value satisfies the trigger conditions.

Supports configuring four threshold levels: Fatal, Critical, Major, Warning, as well as the Normal recovery condition.

Level Configuration Description
Fatal When Result >= [value] Highest severity alert; requires immediate handling
Critical When Result >= [value] High severity alert; requires prioritized handling
Major When Result >= [value] Medium severity alert; requires attention
Warning When Result >= [value] Low severity alert; should be noted
Normal No event generated for [N] detection cycles After the detection rule takes effect, if the detection result recovers from abnormal (Fatal, Critical, Major, Warning) to normal within the configured number of detection cycles, a recovery alert event is triggered.
❗️ Recovery alert events are not subject to alert muting. If no recovery detection count is configured, the alert event will not recover and will remain in Events > Unrecovered Events List.

For more details, refer to Event Level Description.

Trigger Prerequisites

Enabled by default. As the entry threshold for change detection, the change detection rule is evaluated only when the detection value meets the threshold set in the prerequisites.

  • Format: execute the following evaluation only when the detection value for [time period] satisfies [operator] [threshold]
  • Supported operators: >, >=, <, <= (> selected by default)
  • When this configuration is disabled, the system directly evaluates the change detection rule without an entry threshold.

Change Direction

Trigger an event when the change direction is [direction]:

  • Upward: detects changes where the data increases (rises)
  • Downward: detects changes where the data decreases (falls)
  • Upward or downward: detects changes where the data fluctuates in both directions (increases or decreases)

Bulk Alert Protection

Enabled by default.

When the number of alerts generated by a single detection exceeds the preset threshold, the system automatically switches to a status-based aggregation strategy: instead of processing alert objects one by one, it generates and pushes a small number of summary alerts based on event status.

This ensures timely notifications while significantly reducing alert noise and avoiding the risk of timeouts caused by processing an excessive number of alerts.

When this toggle is enabled, Event Details of such events generated after the monitor detects anomalies will not display history records or related events.

Data Gap

Handling strategy when the detection metric returns no query results within the detection window:

Option Description
Do not trigger events (default) Linked to the time range of the detection window; determines whether to generate an event based on the query results of the detection metric within the most recent minutes. Suitable for scenarios where missing data is acceptable.
Treat query results as 0 Linked to the time range of the detection window; treats the query results of the detection metric within the most recent minutes as 0 and compares them against the thresholds configured in Trigger Conditions above to determine whether an anomaly event is triggered.
Custom fill and trigger events Supports custom fill values for the detection window and triggers the following event types respectively: Data Gap Event, Fatal Event, Critical Event, Major Event, Warning Event, and Recovery Event.

❗️ When selecting this strategy, it is recommended that the custom data gap duration be configured as ≥ the interval of the detection window; if the configured duration is ≤ the detection window interval, both data gap and anomaly conditions may be satisfied simultaneously, in which case the data gap handling result takes priority.

When trigger conditions, data gap, and info generation are configured together, the trigger priority is: Data Gap > Trigger Conditions > Info Event Generation.

That is, first determine whether a data gap exists, then whether a threshold is triggered, and finally whether to generate an info event.

Info Generation

When this option is enabled, configure the Info Generation Conditions. Only when the detection result does not trigger any of the "Fatal", "Critical", "Major", or "Warning" thresholds and satisfies the info generation condition will the system create an "Info" event.

It is suitable for scenarios where normal state changes or low-priority information needs to be recorded.

Subsequent Configuration

After completing the detection configuration above, continue with:

  1. Event Notification: define the event title, content, notification members, data gap handling, and linked incidents;

  2. Alert Configuration: select the alert policy, and set notification targets and muting periods;

  3. Link: link dashboards for quick navigation to view data;

  4. Permission: set operation permissions to control who can edit or delete this monitor.