Outlier Detection¶
Purpose of This Document
This document is the second step in the detection rule configuration process. After completing the configuration, return to the main document to continue with the third step: Event Notification.
The system uses algorithms to analyze the metrics or statistical data of detection targets within a specific group to identify significant outlier deviations. If the detected inconsistencies exceed a preset threshold, the system generates an outlier detection anomaly event for subsequent alert tracking and analysis. This method helps identify and handle potential anomalies in a timely manner, improving monitoring accuracy and response speed.
It is suitable for configuring appropriate distance parameters based on the characteristics of metric data so that alerts are triggered when data deviates significantly from the normal range. For example, when the memory usage of a host is significantly higher than that of other hosts in the same group, the system can promptly issue an alert to quickly identify and respond to potential performance issues.
Detection Configuration¶
Detection Frequency¶
Set the time interval for performing detection, which automatically matches the selected detection range.
- Selected by default: 5 minutes
Detection Range¶
Set the data time range queried for each detection (❗️ The detection range must be greater than or equal to the detection frequency and must match the actual data reporting interval to avoid missed or false detections).
| Detection Range (dropdown options) | Detection Frequency |
|---|---|
| 15m | 5m |
| 30m | 5m |
| 1h | 15m |
| 4h | 30m |
| 12h | 1h |
| 1d | 1h |
Detection Metric¶
Define the detection data source and aggregation method based on DQL (❗️ Avoid selecting high-cardinality fields as detection dimensions. Improper configuration with overly permissive trigger conditions may cause frequent alerts. The maximum number of records returned by a query is 100,000).
Configuration Elements¶
| Configuration Item | Description |
|---|---|
| Workspace | Defaults to the current workspace; you can switch to other authorized workspaces. After authorization, you can use detection metrics from other workspaces under the current account to create monitors. Once the rule is created successfully, cross-workspace alert configuration is enabled. Note that when you select another workspace, the detection metric dropdown list only shows the data types that the current workspace is authorized to use |
| Data Source Type | Metrics, Logs, Infrastructure, Resource Catalog, Events, APM, RUM, Network, Profile, etc. |
| Query Method | Simple Query, Expression Query |
| Filter Conditions | Filter detection metric data based on metric labels to limit the data range being detected; supports adding one or more label filters; supports fuzzy match and fuzzy no-match filter conditions |
| Aggregation Algorithm | Avg by (average), Min by (minimum), Max by (maximum), Sum by (sum), Last (last value), First by (first value), Count by (number of data points), Count_distinct by (number of distinct data points), p50 (median), p75 (value at the 75th percentile), p90 (value at the 90th percentile), p99 (value at the 99th percentile) |
| Detection Dimensions | Any string-type (keyword) field in the configured data can be selected as a detection dimension. Currently, up to three fields can be selected as detection dimensions. By combining multiple detection dimension fields, you can define a specific detection target. The system determines whether the statistical metrics corresponding to a detection target meet the thresholds of the trigger condition, and generates an event if the condition is met.(For example, if the detection dimensions host and host_ip are selected, the detection target can be {host: host1, host_ip: 127.0.0.1}.) |
| Alias | Custom name for the detection metric |
Click to view Query Method Details.
Trigger Conditions¶
Configure trigger conditions and severity levels. When the query returns multiple values, an event is generated if any value meets the trigger condition.
Outlier detection uses the DBSCAN algorithm to identify anomalies and supports configuring Critical-level thresholds and Normal recovery conditions.
| Severity | Configuration | Description |
|---|---|---|
| Critical | DBSCAN algorithm, distance [value] |
Uses the DBSCAN algorithm to detect outliers. The distance parameter specifies the maximum distance between two samples for one sample to be considered a neighbor of another (float, default=0.5). ❗️ You can optionally configure any float value within range(0-3.0). If not configured, the default distance parameter is 0.5. The larger the distance, the fewer outliers are detected. A distance that is too small may detect a very large number of outliers, while a distance that is too large may result in no outliers being detected. Set an appropriate distance parameter based on the characteristics of your data. |
| Normal | No events generated after [N] detections |
After the detection rule takes effect, if the data detection result returns to normal (from Critical) within the configured number of detection cycles, a recovery alert event is triggered. ❗️ Recovery alert events are not subject to alert silence. If the number of detections for recovery alert events is not configured, the alert event will not recover and will remain in Events > Unrecovered Events List |
For more details, see Event Level Description.
Bulk Alert Protection¶
Enabled by default.
When the number of alerts generated by a single detection exceeds the preset threshold, the system automatically switches to a status-based aggregation strategy: instead of processing alert targets one by one, it generates and pushes a small number of summary alerts based on event status.
This ensures timely notifications while significantly reducing alert noise and avoiding timeout risks caused by processing too many alerts.
When this switch is enabled, event details generated after the monitor detects anomalies will not display historical records or related events.
Data Gaps¶
Strategy for handling empty query results for the detection metric within the detection range:
| Option | Description |
|---|---|
| Do not trigger events (default) | Linked to the time range of the detection interval, determines whether to generate an event based on the query result of the detection metric over the most recent few minutes. Suitable for scenarios where missing data is acceptable |
| Treat query results as 0 | Linked to the time range of the detection interval, treats the query result of the detection metric over the most recent few minutes as 0 and compares it again with the thresholds configured in Trigger Conditions above to determine whether to trigger an anomaly event |
| Custom fill and trigger events | Supports custom fill values for the detection interval and independently triggers the following event types: Data Gap Event, Critical Event, Important Event, Warning Event, and Recovery Event. ❗️ When selecting this strategy, it is recommended that the custom data gap duration be configured as ≥ the time interval of the detection range; if the configured duration is ≤ the detection range interval, data gaps and anomalies may be satisfied simultaneously, in which case the data gap handling result takes priority |
When trigger conditions, data gaps, and information generation are configured together, the triggering priority is evaluated as follows: Data gaps > Trigger conditions > Information event generation.
In other words: first determine whether data is missing, then whether the threshold is triggered, and finally whether an information event is generated.
Information Generation¶
After enabling this option, the system writes all detection results that do not match the trigger conditions above as "info" events.
Suitable for scenarios where normal state changes or low-priority information needs to be recorded.
Subsequent Configuration¶
After completing the detection configuration above, continue to configure:
-
Event Notification: define the event title, content, notified members, data gap handling, and linked incidents;
-
Alert Configuration: select an alert policy, configure notification targets and the mute period;
-
Link: link dashboards for quick navigation to view data;
-
Permissions: set operational permissions to control who can edit/delete this monitor.