Monitoring¶
TrueWatch Monitoring continuously runs detection on observability data in the workspace, such as metrics, logs, traces, RUM, and objects. When anomalies are detected, the system generates events and sends notifications to designated targets based on the associated alert policy. You can also use mute rules to control notification noise and use SLO to continuously measure service stability.
How Monitoring Works¶
The monitoring workflow consists of four stages: detection, events, alerting, and governance:
- Run detection: Monitors query and analyze observability data at the configured frequency to determine whether anomalies exist;
- Generate events: When detection conditions are met, the system generates an event or event report and aggregates them in Events;
- Send alerts: The Alert Policy sends alert information to designated Notification Targets based on configuration such as event severity, notification time, and aggregation method;
- Continuous governance: Use Mute Management to control alert notifications in specific scenarios and track service objectives and error budgets with SLO.
Events and Alert Notifications
Events are data records generated by monitors after detecting anomalies; they preserve the anomaly state and analysis results. Alert notifications are messages sent out by the alert policy based on events and notification rules. Mute and repeated alert configurations affect notification delivery, but events can still be viewed in Events.
Choose a Detection Method¶
Select the appropriate monitor based on how anomalies are determined and your configuration needs:
| Detection Type | How to Configure and Run | Use Cases |
|---|---|---|
| Rule Detection | Manually select the data scope, query method, and detection rules; configure multiple detection modes such as threshold, sudden change, interval, and outlier. | You have defined the metrics, judgment conditions, and thresholds to monitor and need precise control over the detection logic. |
| AI Smart Monitor | Use a natural language prompt to generate one or more specific detection rules; after saving, the monitor runs continuously according to fixed query conditions, algorithms, and thresholds. | You have clear monitoring objectives and expected anomaly behavior but want AI to help generate the specific rules. |
| Built-in Smart Monitoring | Select a type such as Host, Logs, Applications, RUM, Kubernetes, or Cloud Billing, load the built-in system rules, and adjust thresholds as needed. | You need to quickly cover common monitoring targets and typical anomaly scenarios. |
| AI Monitor | Use a natural language prompt directly as a continuously active detection rule; at each detection run, AI queries and analyzes the data and generates an event report with conclusions and evidence. | You need to combine business semantics, multiple signals, and context for comprehensive judgment. |
How to distinguish an AI Smart Monitor from an AI Monitor
An AI Smart Monitor uses a prompt to generate fixed rules, which take effect after saving; an AI Monitor uses the prompt directly as the detection rule, and AI analyzes data according to the prompt at each detection run. If you need explicit thresholds and fixed judgment logic, use an AI Smart Monitor; if you need ongoing semantic analysis and comprehensive judgment, use an AI Monitor.
Configuration and Processing Workflow¶
1. Create a Monitor¶
Go to Monitoring and choose Rule Detection, Smart Monitoring, or AI Monitor based on your detection target:
- Use monitor templates to quickly create monitors for common technology stacks and business components;
- Create a custom monitor and configure the query, detection rules, and trigger conditions yourself;
- Create a smart monitor to use AI to generate rules or load built-in rules;
- Create an AI Monitor and use a prompt to describe the analysis target, judgment clues, and expected results.
2. Configure Alerts¶
Monitors detect anomalies, and alert policies decide how to notify. When creating or editing a monitor, you can associate an existing alert policy, or go to Monitoring > Alert Policy for unified management.
When creating an alert policy, you can configure the following:
- The event scope and severity to handle;
- Notification targets, notification methods, and notification time;
- Repeated alert, alert escalation, and alert aggregation methods.
3. View and Handle Events¶
After a monitor detects an anomaly, you can view the results through the following entry points:
- Click View Related Events in the monitor list to view events generated by the current monitor;
- Go to Events to centrally filter, analyze, and handle events in the workspace;
- Open the event details to view information such as the event content, analysis report, extended fields, alert notifications, and history.
4. Manage Alert Noise and Service Objectives¶
- Mute Management: Suppress alert notifications by monitor, alert policy, label, or custom condition during maintenance windows or known anomalies;
- SLO: Continuously track whether service availability or performance objectives are met and use error budgets to evaluate service stability.
Features¶
| Feature | Purpose |
|---|---|
| Monitor | Create monitors using templates or custom rules. |
| AI Monitor | Use natural language prompts to continuously analyze observability data. |
| Smart Monitoring | Use AI to generate rules or select built-in system rules. |
| Alert Policy | Configure notification conditions, targets, time, escalation, and aggregation. |
| Notification Targets | Manage notification channels such as email, SMS, bots, and Webhooks. |
| Mute Management | Suppress alert notifications under specified conditions and within time ranges. |
| SLO | Set and track service objectives and error budgets. |