Monitoring and adjusting ML WAF protection
This section describes how to monitor changes in ML WAF (Yandex Malicious Score) behavior after a model update and adjust the protection settings. Since the model is fine-tuned regularly, the score value for the same request may change. If score reaches or exceeds the specified anomaly threshold, the verdict for the request may change. The logs do not show the model revision, but changes in score and verdicts can indicate changes in model behavior. A threshold of 90 guarantees stable verdicts. Start with this value, monitor rule matches, and adjust the settings as needed.
To ensure your infrastructure is always protected:
- Enable detailed logging.
- Set baseline protection metrics.
- Narrow down the change scope.
- Adjust the protection scope to the minimum necessary.
- Assess the impact using security and availability metrics.
- When to contact support.
Enable detailed logging
Configure logging via Smart Web Security and enable logging for:
- Requests with the
DENYandCAPTCHAverdicts. - Percentage of requests with the
ALLOWaction. For allowed traffic, you can use a sampling rate from1to100percent. The higher the percentage, the more information you get, but the more logs are generated.
Use the following log fields to analyze ML WAF rule matches:
actionanddry_run_matched_rule_verdict: Final and dry-run verdict for the request.waf_applied_rule_set_id: Rule set that made the verdict.waf_matched_rulesanddry_run_waf_matched_rules: WAF rules that matched, including those in dry-run mode.rule_id,rule_set_id, andrule_group_id: Rule, rule set, and rule group IDs.score: Anomaly score for the request.matched_data_variable,matched_data_key, andmatched_data_value: Part of the request containing the anomaly.waf_matched_exclusion_rules: Exclusion rules that matched.
Set baseline protection metrics
Before changing the ML WAF settings, set the following metrics separately for each attack group:
- Number and percentage of requests with the
ALLOW,DENY, andCAPTCHAverdicts. - Distribution of
scorevalues. - Top routes, request parameters, and request parts triggering rule matches.
- Percentage of
4xxand5xxresponses, service availability, and business metrics for critical scenarios. - Current WAF profile configuration, including the rule set ID and version.
Baseline metrics help distinguish the impact of a model update from seasonal traffic fluctuations, changes in traffic composition, and your own configuration changes.
Narrow down the change scope
When a false positive is detected:
- Use the application logs to confirm that the request is legitimate.
- Identify the specific ML WAF attack group, rule, route, and request part for which the rule matched.
- Temporarily increase the anomaly threshold for the affected group or switch the rule back to Only logging mode.
- Once the false positive is confirmed, create an exclusion rule.
- Enable logging for the exclusion and check the logs to make sure it only applies to expected traffic.
Warning
Do not disable the entire WAF profile or exclude the entire request if the false positive can be isolated to a single rule, route, or request field. A broad exclusion can create an uncontrolled protection bypass.
Narrow down your exclusion rule by combining multiple criteria:
- Specific WAF rule.
- Route, host, or HTTP method.
- Specific part of the request: HTTP request body, cookie, HTTP header, or query string parameters.
- Specific parameter or header, if known.
Adjust the protection scope to the minimum necessary
Start with the recommended anomaly threshold of 90 and enable attack groups one at a time.
Do not combine the following in a single change:
- Enabling a new attack group.
- Lowering the anomaly threshold.
- Expanding the profile scope.
When lowering the anomaly threshold, consider the risk of false positives: the lower the threshold, the higher the protection sensitivity.
Assess the impact using security and availability metrics
After each change, monitor security and availability metrics.
Security metrics:
- Number and breakdown of ML WAF rule matches.
- Distribution of
scorevalues. - Suspicious traffic and incidents.
- Changes in coverage across attack groups.
Availability and business metrics:
- Percentage of
4xxand5xxresponses. - Authorization and payment errors.
- Successful API requests and integrations.
- User support requests.
- Conversion rates for critical user scenarios.
A lower percentage of blocked requests does not always mean better protection. It may be due either to fewer false positives or more attacks getting through.
When to contact support
Contact Yandex Cloud support
- Verdicts started changing at scale without any configuration changes on your side.
- The issue manifested simultaneously on unrelated routes.
- The matches cannot be isolated to a single rule or request field.
- Mitigating the issue safely requires a broad exclusion or disabling ML WAF.
- A business metric changed, but the logs do not provide enough information to associate it with a specific rule or request.
- The timing of the change coincided with an ML WAF model or rule set update.
Specify the following in your support ticket:
- Change start time and time zone.
security_profile_idandwaf_profile_id.- Rule set ID and version (
ruleSet.idandruleSet.version). rule_idandrule_group_idin question.- Several
alb_request_idandunique_keyvalues from the logs. - Aggregated comparison of metrics before and after the change.
- Actions taken so far.