sk184381 - Connectivity issues between Security Group members in a Maestro Dual-Site configuration with a non-direct site-sync connection

Connectivity issues between Security Group members in a Maestro Dual-Site configuration with a non-direct site-sync connection

Product: Maestro HyperScale Firewall
Version: R81.20, R82, R82.10
OS: Gaia
Platform: Maestro Orchestrator
Last Modified: 2026-07-23

Symptoms

Cause

Consider a Maestro Dual-Site deployment where the Site-Sync connection between the two sites through intermediate Layer-2 switches (see sk168092).
SGM 1

MHO 1_1

MHO 2_1

MHO 1_2

MHO 2_2

Sync-ext

Sync-int

Sync-ext

Sync-int

Sync-ext

Sync-ext

Sync-int

Sync-int

Switch 3

Switch 4

SGM 2

Switch 1

Switch 2

If the connection between the switches goes down, then the Sync-ext interfaces continue to report that the link is up, but the actual path between remote peers is broken. For example, if the link between Switch 3 and Switch 4 fails, MHOs 1_2, 2_2 do not detect it.

SGM 1

MHO 1_1

MHO 2_1

MHO 1_2

MHO 2_2

Sync-ext

Sync-int

Sync-ext

Sync-int

Sync-ext

Sync-ext

Sync-int

Sync-int

Switch 3

Switch 4

SGM 2

Switch 1

Switch 2

Would this link fail:

Sync-ext on 1_2, 2_2 will be seen as up

In the scenario above, the entire deployment is not able to use the faulty link but continues to try to use it as if it were valid.

Solution

This problem was fixed. The fix is included in:

If you choose not to upgrade, Check Point can supply a Hotfix. Contact Check Point Support to get a Hotfix for this issue.
A Support Engineer will make sure the Hotfix is compatible with your environment before providing the Hotfix.
For faster resolution and verification, please collect CPinfo files from the Security Management Server and Security Gateways involved in the case.
Hotfix installation instructions:
Refer to sk168597 - How to install a Hotfix.

How Site-Sync Monitoring mitigates the issue

Each MHO runs the SSM_PMD daemon, which performs two checks:

Detects Link Aggregation Group (LAG) or interface link-down events on the Site-Sync interface

When Site-Sync Monitoring is enabled, the daemon periodically sends Internet Control Message Protocol (ICMP) pings to the remote Maestro Hyperscale Orchestrator (MHO) through the Site-Sync interface.

If the ping test fails to meet the configured success threshold:

Enable Site-Sync Monitoring

Enable Site-Sync Monitoring on (MHOs that do not have a direct Site-Sync connection (typically the “parallel” MHOs):
set maestro configuration security-appliances inter-site monitor state enabled
Note - Enabling this on one MHO automatically enables it on the corresponding parallel MHO in the other site. For example, enabling it on MHO 1_2 also enables it on MHO 2_2, ensuring symmetrical monitoring.

Clish Parameters

In the Clish configuration, aside from state, you can change the following ping settings. The system uses these parameters to determine link health. If the number of successful pings falls below the configured threshold, the link is marked as unreachable.

Parameter Description Default
ping-timeout Timeout (in seconds) to wait for each ping reply 1
ping-count Number of ping packets sent over the Sync-ext interface during each check 2
successful-pings-required Minimum number of successful pings (from ping-count) required for the link to be considered active 1

Logger Settings

Parameter Description Default
interval-length-for-flapping Time (in minutes) during which link state changes are counted. 2
max-flaps Maximum number of link state changes allowed in the interval before the link is considered to be flapping 3
log-timeout-if-flapping Time (in minutes) to suppress repetitive log messages after flapping is detected 30

To prevent excessive logging: when the system detects excessive state changes, and considers the link as flapping (exceeds the value of max_flaps) in the specified interval, then the system:

  1. Logs a single warning in /var/log/messages
  2. Suppresses further logs for the duration defined in log_timeout_if_flapping_min.

FAQs

NOTE

This solution has been verified for the specific scenario, described by the combination of Product, Version and Symptoms. It may not work in other scenarios.