sk181891 - The 'cxld' process consumes the CPU at 70%-100% on VSX Cluster Members

The 'cxld' process consumes the CPU at 70%-100% on VSX Cluster Members

Product: VSX (Traditional)
Version: R81 (EOS), R81.10 (EOS), R81.20
OS: Gaia
Last Modified: 2024-06-05

Symptoms

This CPU utilization is observed in outputs of various commands ('top', 'ps', 'cpview' > CPU > Spikes).

Example from CPview:

Cause

The CXLD processes that run in the context of Virtual Systems write and read the same internal file on the VSX Cluster Member, instead of doing so with a file in the context of each corresponding Virtual System.

Solution

This problem was fixed. The fix is included in:

* As an immediate workaround, follow this procedure:

(Even if you choose to upgrade you are still must to applied this procedure on all members)

  1. Connect to the command line on each VSX Cluster Member.

  2. Log in to the Expert mode.

  3. Examine the state of the Virtual Systems:

    cphaprob state

    For example, there are two VSX Cluster Member:

    • VSX_M_1 - ACTIVE cluster state
    • VSX_M_2 - STANDBY cluster state
  4. On the 'STANDBY' VSX Cluster Member (in our example, VSX_M_2), perform these steps for each Virtual System in the 'STANDBY' state (write down each VS ID):

    1. Go to the context of the Virtual System:

      vsenv <VS ID>

    2. Disable the CPU utilization monitor:

      fw ctl set -f int fwha_cpu_utlization_monitor_enable 0

    3. Identify the Process ID (PID) of the 'cxld' process for this Virtual System:

      cpwd_admin list | grep -E "PID|CXLD"

      Note - The "CTX" column shows the VS ID.

      Example output:

      [Expert@VSX_M_2:0]# cpwd_admin list | grep -E "PID|CXLD"
      APP        CTX        PID    STAT  #START  START_TIME             MON  COMMAND
      CXLD       0          44825  E     1       [12:41:56] 3/1/2024    N    cxld -d
      CXLD       6          10765  E     1       [12:44:32] 3/1/2024    N    cxld -d
      CXLD       12         45470  E     1       [10:52:53] 8/1/2024    N    cxld -d
      CXLD       7          11152  E     1       [12:44:32] 3/1/2024    N    cxld -d
      CXLD       11         11153  E     1       [12:44:32] 3/1/2024    N    cxld -d
      CXLD       4          11154  E     1       [12:44:32] 3/1/2024    N    cxld -d
      CXLD       5          11155  E     1       [12:44:32] 3/1/2024    N    cxld -d
      CXLD       10         11156  E     1       [12:44:32] 3/1/2024    N    cxld -d
      CXLD       2          11157  E     1       [12:44:32] 3/1/2024    N    cxld -d
      CXLD       8          11164  E     1       [12:44:32] 3/1/2024    N    cxld -d
      CXLD       9          11405  E     1       [12:44:33] 3/1/2024    N    cxld -d
      CXLD       3          11407  E     1       [12:44:33] 3/1/2024    N    cxld -d
      [Expert@VSX_M_2:0]#
      

      Example PIDs:

      • PID of CXLD in VS 2 = 11157
      • PID of CXLD in VS 9 = 11405
    4. Terminate the 'cxld' process for this Virtual System:

      kill -9 <PID of CXLD>

    5. Monitor the 'cxld' process for this Virtual System - wait until is starts again:

      watch -d -n 5 'cpwd_admin list | grep -E "PID|CXLD" | grep -E "PID| <VSID> "'

      Notes:

      • The inner command is enclosed in single quotes.
      • There are 3 space characters in front of the VS ID and after the VS ID.

      In the , substitute the required number.

      • The "PID" column must show a non-zero integer.
      • The "STAT" column must show "E".

Example output for VS ID 9 after the 'cxld' process termination:

Every 2.0s: cpwd_admin list | grep -E "PID|CXLD" | grep -E "PID|   9   "

APP        CTX        PID    STAT  #START  START_TIME             MON  COMMAND
CXLD       9          0      T     1       [12:44:33] 3/1/2024    N    cxld -d

Example output for VS ID 9 after the 'cxld' process starts again:

Every 2.0s: cpwd_admin list | grep -E "PID|CXLD" | grep -E "PID|   9   "

APP        CTX        PID    STAT  #START  START_TIME             MON  COMMAND
CXLD       9          49470  E     2       [18:16:00] 17/1/2024   N    cxld -d
  1. Monitor the VSX cluster state - wait for the state of this Virtual System to change to 'STANDBY' again:
  `cphaprob state`
  
  Note - In the output, refer to the bottom section "`Virtual Devices Status on each Cluster Member`".  
  1. Administratively change the cluster state of the "STANDBY" VSX Cluster Member to "DOWN" (in our example, VSX_M_2):

    clusterXL_admin down

  2. On the "ACTIVE" VSX Cluster Member (in our example, VSX_M_1), perform Step 4 for each Virtual System that you changed on the previous VSX Cluster Member (in our example, VSX_M_2).

  3. Administratively change the cluster state of the "DOWN" VSX Cluster Member to "UP" (in our example, VSX_M_2):

    clusterXL_admin up

  4. Examine the state of each VSX Cluster Member and each Virtual System:

    cphaprob state

NOTE

This solution has been verified for the specific scenario, described by the combination of Product, Version and Symptoms. It may not work in other scenarios.