Create a Post
cancel
Showing results for 
Search instead for 
Did you mean: 
Don_Paterson
MVP Gold
MVP Gold

Traditional VSX cluster restore fails - How to do a successful restore using vsx_util reconfigure

Virtual Systems missing on the VSX cluster members after restoring a Traditional VSX Cluster from Gaia backups, and the successful recovery using vsx_util reconfigure.

This scenario was documented to make engineers aware and hopefully help someone in the same situation or just to learn more about the Traditional VSX product. 

Summary

After a restore of the Security Management Server and of both VSX cluster members from Gaia backups, in the order of sk100395, the Management Server database restores the VSX cluster and cluster members, the Virtual Systems and the Virtual Switch.

Each member reports Local restore succeeded. but shows no Virtual Devices - vsx stat -v prints No Virtual Devices.

The configuration files that describe the Virtual Systems are present on the members. The members came back with their Virtual Systems after vsx_util reconfigure was run on the Management Server for each member (including a SIC reset on that member and a reboot). sk100395 does not document this outcome or these extra procedures that were required to complete a successful restore.

Lab

  • Security Management Server, R82.20
  • VSX-01 and VSX-02: VSX cluster VSX-Cluster, Virtual System Load Sharing. Version R82.20
  • Virtual Devices: Virtual Switch VSW, Virtual Systems DMZ-GW and INT-GW, one policy package for each Virtual System (DMZ-Policy, Internal-Policy), and the package VSX-Cluster_VSX for Virtual System 0 (vs0).
  • Scenario:
    • Back up the cluster and the Management Server, manually remove all VSX objects from the database and from the members (the teardown), and delete the VSX Cluster object, then attempt to restore from Gaia backups.
      After the teardown the members run as VSX gateways with Virtual System 0 only.

 

Topology example:

Check Point VSX Topology Diagram.png

Backup

Following the order of sk100395 - Back up the VSX cluster members first, then the Management Server, using  Gaia Clish on each machine:

add backup local
show backup status
show backups


add backup local was run on all three machines.

The .tgz file appears in /var/log/CPbackup/backups

show backup status output: Backup process finished and Local backup succeeded

Each file was copied to a separate Windows host and the MD5 of the copy was compared with md5sum on the machine (equal).

Restore (sk100395 order)

  1. Close SmartConsole. On SMS, in Gaia Clish:
set backup restore local backup_SMS_2026_10_04_17_16_21.tgz
show restore status

 

The command asks no question and the machine restarts by itself. SSH was available again after about four minutes.

show restore status output: Local restore succeeded.

api status showed CPM running and ready.

  1. On VSX-02, then on VSX-01, in Gaia Clish:
set backup restore local backup_VSX-02_2026_10_04_17_16_01.tgz
show restore status

 

Both members output: During backup validation - serval warning have come up: and The current machine doesn't contain the following interfaces: br1/wrp128/wrp192/wrpj128/wrpj192.

The restore ended with Local restore succeeded. within a minute, and the member was reachable again about three to four minutes after its SSH session ended.

 

Findings after the restore

SMS database (mgmt_cli -r true show gateways-and-servers)

All of the following objects and the policy packages are in the SmartConsole: VSX-Cluster, VSX-01, VSX-02, VSW, DMZ-GW, INT-GW, A-SME and SMS. The package VSX-Cluster_VSX.

Members, vsx stat -v

SIC Trust, Virtual Systems [active / configured] 1 / 1, Virtual Routers and Switches 0 / 0, No Virtual Devices.

Members, cphaprob stat

Virtual System Load Sharing, VSX-01 ACTIVE, VSX-02 STANDBY, no PNOTEs.

Members, $FWDIR/state/local/VSX/

local.vsall, local.vs, local.vskeep, local.cmdset, local.sic_name, policy.info and policy.map are present; local.vs lists the creation of VSW, DMZ-GW and INT-GW on both members.

Members, /opt/CPsuite-R82.20/fw1/CTX/

Empty.

vsx_push_configuration_after_reboot.elg on VSX-02

Shows a push at the restart in which vs destroy vsid 1 fails ("Virtual System cannot be removed", "failed to run cpstop for VS (1)").

All of the backups were taken together and restored in the order of sk100395, and the members still came back without Virtual Systems.

An attempt to push the Virtual Switch configuration from SmartConsole (Edit VSW and click OK in the VSW window) failed on both members with "Failed to create Virtual System directories".

 

Solution

Summary: SIC reset, vsx_util reconfigure, and then reboot. One member at a time.

 

Run these steps for VSX-01, then for VSX-02.

  1. Reset SIC on the member. SSH to the member as admin, Expert mode:
cpconfig

 

Enter:

5 (Secure Internal Communication)

y to re-initialize

y to confirm

Enter the one-time activation key twice,

11 (Exit).

 

cp_conf sic state then prints Trust State: Initialized but Trust was not established

vsx stat -v prints SIC Status No Trust.

  1. Reconfigure the member from SMS. In Expert mode on SMS:
vsx_util reconfigure

 

The command prompts, in order:

  • The Management Server address (Enter for localhost)
  • The administrator name (cpadmin, the Management administrator, not the Gaia user)
  • The administrator password
  • The VSX cluster object (VSX-Cluster);
  • The member to reconfigure
  • Are you sure you want to continue [y/n]? (y)
  • The activation key twice (the key typed in step 1)

 

NOTE:

In the first test the IPv4 CoreXL instance number needed to be typed in (use cpview and check the CPU section for CoreXL instance count). In the second test this last question was not asked.

 

  • Ten steps are then carried out automatically by the vsx_util reconfigure command.
    • Certificate Revocation
    • Certificate Replacement
    • Connectivity Check
    • Fetching Configuration
    • Verifying Configuration,
    • Installing policy on VSX-Cluster
    • Converting Gateway to VSX
    • Generating Activation Keys
    • Reconfiguring,
    • Pushing Configuration.

 

The command took about five and a half minutes for VSX-01 and about five minutes for VSX-02 and ended with Reconfigure gateway operation completed successfully and IMPORTANT: Please reboot the gateway.

  1. Reboot the member (reboot).
    vsx stat -v first shows the Virtual Devices as Unknown or fewer than three active; about one to one and a half minutes after that it shows 3 / 3 and SIC Trust for VSW, DMZ-GW and INT-GW.
  2. Install the policies (SmartConsole, or the Management API):
mgmt_cli -r true install-policy policy-package DMZ-Policy access true threat-prevention false targets.1 DMZ-GW

mgmt_cli -r true install-policy policy-package Internal-Policy access true threat-prevention false targets.1 INT-GW

 

Result after both members (second test):

vsx stat -v on each member lists VSW (Virtual Switch), DMZ-GW (DMZ-Policy) and INT-GW (Internal-Policy) with SIC Trust, Virtual Systems 3 / 3, Virtual Routers and Switches 1 / 1

cphaprob stat shows Virtual System Load Sharing, VSX-01 ACTIVE and VSX-02 STANDBY with no PNOTEs

Both policy installations succeeded.

 

Questions for Check Point

  1. Is it expected that a Gaia restore of a VSX cluster member, taken with the Virtual Systems running, returns the member with Virtual System 0 only?
  2. Is vsx_util reconfigure (with a SIC reset) the supported way to bring the Virtual Systems back, and is this documented?

Sources

  • sk100395: How to back up and restore VSX gateway.
  • R82.20 Gaia Administration Guide: "Backing Up and Restoring the System".
  • R82.20 VSX Administration Guide, Command Line Reference: vsx_util reconfigure, vsx fetch.
0 Kudos
7 Replies
Chris_Atkinson
MVP Diamond CHKP MVP Diamond CHKP
MVP Diamond CHKP

Your third question is empty / missing?

Parking version applicabiltiy for the moment the following is an older article that speaks to some of this.

sk101515 - How to Reconfigure a VSX Cluster member 

CCSM R77/R80/ELITE
0 Kudos
Don_Paterson
MVP Gold
MVP Gold

Thanks Chris. 

I removed the 3rd empty question. That was just a Word thing. 

The SK has relevance but would need to state that the procedure (reconfigure specifically) was also for a failed restore attempt (under When to use this procedure) but the question about a restore failing for the product still stands. 

All recommended backup procedures were followed as documented. 

0 Kudos
emmap
MVP Gold CHKP MVP Gold CHKP
MVP Gold CHKP

Virtual Devices should be restored with the backup restore operation, as I understand it. Did the gateways reboot after the restore, or did you manually reboot them? Due to the VSs having been removed, a reboot may have been necessary to complete the restore.

0 Kudos
Don_Paterson
MVP Gold
MVP Gold

They should indeed have been, but were not.

"The restore ended with Local restore succeeded. within a minute, and the member was reachable again about three to four minutes after its SSH session ended."

The Gaia restore automatically reboots the machine/s. I left out some details in my breakdown but I did check that the gateways rebooted as part of the restore (I checked uptime (9 minutes)) before running vsx stat -v.

If a second reboot is required then that is questionable behaviour and is also not documented. I did not try that.

In my first test, VSX-02 was also reverted to a snapshot, which restarts the machine. It ended with no Virtual Devices as well.

My guess would be a VSX configuration push (Network configuration script) failed and there is no recovery programmed in.

 

This is a VM environment MS Hyper-V Windows 2022, which is supported.

https://sc1.checkpoint.com/documents/R82.20/WebAdminGuides/EN/CP_R82.20_RN/Content/Topics-RN/Support...

 

VSNext and Traditional VSX

This table shows the support for VSNext and Traditional VSX

in R82.20:

Platforms

VSNext

Traditional VSX

ElasticXL Cluster

Yes (3)

No

Security Group - Maestro

Yes (4)

Yes (5)

Open Servers (1)

Yes

Yes

Virtual Machines (2)

Yes (3)(6)(7)

Yes (7)

 

 

  1. VSNext and Traditional VSX modes are not supported in Public Cloud.

0 Kudos
emmap
MVP Gold CHKP MVP Gold CHKP
MVP Gold CHKP

OK yea the automatic reboot should have sufficed. This might be better investigated via TAC so that it can be replicated and worked on that way.

0 Kudos
JozkoMrkvicka
Authority
Authority

You have mentioned that coreXL settings were asked. Where ? How ? VS0 doesnt have coreXL enabled and each VS has dedicated and configurable number of IPv4 and IPv6 coreXL in VS object.

What is the size of backup files from both VSX members ? Does the backup file contain relevant files for each VS (in CTX directory) ?

I have feeling that the backup was done only for VS0 and thus no VSs were restored.

Anyway, R82.20 doesnt have any JHF yet and I would consider this version as buggy.

I would try the same restore aproach with all machines running R82 or R82.10 with their latest JHFs.

Kind regards,
Jozko Mrkvicka
0 Kudos
Don_Paterson
MVP Gold
MVP Gold

"Where ?" During vsx_util reconfigure

"How ?" Via CLI during vsx_util reconfigure - after SIC OTP - see attached

 

"What is the size of backup files from both VSX members ?"

  • VSX-01: 156,459,692 bytes (2026-10-02) and 153,836,191 bytes (2026-10-04).
  • VSX-02: 156,541,103 bytes (2026-10-02) and 153,829,191 bytes (2026-10-04).

 

"Does the backup file contain relevant files for each VS (in CTX directory) ?" Yes. Both members' backups contain data for each Virtual System.

CTX contents. All four archives have CTX00001, CTX00002 and CTX00003, on both members.

  • Virtual Systems: under var/opt/CPsuite-R82.20/fw1/CTX/ each has about 1,640 entries (about 26 MB). CTX00001 has about 230 entries (1.3 MB).
  • Other paths: the same three directories also appear under opt/CPsuite-R82.20/fw1, opt/CPshrd-R82.20, var/opt/CPshrd-R82.20 and var/log/opt/CPsuite-R82.20/fw1.

There are a lot of symlinks and hard links pointing back to the VS0 files.

The data is in the backup, but after the restore both CTX paths on the members were empty, and vsx stat -v showed no Virtual Devices.

 

"Anyway, R82.20 doesnt have any JHF yet and I would consider this version as buggy." I would expect a restore to work at least.

"I would try the same restore aproach with all machines running R82 or R82.10 with their latest JHFs." Please share the results with us here.

 

The main point here is that a backup and  restore that is carried out using to the recommended procedures fails, repeatedly and consistently it seems.

Thankfully (and strangely thanks to the complexity of Traditional VSX management and provisioning) there is a restore workaround in the vsx_util reconfigure procedure.

 

0 Kudos

Leaderboard

Epsum factorial non deposit quid pro quo hic escorol.

Upcoming Events

    CheckMates Events