Create a Post
cancel
Showing results for 
Search instead for 
Did you mean: 
HeikoAnkenbrand
MVP Diamond
MVP Diamond
Jump to solution

Converting Tool - ClusterXL to ElasticXL

I’ve opened a new article on this topic.

In the near future, there is supposed to be a tool that allows converting a ClusterXL into an ElasticXL cluster.

This is discussed in the following article: 
Replacing 5800 HA ClusterXL with 9200s. Should I convert to ElasticXL?

What information is available about this tool and when will it be released?

 

Can you already provide any further information about this tool?This would be very relevant for planning with our customers.

➜ CCSM Elite, CCME, CCTE, CCVS ➜ www.checkpoint.tips
40 Replies
Alex-
MVP Silver
MVP Silver

Unless the factory default is changed, you have a potential issue with your ElasticXL cluster.

Removing a cluster member triggers a factory default. If the image is anterior to R82, you will need to somehow find a way to format it and if it's on the other side of the country without LOM, you're in for a treat.

I understand that coming gateways will be shipped with R82 by default, but there are very likely a lot of existing gateways which went to R82 from previous versions with Blink and wrongly assume they can just use a conversion tool to do all the work for them.

0 Kudos
Tobi
Participant
Participant

Will the migration tool trigger the factory default or/and if you remove the member from the existing ElasticXL cluster?

Thats why I would like to test it in the lab, so see what the challenges are.

0 Kudos
Alex-
MVP Silver
MVP Silver

Changing the factory image is one of the sequences of ISOMorphic.

I don't believe a tool on a running unit could do that.

0 Kudos
Bob_Zimmerman
MVP Gold
MVP Gold

What Check Point calls the "factory default" image is just a logical volume in the same volume set as everything else. It could be replaced on a live system, but Check Point doesn't provide tools which do this. I think this is reasonable, since the point of the image is to have a reliable way to restore the box to functionality with minimal access needed. By only replacing the image when you have hands on the box (or LOM access), you're guaranteed to be able to do it again if something goes wrong.

Incidentally, it's also possible to totally rebuild the drives of a live system. You have to free enough RAM to load a sizable RAMdisk image, stop most services, mount the image, maybe hand execution over to a kernel image in the RAMdisk, unmount the drives, repartition the drives, write everything you want to the drives, then reboot. It's not especially complicated (especially compared to live patching of ROP-resistant executables, for instance), but if anything goes wrong, recovery tends to require hands on the box.

Apple did something like this in iOS 10.1, 10.2, and 10.3 with the move to APFS. Upgrading to 10.1 and 10.2 ran a conversion of the live filesystem on the device from HFS+ to APFS to find and log any errors, then it left the original HFS+ filesystem tree in place. Upgrading to 10.3 actually removed the HFS+ filesystem tree, leaving only the APFS tree. They only rewrote the filesystem itself rather than the data indexed by it, but there's no fundamental reason they couldn't have also rewritten the data if the device had enough RAM.

emmap
MVP Gold CHKP MVP Gold CHKP
MVP Gold CHKP

Removing an SGM doesn't do a full FCD to the stored FCD image, it's a 'light' FCD that just cleans the config. So you don't end up downgrading versions if you have in-place ugpraded it before.

0 Kudos
Alex-
MVP Silver
MVP Silver

I went to look and the story is a bit different indeed. The box was fresh installed with R82 from R81.20 and then joined in the cluster.

Somehow, it didn't pick up even after membership removal so we had to do an FCD, which was then R81.20 because of the fresh install and not ISOMorphic. So updating the FCD image still remains a time saver in such situations.

Thanks for the clarification.

0 Kudos
Don_Paterson
MVP Gold
MVP Gold

The migration script is probably at fault and not handling the blink image naming as it needs to. 

 

The probablen root cause is the version-parsing routine assumes every relevant package's key ends in an integer take number, and a Blink image's package key doesn't.

It's not that Blink upgrades are unsupported in principle — it's that the tool's take-number parser wasn't written to tolerate the Blink package naming.

The "upgraded from R81.10 via Blink" detail is the trigger, but not because of the R81.10 origin per se.

It is likely because the Blink packaging leaves a package record (BLINK_R82_T779_JHF_T107_GW) in the inventory that init_version walks over and tries to int().

Their JHF107 is fine; the take-number check (validate_cluster_mems_vers) would even pass if it got that far.

The failure is upstream of validation, in the initial data collection, which is why it dies at "Collecting data about the migrated object" rather than at a named validation gate.

What should resolve it, in rough order of preference:

The clean fix is Check Point's responsibility.

The parser needs to skip or safely handle package keys that don't end in an integer (a try/except ValueError around that int(), or filtering to only category: jumbo/hotfix packages before parsing).

It might take a TAC case but hopefully they'll pick up this feedback and test.

Given T107 is current and Blink upgrades are ecommon, this could affect a lot of people, so it should be fixed. 

The practical workaround you can test in the lab: get the gateways onto a package inventory that doesn't carry the Blink wrapper package — i.e. a clean R82 install plus Jumbo applied the normal (non-Blink) way, rather than a Blink image. On a lab that's a rebuild; on production it's obviously not casual. But it confirms the diagnosis and unblocks them.

Checking exactly what's in the inventory:

da_cli packages_info status=installed | grep -i packageKey

 

In my fist lab test it failed and I needed to patch the gateways. 

After the JHFA the script failed again and reported mismatched JHFA but they were the same. I just had to wait and try again...

genisis__
MVP Silver
MVP Silver

Don,
Good to hear, but not sure I would want to do this in live client; what's the backout? i.e snapshot everything including the manager and the roll the whole lot back, or is there a neater way to do this?

 

0 Kudos
Don_Paterson
MVP Gold
MVP Gold

 

Go with whatever the SK advises (yes, snapshot first), and pay attention to the script output. 

https://support.checkpoint.com/results/sk/sk183894

0 Kudos
_Val_
Admin
Admin

You only need one production IP address per clustered interface, nothing else. Play with it in the lab, see the demos, ask your partner to get you access to the ElasticXL lab if you can. It is a new way to build a cluster.

0 Kudos
Don_Paterson
MVP Gold
MVP Gold

I have created a new post to ask about hardware refresh in the migration scenario.

https://community.checkpoint.com/t5/Firewall-and-Security-Management/The-Scalable-Platform-Migration...

 

0 Kudos

Leaderboard

Epsum factorial non deposit quid pro quo hic escorol.

Upcoming Events

    CheckMates Events