Create a Post
cancel
Showing results for 
Search instead for 
Did you mean: 
Richard_Orton
Participant
Participant

Rebuilding 2 Gateways on ESXi - no layer 2 - ESXi arp cache?

Hi All

Have a pair of firewalls in a cluster running locally on ESXi. Running R81.10 and want to upgrade to R82

There are no spare IP addresses in the subnet.

We cant do an in place upgrade due to error:

The following results are not compatible with the package:

- Your partitions are not in Check Point standard format, and an upgrade is not possible. Reinstall your system and verify correct partitioning.

The problem we have, is when we do a clean build of the standby device with the primary running, we cant get basic layer 2 connectivity between the primary and the standby.

If this was a physical switched based network, I would be clearing the arp cache on the switch, but as this is ESXi, I don’t even know if there is such a thing – virtual switch on the ESXi host, does it have an arp cache and can it clear?

Running arp -a on the primary came back with incomplete for the standby device IP

We had TAC on a call while we attempted a rebuild of the standby device, and they were as stumped as we were at the basic lack of connectivity. What we did notice was that approx. 4 hours after we started the change, we started to see packets sent and received on the newly built standby member, but had to regress at that point. This also felt like a 4 hour arp cache timeout.  If we restored the standby vm and powered it up, it came up fine, connection to the primary fine, and worked as a cluster.

 

Anyone seen anything similar in ESXi, is there an arp cache on the virtual switches that sit on the ESX host?

Thanks

0 Kudos
9 Replies
Chris_Atkinson
MVP Platinum CHKP MVP Platinum CHKP
MVP Platinum CHKP

What NIC type was used?

Assume all the prerequisites from sk101214 are in place?

CCSM R77/R80/ELITE
Richard_Orton
Participant
Participant

I need to check on the nic, im pretty sure was (and current running ones) are vmxnet3

Thanks for the sk, will digest that one and compare

0 Kudos
Chris_Atkinson
MVP Platinum CHKP MVP Platinum CHKP
MVP Platinum CHKP

Were you able to identify a likely cause as yet?

CCSM R77/R80/ELITE
0 Kudos
Richard_Orton
Participant
Participant

A colleague who manages the ESXI infrastructure has noted that from the SK, mac learning is disabled.

We are looking to get a window to reset and see if this resolves, thanks for everyone's input so far, will post back with any updates but may be a few weeks due to change freezes etc. 

Bob_Zimmerman
MVP Gold
MVP Gold

Are you installing the new OS in the same VM, or are you building a new VM and trying to swap it into place?

Check your interface ordering. It's possible the newer OS brings them up in a different order, leading to the interfaces not being connected how you expect (that is, eth0 on your old VM might be eth3 on your new VM). The PCIe addresses should be consistent, so you should be able to confirm this using 'ethtool -i eth0' for each interface on each member.

Richard_Orton
Participant
Participant

We actually tried both, the install into the existing VM didnt go to plan due to some discrepancy with the BIOS and couldn't get the ISO to boot, so eventually built a new VM and this is where we got to - we did make sure that when manually allocating the interfaces in ESXi that they were allocated in the same order, and mac address checks confirmed that each interface inside the OS matched what we expected the mac to be (so was happy eth1 was eth1 etc)  

0 Kudos
Bob_Zimmerman
MVP Gold
MVP Gold

If you're using VMware NSX, a new VM doesn't get all the same rules as the existing VM. They're unique objects. Doesn't seem likely to be your problem, but worth confirming.

The MAC learning thing seems like the most promising path.

0 Kudos
Duane_Toler
MVP Silver
MVP Silver

Are you attaching the new cluster member to the same portgroup ?  This is your layer 2 service on ESXi (well, it's the top half of layer 2). This is where your VLAN (port group ID) is set.  The port group attaches the VLAN to the vSwitch (the bottom half of layer 2).  The vSwitch attaches to the physical NICs on the host.

If you're missing ARP entries for all hosts on the LAN (not just the cluster peer), then you have a port group configuration problem.  If you're configuring a new port group, then make sure it matches the settings of the known-working port group.  Lots of options in here (check Forged Trasnmit, too).

On your uplinking switch, you can also check to see if the LAN switch is learning the MAC addresses in the spanning tree for your VLAN and on the intended ports (show mac addr vlan X,  or show mac addr int ethernetX/0/Y, or whatever).

In the OS of the gateway, make sure you have the interface up and active:  set interface ethX state on; I got burned on that once myself, too.  Simple fix, thankfully.

Check these and let us know how it goes.

 

--
Ansible for Check Point APIs series: https://www.youtube.com/@EdgeCaseScenario and Substack
0 Kudos
emmap
MVP Gold CHKP MVP Gold CHKP
MVP Gold CHKP

I have had some issues with new VMs not being able to talk out, but usually if I ping out from the new VM the various ARP mechanisms update and it starts working. If you're using ESX vswitches, sometimes putting them into promiscuous mode can help?

Leaderboard

Epsum factorial non deposit quid pro quo hic escorol.

Upcoming Events

    CheckMates Events