Create a Post
cancel
Showing results for 
Search instead for 
Did you mean: 
afiqmohdyusof
Explorer

Running VSX on cluster with different appliance

Greetings,

We are planning to replace a cluster of VSX systems currently operating on a 23800 appliance with a 9700 appliance. Both the 23800 and 9700 will be running on R82.


Our constraint is that the customer requires zero downtime and no interruption for all the VS systems during this technical refresh.


From our testing, we successfully onboarded one 9700 appliance into the cluster. We encountered no issues during policy installation. However, we are experiencing difficulties when running 'vsx_util reconfiguration' for the 9700. We encountered the following error:


"Internal Error - Failed to commit changes in the OS... Unable to obtain information regarding interface eth2-01
Aborting reconfiguration ...

On the 23800 system, we are using interfaces labeled eth2-XX, while on the 9700 system, we are using interfaces labeled ethX, as they have different interface naming conventions. May I inquire if it is possible to modify the interface name on the 9700 to match that of the 23800? 

0 Kudos
6 Replies
_Val_
Admin
Admin

Sorry to rain on your parade, but "vsx_util reconfigure" approach will not work, and Zero Downtime will not work either.

You need to rebuild the full cluster while keeping the old one in production. You will need to manually build a new cluster from scratch and replicate the exact topology for a smooth switchover.

I did exactly that more than once with my customers. All you need is to isolate the new cluster on the network side, either physically or logically, to avoid any duplicate IP/route issues, configure it to mimic the old VSX topology, push policy on it, and then perform the network swap procedure to cut off the old cluster and send traffic to the new one. 

Short downtime is expected, but if you don't make any mistakes, traffic should be reestablished through the new cluster. 

This also gives you a clear rollback option: if something goes wrong, you can reconnect the old cluster and isolate the new one again.

Before you go this way, make sure you freeze changes and make an actionable backup of your management server.

Bob_Zimmerman
MVP Gold
MVP Gold

There is not a way to do this with no downtime. Flatly not possible.

Building a whole new cluster as Val suggested is probably better. Just make parallel VSs and push the same policies to them. Policy push goes through VS0, so as long as you aren't using VS0 to pass traffic, it should work with separate "management addresses" on the new cluster. Build it with bonds. This approach lets you cut over one VS at a time by adding the only the relevant VLANs to the trunks.

If you don't want to go that route for whatever reason, I would deal with the interface naming problem first. With VSX, the firewall application cares very deeply about the exact interface names, so you should only let it know about bonds. It deals with exactly this interface naming problem, since the OS can back the bond with whatever physical interface name(s) you want. A bond with one interface doesn't need any special support from the switch side. You can use 'vsx_util change_interfaces' on the management to deal with this. It moves all references to one physical interface to another physical interface. For example, you can move everything from eth2-01 to bond5. There are two constraints: the new interface must already exist (you can't create the bond as part of the process), and there is hard downtime for traffic through an interface when changing the interface (Edit: just remembered a third constraint: the new interface can't be used for anything already). Be sure to remove the non-bond interfaces from the VSX cluster object's Physical Interfaces section.

Once you've moved the interfaces to names the new boxes can support, you can do the vsx_util reconfigure to have the new boxes take over the old boxes in the cluster, but they won't be able to sync (the 23800 has 24c48t while the 9700 has 16c32t; can't sync from more to fewer). When you fail over, you'll have hard downtime.

JozkoMrkvicka
Authority
Authority

Sync can possibly work as each VS has static number of IPv4 and IPv6 cores defined in SmartConsole (by default 1/1). CoreXL is disabled on VS0. It doesnt take into consideration how many cores has new hardware, if at least the total number of IPv4 and IPv6 assigned cores to every VS is lower than cores available on new VSX hardware.

To judge if sync will work, we need to know how many VSs are configured and what is coreXL assignment of each VS.

Kind regards,
Jozko Mrkvicka
Bob_Zimmerman
MVP Gold
MVP Gold

Fair point. Hmm. I haven't personally tested that. It would need per-VS sync, but that's needed for VSLS anyway, and VSLS has been the default for quite a while.

In any case, renaming the interfaces is a must, and that involves outages. SmartConsole should really pop up a warning if you try to tell a VSX object about a "physical interface" which isn't a bond. It's almost always a bad idea for just this reason.

Martijn
MVP Platinum
MVP Platinum

Totally agree with Bob and Val.

What you can do to minimize interruptions, is to temporarily (couple of minutes) disable statefull inspection in the Global Properties and install policy on the new cluster before migrating from the old VSX cluster to the new one.

I know, security-wise not something you want, but it can help with a smooth migration.

Current sessions aren't dropped by out-of-state on when the new cluster is active. If all is OK and all tests are succesfull, you can enable statefull inspection again. But it is important you leave it of for just the migration and do not forget the turn in back on.

Martijn

 

JozkoMrkvicka
Authority
Authority

The names of all interfaces must match between old and new VSX. The only solution is to use bonds for everything, even if only single physical interface is used. Bond must have at least 1 member. For example bond5 can have only eth6-04 interface as part of bond on old VSX, but on new VSX, the bond5 can have only eth1 as part of bond. As far as bond5 is used in Physical Interfaces on VS0, it doesnt matter what are members between old and new VSX. Before 'vsx_util reconfigure', you just need to create bond5 with eth1 on new VSX hardware.

This aproach requires huge time planning including outages during transfering physical interface into bond. But once you are done, you can use any hardware during future VSX replacements (depending on NICs and type of HW).

Kind regards,
Jozko Mrkvicka

Leaderboard

Epsum factorial non deposit quid pro quo hic escorol.

Upcoming Events

    CheckMates Events