Fault tolerance



How OpenVidu Pro provides fault tolerance 🔗

Fault tolerance in OpenVidu Pro is provided through the presence of multiple Media Nodes (see OpenVidu Pro architecture). An OpenVidu Pro cluster with at least 2 Media Nodes ensures that if one Media Node goes down for any reason, the other can take over the affected OpenVidu sessions. This is represented in the image below.

The points below summarize the functioning of OpenVidu Pro in terms of fault tolerance upon a Media Node crash:

  1. The Master Node keeps a persistent, full-duplex connection with Media Nodes through WebSocket.
  2. Whenever the Master Node detects a disconnection of a Media Node, it starts a reconnection process that grants 3 seconds to successfully reconnect. At this point, it is possible that videos on the client side are frozen, if the disconnection of the Media Node was an actual crash of the media routing process.
  3. If the Master Node succeeds in reconnecting to the Media Node within the allowed time interval, it is considered a one-time issue and no further action is taken. But if the reconnection is not possible, then every OpenVidu session hosted in the crashed Media Node is closed and every participant will receive the proper event to allow the application to rebuild the session. This is further explained in the next section Making your OpenVidu app fault tolerant.



Making your OpenVidu app fault tolerant 🔗

OpenVidu Pro delegates the recovery of the sessions to the application in the event of a Media Node crash. The application should simply re-create the crashed session, which translates in nothing more than repeating the normal process of joining users to a session:

  • Initialize the Session in OpenVidu Server from your application's backend.
  • Create a new Connection for the Session from your application's backend.
  • Return the Connection's token to your application's frontend so it can use it to call Session.connect.

The key part is letting the application's frontend know when to ask the application's backend for a new token to re-connect to a recently crashed session. The application's frontend must listen to sessionDisconnected event and identify its reason. If it is nodeCrashed then the application's frontend just needs to ask the application's backend for a new token for a session with the same identifier as the previous one. This is reproduced in the snippet below, in a JavaScript code using openvidu-browser library.

var OV = new OpenVidu();
var session = OV.initSession();

session.on('sessionDisconnected', event => {

  if (event.reason === 'nodeCrashed') {

    // User was evicted from the session upon a node crash
    console.warn('Your session has been closed due to a node crash!');

    // HERE THE CLIENT SHOULD RE-RUN THE PROCESS OF CONNECTING TO THE SESSION AS NORMAL.
    // THE SESSION SHOULD KEEP THE SAME IDENTIFIER.

  } else {

    // User left the session for any conventional reason
    console.log('You left the session!');

    // HERE THE CLIENT SHOULD DO WHATEVER IS NECESSARY UPON A NORMAL SESSION CLOSURE.

  }
});

Your application's backend can also receive the nodeCrashed CDR event if you want. Listening to this CDR event is really not necessary for achieving fault tolerance and re-building sessions after a Media Node crash, but you can still use it for custom logic and monitoring purposes.

If you want to see an example of an application that automatically reconnects users after a node crash, take a look to the simple openvidu-high-availability tutorial.