A load balancer can provide many load balancing methods, which are what we commonly call scheduling methods or algorithms:

Round Robin
This method cyclically distributes received requests to each machine in the server cluster, i.e., the available servers. If this method is used, all servers marked for the virtual service should have similar resource capacity and host applications with similar loads. If all servers have the same or similar performance, choosing this method will make the server load equal. Based on this premise, round-robin scheduling is a simple and effective way to distribute requests. However, for environments where servers differ, choosing this method means that weaker servers will also receive round-robin requests in the next cycle, even if the server is already unable to handle the current request. This may cause weaker servers to become overloaded.
Weighted Round Robin
This algorithm solves the shortcomings of the simple round-robin scheduling algorithm: incoming requests are distributed in order to the servers in the cluster, but the weights pre-assigned to each server are taken into account. The administrator simply defines each server's weight based on its processing capability. For example, the most powerful server A is given a weight of 100, while the least powerful server is given a weight of 50. This means that before server B receives its first request, server A will receive 2 consecutive requests, and so on.
Least Connection
Neither of the above two methods considers that the system cannot identify how many connections are being maintained at a given time. Therefore, it may happen that server B receives fewer connections than server A but is already overloaded, because users on server B hold their connections open for a longer time. That is to say, the number of connections, i.e., the load on a server, is cumulative. This potential problem can be avoided by the "Least Connection" algorithm: incoming requests are distributed according to the number of connections currently open on each server. That is, the server with the fewest active connections automatically receives the next incoming request. The basic principle is the same as simple round robin: all servers with virtual services should have similar resource capacity. It is worth noting that in low-traffic configurations, the traffic on each server is not the same, and the first server will be preferred. This is because if all servers are identical, the first server is prioritized; until the first server has continuous active traffic, otherwise the first server will always be the one selected first.
Least Connection Slow Start Time
For the least-connection and weighted least-connection scheduling methods, when a server has just joined the production environment, a time period can be configured for it. During this period, the number of connections is limited and increases slowly. This provides a "transition time" for the server to ensure that it does not become overloaded immediately after startup due to too many assigned connections. This value is set in the L7 configuration interface.
Weighted Least Connection
If servers have different resource capacities, the "Weighted Least Connection" method is more suitable. The number of active connections determined by the weights customized by the administrator based on server conditions generally provides a very balanced utilization of servers, because it combines the advantages of both least connection and weighting. Typically, this is a very fair distribution method because it uses the ratio of connections to server weight; the server with the lowest ratio in the cluster automatically receives the next request. However, please note that when using this method in low-traffic conditions, refer to the caveats in the "Least Connection" method.
Agent Based Adaptive Balancing
In addition to the above methods, the load balancer contains adaptive logic to periodically monitor server status and the weights of those servers. For the very powerful "Agent Based Adaptive Balancing" method, the load balancer periodically checks the load of all servers in this way: each server must provide a file containing a number from 0 to 99 that indicates the actual load of that server (0 = idle, 99 = overloaded, 101 = failed, 102 = disabled by administrator), and the server obtains this file via HTTP GET; at the same time, for servers in the cluster, providing their own load in the form of a binary file is also one of the server's tasks. However, there is no restriction on how servers calculate their own load. Based on the overall load of the servers, two strategies can be chosen: in normal operation, the scheduling algorithm calculates a weight ratio based on the ratio between the collected server load value and the number of connections assigned to that server. Therefore, if a server is overloaded, its weight will be transparently readjusted by the system. As with the weighted round-robin scheduling method, incorrect distributions can be recorded so that different weights can be effectively assigned to different servers. However, in a very low-traffic environment, the load values reported by servers will not establish a representative sample; distributing load based on these values will lead to loss of control and command oscillation. Therefore, in this case it is more reasonable to calculate load distribution based on static weight ratios. When the load of all servers is below the administrator-defined lower limit, the load balancer automatically switches to weighted round-robin to distribute requests; if the load is greater than the administrator-defined lower limit, the load balancer switches back to the adaptive mode.
Fixed Weighted
The highest weight is used only when the weight values of other servers are all very low. However, if the server with the highest weight goes down, the server with the next highest priority will serve the client. In this method, the weight of each real server needs to be configured based on server priority.
Weighted Response
Traffic scheduling is done through weighted round robin. The weights used in weighted round robin are calculated based on the response times of server availability checks. Each availability check is timed to record how long it took to respond successfully. Note, however, that this method assumes server heartbeat checks are based on machine speed, but this assumption may not always hold true. The sum of all servers' response times on the virtual service is added together, and this value is used to calculate the weight of an individual physical server; this weight value is calculated approximately every 15 seconds.
Source IP Hash
This method generates a hash value from the source IP of the request and uses this hash value to find the correct real server. This means that for the same host, the corresponding server is always the same. With this method, you do not need to save any source IPs. However, note that this method may lead to unbalanced server load.
Source: http://www.codeceo.com/article/balanced-algorithm.html