Failover between upstreams
This recipe builds a client-side etemenanki-app that sends your traffic through one of three upstream proxy servers. It uses the first server while that server accepts connections, moves new connections to the second when the first stops answering, and moves back once the first recovers. A variant at the end spreads connections over all healthy servers instead.
Read it if you run more than one upstream server and want the client to route around an outage without you editing the config. It also explains the health probe in detail, because the probe decides everything a balancer does, and it checks less than you might expect.
What you build
Section titled “What you build”flowchart LR
App["Application"] --> In["socks-in 127.0.0.1:1080"]
In --> R{"Route"}
R -- "private addresses" --> Direct["direct"]
R -- "everything else (default)" --> B["balancer upstream"]
B -- "1st choice" --> P["primary: Trojan"]
B -- "2nd choice" --> S["secondary: VLESS"]
B -- "3rd choice" --> T["tertiary: VMess"]
Prober["Health prober"] -. "TCP connect" .-> P
Prober -. "TCP connect" .-> S
Prober -. "TCP connect" .-> T
- Three outbounds,
primary,secondaryandtertiary, each reach a different server. The protocols differ on purpose to show that members do not have to match. Yours can all use the same protocol. - A
[[balancer]]namedupstreamgroups them. Withstrategy = "failover"the order of itsoutboundslist is the priority. [route].default = "upstream"sends every flow that no rule catches to the balancer, so the balancer is the fallback for all traffic.- A background prober opens a TCP connection to each member’s
serverandporton a timer. The balancer only picks members whose last probe succeeded.
Before you start
Section titled “Before you start”The balancer’s probe only checks that a server accepts TCP connections. It does not check your password, UUID, TLS name or WebSocket path (see What the probe does not check). Make sure each upstream works on its own first, so that a healthy probe really means a working path.
-
Write a config with only the first upstream outbound and
[route].defaultpointing at it. Copy its[[outbound]]block from the recipe for its protocol, such as VLESS over WebSocket and TLS or VMess over gRPC and TLS. -
Check the config and start it:
Terminal window etemenanki-app --test -c config.tomletemenanki-app -c config.toml -
From another terminal, send a request through it:
Terminal window curl --socks5-hostname 127.0.0.1:1080 https://example.com/ -o /dev/null -w '%{http_code}\n'A status code such as
200means the whole path works: TCP, TLS, the proxy protocol and your credentials. -
Stop the proxy with Ctrl-C, point
[route].defaultat the next outbound, and repeat until every upstream has passed on its own.
The configuration
Section titled “The configuration”# The "Failover between upstreams" recipe: a local SOCKS proxy that sends# everything through one of three upstream servers. "primary" is used while it# accepts TCP connections, then "secondary", then "tertiary". Private addresses# go out directly.# Replace the server names, the password and the UUIDs with your servers' values.
[log]level = "info"
[[inbound]]tag = "socks-in"protocol = "socks"listen = "127.0.0.1"port = 1080
[[outbound]]tag = "direct"protocol = "freedom"
# Member 1: Trojan over TLS.[[outbound]]tag = "primary"protocol = "trojan"server = "proxy1.example.com"port = 443
[outbound.stream]network = "tls"
[outbound.stream.tls]server_name = "proxy1.example.com"
[outbound.settings]password = "replace-with-a-long-random-password"
# Member 2: VLESS over WebSocket and TLS.[[outbound]]tag = "secondary"protocol = "vless"server = "proxy2.example.com"port = 443
[outbound.stream]network = "ws"security = "tls"
[outbound.stream.ws]path = "/vless"
[outbound.stream.tls]server_name = "proxy2.example.com"
[outbound.settings]id = "11111111-2222-3333-4444-555555555555"
# Member 3: VMess over gRPC and TLS.[[outbound]]tag = "tertiary"protocol = "vmess"server = "proxy3.example.com"port = 443
[outbound.stream]network = "grpc"security = "tls"
[outbound.stream.grpc]service_name = "tunnel"
[outbound.stream.tls]server_name = "proxy3.example.com"
[outbound.settings]id = "11111111-2222-3333-4444-666666666666"
# All three members carry UDP, so the balancer can take UDP flows too.# List order is the priority under "failover".[[balancer]]tag = "upstream"outbounds = ["primary", "secondary", "tertiary"]strategy = "failover" # the default, written out for clarityprobe_interval = 15 # seconds between probes of each member (default 30)probe_timeout = 5 # seconds a probe may take (default 5)
[route]# Set explicitly: without it the default would be "direct", the first outbound.default = "upstream"
[[route.rule]]outbound = "direct"cidr = ["10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16", "127.0.0.0/8", "fc00::/7"]What each part does:
directcomes first among the outbounds only because it is a common layout. It does not matter for the balancer, but it does matter for the default route: without[route].default, flows that match no rule go to the first[[outbound]], which here isdirect. A balancer never becomes the default on its own, so this config setsdefault = "upstream"explicitly.- The three members are ordinary outbounds. Nothing in an
[[outbound]]block changes when it becomes a balancer member, and you can still route to a member’s own tag directly in a rule. - All three members carry UDP. Trojan, VLESS and VMess outbounds carry UDP, so the balancer can take UDP flows as well as TCP. An
httporshadowsocksmember would silently drop every UDP flow the balancer hands it; see UDP to an outbound without datagrams if you need one. [[balancer]]lists the members in priority order.probe_interval = 15halves the default of 30 seconds, so a dead server is noticed sooner. Choosing the probe timings explains the trade-off.- The
directrule keeps private and loopback destinations away from the upstreams. Rules are checked in order before the default applies, so this rule wins for those addresses. Acidrrule only sees destinations that arrive as an IP address: a name such asnas.lanthat resolves to192.168.1.10still goes to the balancer (see Domains are not resolved for routing).
Balancer settings
Section titled “Balancer settings”A [[balancer]] accepts exactly these keys. Any other key stops the config from loading, for example unknown field `probe_intervall`, expected one of `tag`, `outbounds`, `strategy`, `probe_interval`, `probe_timeout` .
| Key | Type | Required | Default | Description |
|---|---|---|---|---|
tag | string | yes | — | Name of the balancer. [route].default and [[route.rule]].outbound refer to it exactly as they refer to an outbound. It must not repeat any [[outbound]] tag or the tag of another balancer; both fail with balancer tag <tag> collides with an outbound tag. |
outbounds | array of strings | yes | — | The member outbounds, by exact tag. At least one (a balancer needs at least one outbound). Only [[outbound]] tags are accepted: a missing tag or the tag of another balancer fails with balancer <tag> references unknown outbound tag: <member>. Every member needs a server and port that a TCP connect can probe. hysteria2 members are always refused, and freedom, blackhole and wireguard members are refused because they normally have no server and port; both fail with balancer <tag>: outbound <member> has no upstream a TCP health probe can reach, so it cannot be balanced. Under failover the list order is the priority. A tag listed twice is accepted and counts twice. |
strategy | string (enum) | no | "failover" | How a healthy member is chosen for each new flow: failover takes the first healthy member in outbounds order, round_robin takes each healthy member in turn. Matched exactly and case-sensitively; anything else, such as "roundrobin" or "Failover", fails with unknown balancer strategy "<value>" (expected "failover" or "round_robin"). |
probe_interval | u64 | no | 30 | Seconds to wait after one health probe of a member finishes before the next one starts. Each member is probed on its own schedule, and the first probe runs as soon as the configuration starts. 0 is accepted and makes the probes run back to back with no pause, opening connections to the upstream continuously. |
probe_timeout | u64 | no | 5 | Seconds one probe may take, name resolution included, before the member counts as down. 0 is accepted but leaves a probe only until the next timer tick, about a millisecond, so a member is marked down unless its connect completes almost at once. Use at least 1. |
probe_interval and probe_timeout are whole seconds written as TOML integers. A string such as "30s" or a negative number is a parse error.
Run it and watch a failover
Section titled “Run it and watch a failover”-
Check the config.
--testbuilds everything, balancer included, but does not start the prober, so it does not tell you whether the servers are reachable:Terminal window etemenanki-app --test -c config.tomlConfiguration OK. -
Start it:
Terminal window etemenanki-app -c config.tomlThe prober starts with the configuration and probes every member at once. A member that answers its first probe logs nothing, because every member already counts as healthy before its first probe. A member that fails it logs a line at
infolevel:balancer member tertiary is now down -
Make
primaryunreachable. Stop the proxy service on that server, or block it from the client for a moment. On Linux, for example:Terminal window sudo iptables -I OUTPUT -p tcp -d proxy1.example.com --dport 443 -j REJECTiptablesonly covers IPv4. Ifproxy1.example.comalso has an IPv6 address, add the same rule withip6tables, or the probe connects over IPv6 andprimarystays healthy.Within one
probe_intervalthe log shows:balancer member primary is now downFrom then on, new connections go to
secondary. Run thecurlcommand from Before you start again to confirm. If your upstreams have different exit addresses, an IP echo service shows which one carried the request. etemenanki-app does not log which member a flow went to. -
Undo the block:
Terminal window sudo iptables -D OUTPUT -p tcp -d proxy1.example.com --dport 443 -j REJECTRemove the
ip6tablesrule too if you added one.Within one
probe_intervalthe log showsbalancer member primary is now up, and new connections go back toprimary.
Connections that were already open when a member went down are not moved. A TCP connection stays on the member it started on until it closes, and a UDP association keeps the member its sub-link to the balancer opened with. Only new connections follow the balancer’s current choice.
Spreading load with round_robin
Section titled “Spreading load with round_robin”If your upstreams are equivalent and you would rather use all of them than keep some on standby, change the strategy:
[[balancer]]tag = "upstream"outbounds = ["primary", "secondary", "tertiary"]strategy = "round_robin"probe_interval = 15probe_timeout = 5Each new connection now goes to the next healthy member in turn. A member that goes down drops out of the rotation, and it rejoins when it answers again.
failover |
round_robin |
|
|---|---|---|
| Which member a new connection gets | The first healthy member in outbounds order |
The next healthy member in rotation |
| What the list order means | Priority | Only the rotation order |
| Load on the servers | All on one server, the others idle | Spread across every healthy server |
| When a member recovers | New connections go back to it at once | It rejoins the rotation |
| Every member down | The first member in outbounds |
The first member in outbounds |
| Good for | A preferred server with standbys | Several equivalent servers |
A few things to know about round_robin:
- It rotates per connection, not per destination. Two connections to the same website can leave through different servers. Sites that tie a login session to your IP address may notice, so use
failoverif your upstreams have different exit addresses and that matters. - The rotation skips members that are down. One shared counter walks over the members that are healthy at that moment, so when the healthy set changes the rotation continues over the new set.
- Weights are possible only by repetition. Listing a tag twice, as in
outbounds = ["primary", "primary", "secondary"], gives it two slots in the rotation. Each slot is a separate member with its own probe, so that server is also probed twice as often. There is no weight setting.
You can also keep both: a failover balancer as [route].default and a separate round_robin balancer that a [[route.rule]] points at. Each balancer probes its own members, so an outbound that belongs to both is probed twice. Balancers cannot contain other balancers.
How the health probe works
Section titled “How the health probe works”Each balancer runs one background task per member. The task repeats a single cycle for as long as the configuration runs:
- Resolve the member’s
serverwith the resolver configured in[dns](see DNS), the same resolver and cache the outbound itself uses. An IP address needs no lookup. - Try a plain TCP connection to each resolved address in turn, and close it as soon as one connects. Nothing is sent over it.
- If a connection succeeded within
probe_timeoutseconds, the member is healthy. A lookup failure, a refused connection or the timeout makes it down. The timeout covers the lookup and all addresses together. - Log a line if the state changed, then wait
probe_intervalseconds and start again.
stateDiagram-v2 [*] --> Healthy: config starts, before the first probe Healthy --> Down: one probe fails or times out Down --> Healthy: one probe connects Healthy --> Healthy: probe connects Down --> Down: probe fails
Members start out healthy. If they started out down, the balancer would have nothing to send traffic to until the first round of probes finished. There is no threshold in either direction: one failed probe marks a member down and one successful probe marks it up again.
The prober stops when the configuration stops, and it runs whether or not anything is routed to the balancer.
What the probe does not check
Section titled “What the probe does not check”The probe answers one question: does something accept TCP connections at this member’s server and port? That is the failure a balancer exists to route around, an unreachable or stopped server. It does not use the member’s [outbound.stream] settings or its credentials, because checking those would mean running a second proxy client inside the prober.
Some finer points:
- The probe ignores
address_family. It tries every address the resolver returns, IPv4 and IPv6 alike, in the order the resolver returns them. A member restricted toipv4_onlycan probe healthy over IPv6 while its traffic can only use IPv4. - One address that swallows packets can take the whole timeout. The addresses are tried one after another within a single
probe_timeout. If the first address never answers, the probe times out before it reaches the second, and the member counts as down even though the second address works. Pointserverat a name or address whose first answer is reachable from the client. - Servers see the probes. Each probe is a TCP connection that closes without sending a byte. Proxy servers often log these as failed or empty handshakes. With three members and
probe_interval = 15, each server sees about four such connections a minute from each client.
When every upstream is down
Section titled “When every upstream is down”When no member is healthy, the balancer sends new connections to the first member in outbounds, primary in this recipe, instead of dropping them. Probing can be wrong for a minute, for example when the client’s own network blips, and a connection sent to a server that may have recovered has some chance to work, while a dropped one has none.
In practice, when all your upstreams are really down:
- every new connection that reaches the balancer fails, and client applications see connection errors;
- the log shows a
balancer member … is now downline for each member; the failed connections themselves are logged only atdebuglevel, one… connection from … ended: …line each; - traffic that a rule sends elsewhere, such as the private ranges sent to
directhere, keeps working; - as soon as any member answers a probe, new connections go to it.
The balancer never falls back to direct on its own. If you want traffic to leave directly when every upstream is gone, that is a routing decision you have to write yourself, and it usually defeats the purpose of the proxy.
Choosing the probe timings
Section titled “Choosing the probe timings”Two numbers decide how the balancer behaves, and they pull in opposite directions.
How long until a dead server is noticed. A probe runs probe_interval seconds after the previous one finished. If a server dies right after a successful probe, the next probe starts up to probe_interval seconds later. A server that refuses connections is marked down almost at once. A server whose packets vanish, the usual case when a host or its network is down, keeps the probe waiting for the full probe_timeout. So the worst case is about probe_interval + probe_timeout seconds, during which new connections to that server fail. Recovery is noticed by the first probe after the server comes back: within about probe_interval seconds when the failing probes were refused, and within about probe_interval + probe_timeout seconds when they timed out.
How many probes the servers see. Each member gets roughly one TCP connection every probe_interval seconds from every client that runs this config.
| Profile | probe_interval |
probe_timeout |
Worst case to mark a silent server down | Probes per member per minute |
|---|---|---|---|---|
| Default | 30 |
5 |
about 35 s | about 2 |
| This recipe | 15 |
5 |
about 20 s | about 4 |
| Quick | 5 |
3 |
about 8 s | about 12 |
Some guidance:
- Keep
probe_timeoutat 3 seconds or more on real networks. A probe is one TCP handshake, and when the first SYN packet is lost, Linux sends it again after 1 second and again 2 seconds later. Withprobe_timeout = 1, a single lost packet marks a healthy server down, and because there is no threshold,failovertraffic then bounces between members. Every probe of a member named by a domain also includes the lookup, which is slow whenever it misses the resolver’s cache. - Do not set
probe_timeouthigher than it needs to be. It only matters for silent failures, and it adds directly to the time those take to notice. - Pick
probe_intervalby how long you can tolerate failing connections. Below about 5 seconds the gain is small and the probe traffic grows quickly, especially with many clients. - Never use
0for either. Both are accepted.probe_interval = 0makes each member’s probes run back to back with almost no pause, a constant stream of connections to your servers.probe_timeout = 0gives a probe only until the next timer tick, about a millisecond, so any member that is not on the local network is marked down on every probe. When that happens to all members, the balancer treats them all as down and traffic always goes to the first member: there is no failover at all.
Why some outbounds cannot be members
Section titled “Why some outbounds cannot be members”A member must have a TCP endpoint for the probe to connect to. The config refuses any member that does not, rather than treating it as always healthy, which would hide the outage the balancer exists to catch.
| Member protocol | Accepted | Why |
|---|---|---|
socks, http, trojan, vless, vmess, shadowsocks |
yes | Each dials a TCP server and port, which the probe can connect to |
hysteria2 |
no | It has a server and port, but a Hysteria 2 server listens on UDP only. A TCP probe would fail every time, mark the member down on its first probe and never mark it up again |
wireguard |
no | It has no server; its peer is settings.endpoint, which is UDP |
freedom, blackhole |
no | There is no upstream server to probe |
The protocol aliases count too: hysteria and hy2 are refused like hysteria2, direct like freedom and block like blackhole.
A refused member stops the config from loading. --test reports it as:
configuration invalid: balancer upstream: outbound direct has no upstream a TCP health probe can reach, so it cannot be balancedA normal start exits with the same message after failed to start:, and a reload that introduces it is refused and keeps the running config (see Reloading).
If you want a Hysteria 2 or WireGuard server as a fallback, a balancer cannot do it. Route to that outbound with a rule or make it [route].default on its own; switching to or from it is then a config change. Adding direct as a last-resort member is refused for the same reason, and see When every upstream is down for why you probably do not want it.
Reloading
Section titled “Reloading”etemenanki-app reloads when you save the config file. The balancer is rebuilt with everything else: the old probers stop, every member starts out healthy again, and the new probers probe at once. A reload also closes every open connection. The reload log line lists changes to inbounds, outbounds, the route and the log section only, so a save that changes only a [[balancer]] block logs config reload: no changes although the new settings do take effect. See Hot reload.
Common errors
Section titled “Common errors”Each of these fails etemenanki-app --test, where the message follows configuration invalid:, and a normal start, where it follows failed to start:. On a reload it follows reload: build failed, keeping current config: and the previous config keeps running. A mistyped key or a wrong value type in [[balancer]] is a TOML parse error instead, such as unknown field `probe_intervall` or invalid type: string "30s", expected u64, reported after TOML parse error at line …. On a reload a parse error follows reload: parse failed, keeping current config:, and the previous config keeps running as well.
| Message | Cause | Fix |
|---|---|---|
balancer <tag>: outbound <member> has no upstream a TCP health probe can reach, so it cannot be balanced |
A hysteria2, wireguard, freedom or blackhole member |
Remove it from outbounds and route to it directly |
balancer <tag> references unknown outbound tag: <member> |
A member tag is misspelt, not defined, or is another balancer | List only [[outbound]] tags |
balancer tag <tag> collides with an outbound tag |
The balancer’s tag repeats an outbound or another balancer |
Give it a unique name |
a balancer needs at least one outbound |
outbounds = [] |
List at least one member |
unknown balancer strategy "<value>" (expected "failover" or "round_robin") |
A misspelt or differently cased strategy | Write failover or round_robin exactly |
route references unknown outbound tag: <tag> |
[route].default or a rule names a balancer that does not exist |
Fix the tag or define the balancer |
And one mistake that produces no error: leaving out [route].default. The config loads, the balancer probes its members, and all traffic goes to the first [[outbound]] instead, which is direct in this recipe.