Balancers
A balancer is a named group of outbounds. You route to its tag the same way you route to an outbound, and etemenanki-app picks one member for each new flow. A background prober checks every member with a TCP connect at a fixed interval, and the pick prefers members whose last probe succeeded.
Use a balancer when you have more than one upstream proxy server and want traffic to keep flowing when one of them becomes unreachable (failover), or want to spread connections across several servers (round_robin).
A minimal example
Section titled “A minimal example”Two Trojan servers, with primary preferred while it is reachable:
[[inbound]]tag = "socks-in"protocol = "socks"listen = "127.0.0.1"port = 1080
[[outbound]]tag = "direct"protocol = "freedom"
[[outbound]]tag = "primary"protocol = "trojan"server = "proxy1.example.com"port = 443
[outbound.stream]network = "tls"
[outbound.settings]password = "replace-with-a-long-random-password"
[[outbound]]tag = "backup"protocol = "trojan"server = "proxy2.example.com"port = 443
[outbound.stream]network = "tls"
[outbound.settings]password = "replace-with-a-long-random-password"
# Failover: "primary" while it answers, "backup" otherwise.[[balancer]]tag = "proxy"outbounds = ["primary", "backup"]
[route]default = "proxy"
[[route.rule]]outbound = "direct"cidr = ["10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16"]Private addresses go out directly. Everything else goes to the balancer proxy. It uses primary while a TCP connect to proxy1.example.com:443 succeeds, and backup while it does not. strategy defaults to failover. Each member is probed with a 5-second timeout, and the next probe starts 30 seconds after the previous one ends.
[route].default is set explicitly on purpose. Without it the default is the first [[outbound]] in the file, which here is direct. A balancer is never the implicit default.
Settings
Section titled “Settings”Each [[balancer]] is a top-level array-of-tables entry, next to [[outbound]]. It accepts exactly these keys. Any other key is a parse error, for example unknown field `probe_intervall`, expected one of `tag`, `outbounds`, `strategy`, `probe_interval`, `probe_timeout` .
| Key | Type | Required | Default | Description |
|---|---|---|---|---|
tag | string | yes | — | Name of the balancer. [route].default and [[route.rule]].outbound refer to it exactly as they refer to an outbound. It must not repeat any [[outbound]] tag or the tag of another balancer; both fail with balancer tag <tag> collides with an outbound tag. |
outbounds | array of strings | yes | — | The member outbounds, by exact tag. At least one (a balancer needs at least one outbound). Only [[outbound]] tags are accepted: a missing tag or the tag of another balancer fails with balancer <tag> references unknown outbound tag: <member>. Every member needs a server and port that a TCP connect can probe. hysteria2 members are always refused, and freedom, blackhole and wireguard members are refused because they normally have no server and port; both fail with balancer <tag>: outbound <member> has no upstream a TCP health probe can reach, so it cannot be balanced. Under failover the list order is the priority. A tag listed twice is accepted and counts twice. |
strategy | string (enum) | no | "failover" | How a healthy member is chosen for each new flow: failover takes the first healthy member in outbounds order, round_robin takes each healthy member in turn. Matched exactly and case-sensitively; anything else, such as "roundrobin" or "Failover", fails with unknown balancer strategy "<value>" (expected "failover" or "round_robin"). |
probe_interval | u64 | no | 30 | Seconds to wait after one health probe of a member finishes before the next one starts. Each member is probed on its own schedule, and the first probe runs as soon as the configuration starts. 0 is accepted and makes the probes run back to back with no pause, opening connections to the upstream continuously. |
probe_timeout | u64 | no | 5 | Seconds one probe may take, name resolution included, before the member counts as down. 0 is accepted but leaves a probe only until the next timer tick, about a millisecond, so a member is marked down unless its connect completes almost at once. Use at least 1. |
probe_interval and probe_timeout are whole seconds written as TOML integers. probe_interval = "30s" fails with invalid type: string "30s", expected u64, and a negative number fails with invalid value: integer `-1`, expected u64.
Choosing a member
Section titled “Choosing a member”The balancer picks a member when a flow is dispatched to it, not when the route is built. A member that goes down therefore stops receiving flows as soon as the prober notices, without a reload.
failover (default) |
round_robin |
|
|---|---|---|
| Picks | The first healthy member in outbounds order |
The next healthy member in rotation |
| List order | Is the priority | Only sets the rotation order |
| When a preferred member recovers | New flows go back to it | It rejoins the rotation |
| Typical use | A primary server with one or more standbys | Several equivalent servers |
| Every member down | The first member in outbounds |
The first member in outbounds |
flowchart TB
F["A new flow is routed to the balancer"] --> H{"Is any member healthy?"}
H -- no --> First["The first member in outbounds"]
H -- yes --> S{"strategy"}
S -- failover --> FO["The first healthy member, in list order"]
S -- round_robin --> RR["The next healthy member in rotation"]
First --> D["Dial through the chosen outbound"]
FO --> D
RR --> D
Some consequences:
- Selection is per flow. A TCP connection stays on the member it started on until it closes, even if that member later goes down or a higher-priority member comes back. Only new connections move.
- For UDP, the choice is made when the association’s sub-link to the balancer opens. The association keeps that member until the sub-link ends. See How UDP reaches an outbound.
- When every member is down, the first member is used anyway. A flow sent to an upstream that may have recovered has a chance; a dropped flow has none. The balancer never becomes a black hole because probing had a bad minute.
- A failed dial is not retried on another member. By the time the dial fails, the flow has been handed to that member’s outbound and cannot be replayed elsewhere. The client sees the failure and retries on its own.
round_robinrotates one shared counter over the members that are healthy at that moment. When the healthy set changes, the rotation continues over the new set. Listing a tag twice gives it two slots in the rotation.
Health checks
Section titled “Health checks”Each balancer probes each of its members in a separate background task:
- It resolves the member’s
serverwith the resolver configured in[dns](see DNS). - It opens a plain TCP connection to
server:port, trying each resolved address in turn, and closes it as soon as it connects. It does not apply the member’s[outbound.stream]settings. - If a connect succeeds within
probe_timeoutseconds, name resolution included, the member is healthy. A resolution failure, every address refusing the connection, or the timeout makes it down. - It waits
probe_intervalseconds from the end of that probe, then probes again.
The first probe runs as soon as the configuration starts. Until a member’s first probe completes, it counts as healthy: marking every member down at start-up would send nothing anywhere for a whole interval.
A change of state is logged at info level:
balancer member primary is now downbalancer member primary is now upThe failover in the minimal example plays out like this:
sequenceDiagram participant P as Prober participant A as primary participant B as backup participant R as Balancer proxy P->>A: TCP connect succeeds Note over R: new flows go to primary P->>A: TCP connect refused or timed out Note over P,R: log - balancer member primary is now down Note over R: new flows go to backup P->>A: TCP connect succeeds again Note over P,R: log - balancer member primary is now up Note over R: new flows go back to primary
How quickly traffic moves
Section titled “How quickly traffic moves”The prober only notices a change on its next probe. Between an upstream going away and the balancer marking it down, new flows still go to it and fail. With the defaults, that window is at most about 35 seconds: the 30-second probe_interval, plus up to the 5-second probe_timeout if the upstream silently drops packets instead of refusing the connection. Recovery is noticed on the same schedule.
Lower probe_interval for faster reaction, at the cost of one TCP connection per member per interval to each upstream. probe_interval = 1 and probe_timeout = 1 are the smallest useful values.
A few more details:
- The probe ignores the member’s
address_family. It tries every address the resolver returns, IPv4 and IPv6, in the resolver’s order, and the first one that connects makes the member healthy. A member set toipv4_onlycan probe healthy over IPv6 while its traffic can only use IPv4. - One slow address can fail the whole probe. The addresses are tried one after another, and only the overall
probe_timeoutbounds them. If the first address silently drops packets, its connect uses up the whole timeout and the member is marked down, even when a later address would have answered. - Upstream servers see the probes. Each probe is a TCP connection that closes without sending anything, one per member every
probe_interval. - Members are per balancer. An outbound listed in two balancers is probed twice, once by each, and each balancer keeps its own view of its health.
- Probes run even if nothing routes to the balancer.
--testdoes not probe.etemenanki-app --test -c config.tomlchecks that the balancer is well formed, not that its members are reachable.
Which outbounds can be members
Section titled “Which outbounds can be members”A member must have a TCP endpoint for the probe to connect to.
| Protocol | Member | Why |
|---|---|---|
socks, http, trojan, vless, vmess, shadowsocks |
yes | They dial a TCP server and port |
freedom (direct), blackhole (block) |
no | No upstream server and port to probe |
wireguard |
no | It has no server and port; the peer endpoint in its settings is UDP |
hysteria2 (hysteria, hy2) |
no, always | It has a server, but listens on UDP only, so a TCP probe would mark it down for good |
A refused member stops the configuration from loading. etemenanki-app --test reports:
configuration invalid: balancer proxy: outbound direct has no upstream a TCP health probe can reach, so it cannot be balancedThe check looks at whether the member has server and port, not at its protocol, except for hysteria2, which is refused by name. A freedom, blackhole or wireguard outbound that carries a stray server and port (which those protocols otherwise ignore) is accepted and probed at that address, while its traffic still goes direct, is dropped, or goes to the WireGuard peer. Do not rely on this.
Using a balancer in routing
Section titled “Using a balancer in routing”A balancer tag works wherever an outbound tag does:
- in
[route].default, to make it the fallback for flows that match no rule; - in
[[route.rule]].outbound, to send matching flows to it.
It cannot appear in another balancer’s outbounds: balancers do not nest. A route that names a tag that is neither an outbound nor a balancer fails with route references unknown outbound tag: <tag>. See Routing for how rules are matched.
This example spreads example.com traffic over two VLESS servers in turn, and probes them more often than the default:
[[inbound]]tag = "socks-in"protocol = "socks"listen = "127.0.0.1"port = 1080
[[outbound]]tag = "direct"protocol = "freedom"
[[outbound]]tag = "edge-a"protocol = "vless"server = "proxy1.example.com"port = 443
[outbound.stream]network = "ws"security = "tls"
[outbound.stream.ws]path = "/ws"
[outbound.settings]id = "11111111-2222-3333-4444-555555555555"
[[outbound]]tag = "edge-b"protocol = "vless"server = "proxy2.example.com"port = 443
[outbound.stream]network = "ws"security = "tls"
[outbound.stream.ws]path = "/ws"
[outbound.settings]id = "11111111-2222-3333-4444-555555555555"
[[balancer]]tag = "spread"outbounds = ["edge-a", "edge-b"]strategy = "round_robin"probe_interval = 10probe_timeout = 3
[[route.rule]]outbound = "spread"domain_suffix = ["example.com"]direct stays the default because it is the first outbound and [route].default is not set.
Reloads
Section titled “Reloads”A reload happens when the file’s contents change and the new configuration parses and builds. Every such reload rebuilds all balancers, like the rest of the configuration. The old probers stop with the old configuration and new ones start with the new configuration, so every member starts out healthy again and is probed at once. If the new configuration fails, the old one keeps running with its balancers and their health state.
The reload summary in the log lists inbound, outbound, route and log changes only. A reload that changes only a [[balancer]] is therefore logged as config reload: no changes, even though the new settings take effect and, as on any other reload, the old configuration’s connections are closed. See Hot reload.
Coming from Xray
Section titled “Coming from Xray”[[balancer]]is a top-level table, not part of the routing section.outboundslists exact tags. There is no prefix matching like Xray’sselector.- There are two strategies,
failoverandround_robin, spelt exactly like that. Xray’srandom,leastPingandleastLoadhave no equivalent. - There is no
fallbackTag. When every member is down, the balancer uses its first member. - Health checking is built in and configured on the balancer itself. There is no separate observatory, and the probe is a TCP connect rather than an HTTP request through the proxy.
See Migrating from Xray for the rest of the differences.
Common errors
Section titled “Common errors”All of these stop the configuration from loading. etemenanki-app --test -c config.toml logs configuration invalid: followed by the message, and a normal start logs failed to start: followed by it. On a reload, the message follows reload: parse failed, keeping current config: (for the missing-field errors and other TOML errors) or reload: build failed, keeping current config: (for the rest), and the old configuration keeps running.
| Message | Cause | Fix |
|---|---|---|
balancer tag <tag> collides with an outbound tag |
The balancer’s tag repeats an outbound tag or another balancer’s tag |
Give it a unique name |
balancer <tag> references unknown outbound tag: <member> |
A member is misspelt, not defined, or is itself a balancer | List only [[outbound]] tags |
balancer <tag>: outbound <member> has no upstream a TCP health probe can reach, so it cannot be balanced |
A hysteria2 member, or a freedom, blackhole or wireguard member (which normally has no server and port) |
Remove it, and route to it directly instead |
a balancer needs at least one outbound |
outbounds = [] |
List at least one member |
unknown balancer strategy "<value>" (expected "failover" or "round_robin") |
A misspelt or differently cased strategy | Write failover or round_robin exactly |
missing field `outbounds` / missing field `tag` |
A required key is absent | Add it |
route references unknown outbound tag: <tag> |
A rule or [route].default names a balancer that does not exist |
Fix the tag, or define the balancer |