Skip to content

Hot reload

etemenanki-app reloads its configuration on its own when the config file changes on disk. There is no reload signal and no control API: you save the file, and the running process builds the new configuration, checks it, and replaces the running one if it is valid.

This page explains how the change is noticed, what happens to an invalid file, what a successful reload does to listeners and open connections, which settings a reload applies, and how to edit a live file without surprises. Read it before you change the config of a proxy that carries real traffic. katana reloads its own configuration differently; the katana guide covers that.

Question Answer
What starts a reload? Any event in the directory that holds the config file, when the file’s bytes then differ from the last version etemenanki-app applied or rejected.
How fast? Within about 200 ms of the change.
Invalid new file? Logged. The running configuration keeps serving.
Valid new file? The whole running configuration is replaced: every listener closes and every open connection is dropped, even for a one-character comment change.
Certificate or geodata file changed? Nothing happens until the config file itself changes.
[log].level changed? Reported, but not applied. Restart to change it.
SIGHUP? Terminates the process. It is not a reload signal.

etemenanki-app calls one running configuration a generation. A generation owns everything built from the file: every outbound, balancer and its health prober, the router with its geodata, the DNS resolver and its cache, every listener, and every connection accepted on those listeners.

A reload never patches a generation. It builds a complete new one, and if that succeeds, it stops the old generation and starts the new one. Nothing is carried over from the old generation, so a reload behaves like a restart of the proxy inside the same process, with one difference: when the new file is broken, the old generation keeps serving.

At startup, etemenanki-app starts a file watcher on the parent directory of the file you pass with -c, not on the file itself. Many editors save by writing a new file and renaming it over the old one; a watch on the file would lose track of it after the first save, while a watch on the directory sees every save. When -c is a bare file name such as config.toml, the watched directory is the working directory.

The watcher then works like this:

  1. Any event in that directory wakes it, whichever file it concerns: a file created, opened, written, renamed, removed, or with changed permissions or timestamps.
  2. It waits for the burst to settle: it discards the events already queued, waits 200 ms, and discards the events that arrived meanwhile. An editor’s save, which often produces several events, therefore leads to one reload attempt.
  3. It reads the config file again, by path.
  4. It compares the bytes with the last version it applied or rejected. If they are identical, it does nothing and logs nothing.
  5. Otherwise it parses and builds the new file, as described in the next sections.

Events that arrive during a pass queue up and start another pass afterwards. Reading the config file in step 3 opens it, and that open is itself an event in the watched directory. So once the first event has arrived, every pass schedules the next one, and etemenanki-app goes on reading the config file about five times a second for as long as it runs and the file can be read. Each of those reads finds the same bytes and does nothing. The practical effect is that a change is read within about 200 ms of landing, at whatever moment it lands, which is one more reason to replace the file in one step (see Editing a live config safely).

The byte comparison is what keeps unrelated activity cheap, and it has consequences:

You do Result
Save the config file with any change, including a comment or whitespace Reload.
touch the config file, or save it with identical content Nothing. The bytes are the same.
Read, write, create or delete another file in the same directory The config file is read again and compared; nothing else happens.
Replace a certificate, key, CA or geodata file the config refers to Nothing, wherever that file lives: the config file’s bytes are unchanged. See Changing a certificate or geodata file.
Save the same broken content again Nothing. The broken bytes were already tried.
Delete or rename away the config file reload: cannot read … is logged, and the running generation keeps serving. A failed read opens nothing, so the watcher then waits for the next event in the directory, such as the file coming back.

If the watch cannot be set up, for example because the per-user inotify limits (fs.inotify.max_user_instances, fs.inotify.max_user_watches) are reached, etemenanki-app logs one line and keeps running without hot reload:

ERROR etemenanki_app: config hot-reload disabled: <reason>

The proxy serves the configuration it started with until you restart it. Changes to the file are not picked up, and nothing else tells you so, so check for this line after a start if you rely on reloads.

flowchart TB
  E["Event in the config directory"] --> D["Settle 200 ms, drop queued events"]
  D --> R{"Read the config file"}
  R -- "read error" --> RE["Log reload: cannot read, keep old generation"]
  R -- "ok" --> S{"Same bytes as the last read?"}
  S -- "yes" --> N["Do nothing"]
  S -- "no" --> P{"Parse and build"}
  P -- "error" --> PE["Log the error, remember these bytes, keep old generation"]
  P -- "ok" --> L["Log config reload: summary"]
  L --> C["Stop the old generation: listeners close, connections drop"]
  C --> B["Bind the new inbounds in file order"]
  B --> F["A failed bind is logged, the other inbounds serve"]

The new configuration is built completely before the old generation is touched. Building covers both validation phases described on the configuration file page: parsing the TOML, then building the DNS resolver, every outbound, every balancer, the router and every inbound, which reads every file the configuration refers to. Only binding the listeners and creating TUN devices happen after the switch.

For a moment during that build, both generations exist in memory. With large geodata files the process briefly needs room for two copies of the route data.

If the new file fails to parse or fails to build, etemenanki-app logs the error and keeps the old generation. Nothing stops, and no connection is dropped. The two log lines are:

ERROR etemenanki_app::instance: reload: parse failed, keeping current config: <error>
ERROR etemenanki_app::instance: reload: build failed, keeping current config: <error>

<error> is the same text --test prints for that file. For example, a misspelt key in an [[outbound]]:

ERROR etemenanki_app::instance: reload: parse failed, keeping current config: TOML parse error at line 17, column 1
|
17 | bogus = 1
| ^^^^^
unknown field `bogus`, expected one of `tag`, `protocol`, `server`, `port`, `stream`, `address_family`, `settings`

and a route that names an outbound that does not exist:

ERROR etemenanki_app::instance: reload: build failed, keeping current config: route references unknown outbound tag: nope

etemenanki-app remembers the bytes of the rejected file and does not try them again. Two things follow:

  • Saving the same broken content again, or a touch, does nothing. Fix the file and save it.
  • If the build failed because of something outside the file, such as a missing certificate (No such file or directory (os error 2), which does not name the file), putting the missing file in place is not enough. Change the config file too, for example by editing a comment, so that its bytes differ.

When you save a fixed file, the reload proceeds normally. If you simply restore the file that is already running, the bytes still differ from the rejected ones, so etemenanki-app performs a full reload that logs config reload: no changes and drops every connection, although the configuration is the same.

A file that cannot be read at all is treated differently: reload: cannot read <path>: <error> is logged, with <path> exactly as you passed it to -c. Nothing is remembered, and the next directory event tries again.

ERROR etemenanki_app::instance: reload: cannot read /etc/etemenanki/config.toml: No such file or directory (os error 2)
ERROR etemenanki_app::instance: reload: cannot read /etc/etemenanki/config.toml: Permission denied (os error 13)

The second line typically means that the new file has a different owner or mode from the old one, so the service user cannot read it. Fixing the owner or mode is itself a directory event, so the reload then goes ahead without another edit.

Before it stops the old generation, etemenanki-app logs a one-line summary of what changed:

INFO etemenanki_app::instance: config reload: inbounds ~[http-in]; outbounds +[out] -[direct]

The summary compares the old and new configuration section by section:

Part Printed when Meaning
inbounds +[a,b] An inbound tag is new Added inbounds, in the order of the new file.
inbounds -[a] An inbound tag is gone Removed inbounds, in the order of the old file.
inbounds ~[a] An inbound with the same tag has any different setting, users included Changed inbounds.
outbounds +[…] -[…] ~[…] The same, for outbounds Added, removed and changed outbounds.
route changed Anything under [route] differs, rules and geodata paths included
log changed Anything under [log] differs The new level is not applied; see below.
no changes None of the above The generation is still replaced.

Parts are separated by ; and tags inside brackets by ,. Inbounds and outbounds are matched by tag, so renaming a tag shows as one addition and one removal.

The summary is informational. It does not decide what is rebuilt, and it does not cover every section:

  • A change that only touches [dns] or [[balancer]] is logged as config reload: no changes, and it still takes effect.
  • Moving inbounds or outbounds to a different position in the file is not reported. This matters for outbounds: without [route].default, the first [[outbound]] is the default, so reordering them can change where unmatched traffic goes while the log says no changes.
  • A comment or whitespace change is logged as config reload: no changes, and it still replaces the whole generation.

Every reload replaces the whole generation

Section titled “Every reload replaces the whole generation”

Once the summary is logged, etemenanki-app cancels the old generation and waits for its listeners to close. Cancelling drops every task the generation spawned:

  • Every listener stops accepting and releases its port. A Unix socket inbound removes its socket file.
  • Every open connection on every inbound is closed at once: TCP connections, SOCKS UDP associations, Hysteria 2 circuits and TUN flows alike. etemenanki-app does not wait for them to finish, and clients see the connection drop, whether or not their inbound’s settings changed.
  • Every balancer prober stops, after finishing a probe it has already started. The new generation’s probers start with every member healthy and probe at once.
  • The DNS resolver and its cache go away. The new generation starts with an empty cache.
  • Outbound sessions end with the connections that use them. A WireGuard or Hysteria 2 outbound in the new generation starts a fresh session, with a new handshake, when the first flow needs it.

Two inbound types hold their resources past the cancel, so that the new generation can take the same UDP port or TUN device name, and etemenanki-app waits for them before it binds anything new:

Inbound What the old generation waits for Up to Log line if it runs out
hysteria2 Closes every QUIC connection, waits for the endpoint to go idle, then waits until the UDP port can be bound again 3 s, then another 3 s hysteria2: <address> did not come free within 3s
tun Waits until the TUN device’s descriptor and every TCP flow on it are released 3 s tun: device fd or <n> tcp flows still open after 3s

If the wait runs out, the warning is logged and the new generation tries to bind anyway; if the port or device is still taken, that inbound stays down (see the next section).

While the old generation winds down and the new one binds, no inbound accepts connections. For TCP-based inbounds the gap is only the time it takes to close and reopen the listeners. With a Hysteria 2 or TUN inbound in the old configuration, it can last up to several seconds.

The new generation binds its inbounds one by one, in file order, and logs each success:

INFO etemenanki_app::instance: inbound socks-in listening on 127.0.0.1:1080

A bind failure is handled differently at startup and on reload:

When A bind fails Result
Startup The first failure The process exits with status 1: failed to start: inbound http-in bind 127.0.0.1:1080 failed: Address already in use (os error 98)
Reload Any inbound Logged as inbound http-in bind 127.0.0.1:1080 failed: Address already in use (os error 98). That inbound stays down; the other inbounds start and serve.

On reload, one bad port should not take the whole proxy down, so etemenanki-app keeps going. The configuration with the failed inbound becomes the running configuration all the same, and its bytes are remembered. The failed inbound is not retried on its own: once you have freed the address or fixed the port, save the config file again with a change to its bytes.

--test cannot catch bind failures, because it binds nothing. Typical causes are a port already used by another process, two inbounds in the same file on the same address and port, a port below 1024 without the privilege to bind it, and a TUN device that cannot be created.

A successful reload applies everything in the file except the log level:

Setting Applied on reload
[[inbound]]: every setting, including address, port, protocol, users, transport and TLS Yes
[[outbound]]: every setting Yes
[[balancer]]: outbounds, strategy, probe_interval, probe_timeout Yes (logged as no changes if nothing else changed)
[route]: default, rules, geoip and geosite paths Yes
[dns]: every setting Yes (logged as no changes if nothing else changed); the cache starts empty
Files the config refers to: certificates, keys, CA files, geodata Yes, read again on every successful reload
[log].level No. Reported as log changed, not applied.
The RUST_LOG environment variable No. Read once at startup.
The config file path (-c) No. Fixed for the life of the process.

The log level is set once, when the process starts: from RUST_LOG if it is set and valid, otherwise from [log].level, otherwise info. To change it, restart the process. Relative paths in the file are resolved against the process’s working directory on every reload, just as at startup.

etemenanki-app reads certificates, keys, CA files and geodata files only while building a generation. Replacing one of them on disk does not start a reload, and the running generation keeps using the copy it loaded. To pick up the new file, change the config file after the new file is in place.

A convenient way is a dedicated comment line that your renewal or update job rewrites. For example, keep this as the first line of the config:

/etc/etemenanki/config.toml
# reload-stamp: 0

and run this after each certificate renewal or geodata update:

Terminal window
sed -i "s/^# reload-stamp:.*/# reload-stamp: $(date +%s)/" /etc/etemenanki/config.toml

GNU sed -i writes a new file in the same directory and renames it over the old one, so the proxy never reads a half-written file. The reload replaces the generation, with the connection drop that implies.

Two things can go wrong when you edit the file that a running proxy watches:

  • A half-written file. The watcher reads the file within about 200 ms of any event, and after the first event it reads it about five times a second. If your tool writes the file in several steps, or you copy it over a slow link, etemenanki-app may read a truncated file. A truncated file is usually rejected, but one that happens to end between two tables can still be valid, and it would then be applied without the tables after the cut.
  • A mistake you only notice after saving. An invalid file is rejected safely, but a valid file that does the wrong thing, such as a route pointing at the wrong outbound, is applied at once.

Write the new version next to the live file, check it, and move it into place in one step. The commands assume the etemenanki service user and the /etc/etemenanki layout from Running in production:

  1. Copy the live config, keeping its owner and mode, and edit the copy in the same directory:

    Terminal window
    sudo cp -p /etc/etemenanki/config.toml /etc/etemenanki/config.toml.new
    sudoedit /etc/etemenanki/config.toml.new

    Writing the copy wakes the watcher, but the live file’s bytes are unchanged, so nothing happens. -p matters: a copy that the service user cannot read is refused later with reload: cannot read …: Permission denied (os error 13).

  2. Check the copy with --test, as the service user and from the working directory the service uses, so that file permissions and relative paths are the same as in a reload:

    Terminal window
    sudo -u etemenanki sh -c 'cd /etc/etemenanki && exec /usr/local/bin/etemenanki-app --test -c config.toml.new'

    It prints Configuration OK. on success. --test runs every check a reload runs except binding listeners and creating TUN devices, so a port that is already taken still passes.

  3. Rename the copy over the live file:

    Terminal window
    sudo mv /etc/etemenanki/config.toml.new /etc/etemenanki/config.toml

    A rename within one directory is atomic: the proxy reads either the old file or the new one, never a mix. Do not cp over the live file, which writes it in place.

  4. Check the log for the reload line and one listening on line per inbound:

    Terminal window
    journalctl -u etemenanki -n 20

    Look in particular for bind … failed, which means an inbound is down.

Symptom Cause Fix
Saving the file changes nothing, and no reload line appears The bytes did not change, or the same broken bytes were already rejected, or the watcher is off (config hot-reload disabled) Make a real change; look for an earlier reload: … failed line; restart if the watcher is off
A renewed certificate is not served Certificate files do not trigger a reload Change the config file after the certificate is in place
reload: build failed, keeping current config: No such file or directory (os error 2) A file the config refers to is missing, or a relative path resolves against a different working directory Fix the path, then change the config file again so it is retried
config reload: no changes, and every client reconnected Every successful reload replaces the generation, even for a comment, a [dns] or [[balancer]] change, or a restored file Expected. Batch edits and save once
One inbound missing after a reload, inbound … bind … failed in the log The new address is taken, possibly by another inbound in the same file Free the address or change the port, then save the file again
log changed, but the log level is the same [log].level is read only at startup Restart the process
reload: cannot read …: Permission denied (os error 13) The new file’s owner or mode does not let the service user read it Fix the owner or mode; the change retries the read on its own
The process exited after SIGHUP, for example from an ExecReload=/bin/kill -HUP $MAINPID line in the unit etemenanki-app has no SIGHUP handler, so the signal terminates it Remove the ExecReload= line and save the file instead; see Running in production