Skip to content

etemenanki-app: running, reloading and shutting down

Source files: 25 · checked against Etemenanki 555b7df
  • Etemenanki/app/src/main.rs
  • Etemenanki/app/src/instance.rs
  • Etemenanki/app/src/lib.rs
  • Etemenanki/app/src/config.rs
  • Etemenanki/app/src/api.rs
  • Etemenanki/supervisor/src/lib.rs
  • Etemenanki/supervisor/src/supervisor.rs
  • Etemenanki/supervisor/src/build/apply.rs
  • Etemenanki/supervisor/src/topology/spec_plan/plan.rs
  • Etemenanki/supervisor/src/topology/inbound/mod.rs
  • Etemenanki/supervisor/src/system/listener.rs
  • Etemenanki/supervisor/src/entity/id.rs
  • Etemenanki/supervisor/src/policy.rs
  • Etemenanki/supervisor/src/track/sampler.rs
  • Etemenanki/webclient/src/lib.rs
  • Etemenanki/webclient/src/listen.rs
  • Etemenanki/webclient/src/router.rs
  • Etemenanki/app/tests/integration/e2e_reload.rs
  • Etemenanki/app/tests/integration/e2e_api.rs
  • Etemenanki/app/tests/integration/e2e_hysteria_inbound.rs
  • Etemenanki/app/tests/integration/e2e_unix.rs
  • Etemenanki/app/tests/integration/e2e_tun.rs
  • Etemenanki/app/tests/support/mod.rs
  • Etemenanki/app/tests/unit/api.rs
  • Etemenanki/supervisor/tests/hot_swap.rs

etemenanki-app is a thin binary around one supervisor. app/src/main.rs parses the command line, sets up logging, starts an Instance, starts the REST API when the config has an [api] section, watches the config file and the subscribe file it names, and waits for a signal. app/src/instance.rs owns the load itself: it reads the files, merges and lowers them into a Spec, hands the spec to the supervisor, logs what happened and remembers the sources it last recorded. Every reload goes through the supervisor’s validate, plan, prepare and commit. Sessions are not tied to their listener, so a reload does not close them just because it ran, and a refused reload leaves everything that runs as it was.

This page follows the process from main to exit: the CLI, tracing, --test, the start, the file watcher, Core::reload_with step by step with every log line it can print, what a reload keeps, the API follower and shutdown. It is for changing main.rs or instance.rs, building another front end on Core, or explaining what a reload printed. How a config becomes a spec is on the config page, the subscribe merge on the subscribe page, and the API’s endpoints on the REST API page. What an apply does inside the supervisor is on Planning and applying a change.

Component Owns Does not
app/src/main.rs → main Parsing Args, init_tracing, the --test branch, Instance::start, the first API start, spawning the API follower and the watcher, waiting for SIGINT or SIGTERM, and the shutdown order Read any config key other than [log].level
app/src/main.rs → spawn_watcher, sync_watches, watch_dir The notify watcher, the set of directories it watches, and the task that turns the watcher’s ticks into Instance::reload calls Decide whether the files changed; Core::reload_with does
app/src/main.rs → follow_api, start_api, Api Starting, stopping and rebinding the REST API as the running config’s [api] changes Decide when [api] changed; Core publishes that after an apply
app/src/instance.rs → Core Starting a supervisor from a Read, reloading it through a Lowering, skipping sources equal to the ones it last recorded, logging each outcome, and publishing the running config’s API options Know where the files come from
app/src/instance.rs → Instance The config path, reading the sources from disk, and the list of files worth watching Watch anything itself
app/src/instance.rs → check, check_bytes The dry run behind --test and behind every API route edit Bind a listener or start a task
etemenanki_supervisor::Supervisor Validation, planning, prepare and commit, listeners, sessions, outbounds, shutdown Know about files, TOML, the API or signals: the crate “knows nothing of config files, their formats, or how a reload is triggered” (supervisor/src/lib.rs)

app/src/lib.rs makes the split explicit: the library crate is “the TOML front end of Etemenanki” (api, config, instance, lower, routes, subscribe, transport), and “the etemenanki-app binary is the thin glue on top: CLI, tracing, the file watcher that triggers instance::Instance::reload, and, with an [api] section, the REST API”.

The binary does not:

  • install a SIGHUP handler. A reload is triggered by the file watcher, by POST /v1/reload or by PUT /v1/routes/{group} on the REST API, and by nothing else. SIGHUP keeps its default action and terminates the process at once, without the shutdown sequence (a shell reports exit status 129).
  • write logs anywhere but standard output. tracing_subscriber::fmt() writes to stdout by default, and init_tracing does not change the writer.
  • change any supervisor setting that outlives a spec. It starts the supervisor with Supervisor::builder() and every default (see What a reload keeps).
app/src/main.rs
/// A minimal Xray-core-style proxy runtime.
#[derive(Parser)]
#[command(version, about)]
struct Args {
/// Path to the TOML configuration file.
#[arg(short, long, default_value = "config.toml")]
config: PathBuf,
/// Validate the configuration and exit without binding any listener.
#[arg(long)]
test: bool,
}

The --help output of the pinned binary:

A minimal Xray-core-style proxy runtime
Usage: etemenanki-app [OPTIONS]
Options:
-c, --config <CONFIG> Path to the TOML configuration file [default: config.toml]
--test Validate the configuration and exit without binding any listener
-h, --help Print help
-V, --version Print version

--version prints etemenanki-app 2.1.1: clap puts the binary’s name before the version in app/Cargo.toml. The REST API reports the same version at /v1/version, as the bare env!("CARGO_PKG_VERSION"): {"version":"2.1.1"}. The config path is kept exactly as given. A relative path is resolved against the working directory on every read, and the binary never changes its working directory. The user-facing reference is Command line.

app/src/main.rs
/// How long live connections get to finish on shutdown before they are
/// closed.
const SHUTDOWN_GRACE: Duration = Duration::from_secs(5);

It is used twice: as the time the REST API gets to finish its requests in flight whenever it stops (Api::stop), and as the grace passed to Instance::shutdown, which the supervisor gives live connections.

app/src/instance.rs
#[derive(Debug, thiserror::Error)]
pub enum LoadError {
/// The files could not be read, or what they say could not be lowered.
#[error(transparent)]
Config(#[from] io::Error),
/// The supervisor refused the spec.
#[error(transparent)]
Apply(#[from] ApplyError),
}
impl From<LoadError> for ControlError { /* see the table */ }

Both variants are transparent, so every log line and every API error shows the inner error’s text unchanged. From<LoadError> for ControlError sorts a failure by fault for the REST API, which answers Invalid with HTTP 422 and Failed with HTTP 500:

LoadError ControlError Why
Config(e) with e.kind() of InvalidData or InvalidInput Invalid(e.to_string()) The config’s content is at fault: a TOML error (config::parse_bytes uses InvalidData), a lowering refusal or an [api] error (InvalidInput)
Config(e) of any other kind Failed(e.to_string()) A file that cannot be read, for example NotFound or PermissionDenied
Apply(ApplyError::Bind { .. }) or Apply(ApplyError::Stopped) Failed(e.to_string()) Not the config’s fault: a port held elsewhere, or a supervisor that has shut down
Apply(e), every other variant Invalid(e.to_string()) Validation, a build failure or a disruptive change

The ApplyError variants and their texts are listed on Validation and apply errors.

app/src/instance.rs
pub struct Built {
pub spec: Spec<UserKey>,
pub api: Option<ApiOptions>,
}
pub fn build(cfg: &Config) -> io::Result<Built> {
Ok(Built {
spec: lower(cfg)?,
api: crate::api::options(cfg)?,
})
}
pub type Lowering = Arc<dyn Fn(&Config) -> io::Result<Built> + Send + Sync>;

A Built is everything one config produces: the spec the supervisor runs, and the REST API’s options when the config has an [api] section. build is the app’s Lowering: lower (on the config page) followed by api::options. That returns None without an [api] section. Otherwise it parses [api].listen and applies ApiOptions::new, and it wraps any OptionsError as an InvalidInput error prefixed with [api] :

Rule Checked by Error text
listen defaults to DEFAULT_LISTEN, 127.0.0.1:9090 ApiConfig’s serde default default_api_listen (app/src/config.rs) none
A listen containing / is a Unix socket path; anything else must parse as ip:port Listen::from_str [api] listen "<s>": expected ip:port, or a path (containing a /) for a unix socket
A secret that is set must not be empty, whatever the listen ApiOptions::new [api] secret is empty
A listen that is neither a loopback address nor a Unix socket needs a secret ApiOptions::new (Listen::is_local) [api] listen <addr>: an API reachable from other hosts needs a secret; listen on a loopback address or a unix socket, or set one

An [api] error is therefore a lowering error, found by --test and reported by a reload like any other config mistake. api_options_come_from_the_api_section and an_api_open_to_other_hosts_needs_a_secret (app/tests/unit/api.rs) pin these rules. Listen, ApiOptions and OptionsError are described in full on the REST API page.

Lowering is a shared closure rather than a function so that a Core can be started with a different one. The app passes Arc::new(build). A front end handed the files’ contents rather than paths, such as the mobile FFI, runs a Core of its own with its own Lowering (see the FFI page).

app/src/instance.rs
pub type Read = io::Result<(Sources, io::Result<Config>)>;
app/src/config.rs
#[derive(Clone, PartialEq, Default)]
pub struct Sources {
pub config: Vec<u8>,
pub subscribe: Option<Vec<u8>>,
}
pub fn read_sources(path: &Path) -> io::Result<(Sources, io::Result<Config>)>;
pub fn sources_from(config: Vec<u8>) -> io::Result<(Sources, io::Result<Config>)>;
pub fn effective(parsed: io::Result<Config>, sources: &Sources) -> io::Result<Config>;

A Read is one reading of the files a Core loads. The outer io::Result is the reading itself, which fails when the files cannot be gathered into Sources (the cases are on the config page). The inner one is the config’s parse, handed back with the bytes because the subscribe file’s path is known only once the config has parsed. A config that does not parse is still a set of sources; its parse error comes out of config::effective at the lowering step. Sources holds the raw bytes of the config file and of the subscribe file when the config names one, and compares them byte for byte. The three functions are described on the config page; effective merges the subscribe file (see the subscribe page).

app/src/instance.rs
#[derive(Debug)]
pub enum Reload {
/// The files changed, and were applied.
Applied(ApplyReport),
/// The files are byte for byte the ones last tried, and were not tried
/// again. `failed` is why that last try failed, if it did.
Unchanged { failed: Option<String> },
}
impl Reload {
pub fn report(self) -> Result<ReloadReport, ControlError>;
}

report is what an API client is told:

Reload report() HTTP answer
Applied(report) Ok(ReloadReport::Applied(report)) 200, {"status":"applied", "built":[…], "swapped":[…], "drained":[…], "rebound":[…], "restarted":[…], "removed":[…], "reused":[…]}
Unchanged { failed: None } Ok(ReloadReport::Unchanged) 200, {"status":"unchanged"}
Unchanged { failed: Some(reason) } Err(ControlError::Invalid(...)) with the text <reason> (the files have not changed since this was found) 422

The doc comment on ReloadReport::Unchanged in webclient/src/lib.rs names the common case: “the file watcher may have reloaded them already”. Sources recorded together with a failure are refused again with that failure’s text, as they would be if tried.

app/src/instance.rs
pub struct Core {
supervisor: Supervisor<UserKey>,
lowering: Lowering,
api: tokio::sync::watch::Sender<Option<ApiOptions>>,
last: tokio::sync::Mutex<Last>,
}
struct Last {
sources: Sources,
failed: Option<String>,
}
impl Core {
pub async fn start(
builder: SupervisorBuilder<UserKey>,
read: Read,
lowering: Lowering,
) -> Result<Self, LoadError>;
pub fn supervisor(&self) -> &Supervisor<UserKey>;
pub fn api(&self) -> tokio::sync::watch::Receiver<Option<ApiOptions>>;
pub fn tracker(&self) -> Tracker;
pub async fn reload_with(&self, read: impl FnOnce() -> Read) -> Result<Reload, LoadError>;
pub async fn shutdown(&self, grace: Duration);
}
Field Type Role
supervisor Supervisor<UserKey> The running supervisor. UserKey is UserName, the name the app keys a config user by (see the config page).
lowering Lowering How a merged Config becomes a Built, fixed for the Core’s life
api tokio::sync::watch::Sender<Option<ApiOptions>> The API options of the config running now: None while it has no [api]. Changed only by an applied reload.
last tokio::sync::Mutex<Last> The last recorded sources, and the text of the failure recorded with them, if any. The lock is held for a whole reload, which serialises reloads.

Core is the load itself, wherever the files come from. It holds one Supervisor handle. Supervisor is Clone, and every clone drives the same actor over its command channel; when the last handle is dropped without shutdown, the actor sees the channel close and runs its shutdown with Duration::ZERO. A command whose send fails, or whose reply is dropped, comes back as ApplyError::Stopped (Supervisor::ask). The details are on The supervisor. api() subscribes a new receiver to the watch; the receiver starts having seen the current value, so its changed() waits for the next change. tracker() clones the supervisor’s Tracker, which the REST API reads without going through the supervisor’s actor (see Tracking flows, stats and speed limits). shutdown(grace) is Supervisor::shutdown(grace).

app/src/instance.rs
pub struct Instance {
path: PathBuf,
core: Core,
}
impl Instance {
pub async fn start(path: PathBuf) -> Result<Self, LoadError>;
pub fn api(&self) -> tokio::sync::watch::Receiver<Option<ApiOptions>>;
pub fn config_path(&self) -> &Path;
pub fn tracker(&self) -> Tracker;
pub fn watched_files(&self) -> Vec<PathBuf>;
pub async fn reload(&self) -> Result<Reload, LoadError>;
pub async fn shutdown(&self, grace: Duration);
}

Instance is the app’s Core, reading its files from disk:

  • start(path) is Core::start(Supervisor::builder(), config::read_sources(&path), Arc::new(build)).
  • reload() is core.reload_with with a read closure that calls config::read_sources(&self.path) and rewrites a read error as io::Error::new(e.kind(), format!("cannot read {}: {e}", ...)). The new error keeps the io::ErrorKind, so ControlError still sorts it by kind; the original error survives only as text.
  • watched_files() returns the config path and, when the config file on disk parses and its [subscribe] names a path, that path. It reads the file with config::load rather than asking the running config, “so a newly named subscribe file is followed even while the config naming it fails to apply”. A config that does not parse yields the config path alone.
  • config_path() is the -c path as given. The REST API’s ConfigFile uses it to read and rewrite the file.
app/src/instance.rs
fn summary(report: &ApplyReport) -> String;

summary turns an ApplyReport into the text after config loaded: and config reloaded: . It prints each non-empty group as <label> [<first>, <second>, …], joins the groups with ; , and uses this order:

Order Label Items Printed as
1 built Resource dns, route, outbound <tag>@v<n>, balancer <tag>, user set <tag>, inbound <tag>
2 swapped inbound tag <tag>
3 drained OutboundId <tag>@v<n>
4 rebound inbound tag <tag>
5 restarted inbound tag <tag>
6 removed inbound tag <tag>
7 reused Resource as for built

When every group is empty it prints nothing to run. reused comes last although it is the first field of ApplyReport, so the line leads with what changed. ApplyReport’s doc comment states that every resource of the desired spec appears in reused or built; a start has nothing to reuse, so its line lists every resource under built. Because the app never allows a disruptive change (see Disruptive changes are refused), its lines never have a restarted group. What each group means is on Planning and applying a change. Lines captured from the pinned debug binary (timestamps removed):

INFO etemenanki_app::instance: config loaded: built [dns, outbound direct@v1, route]
INFO etemenanki_app::instance: config loaded: built [dns, outbound direct@v1, outbound n@v1, route]
INFO etemenanki_app::instance: config reloaded: reused [dns, outbound direct@v1, route]
INFO etemenanki_app::instance: config reloaded: built [outbound block@v1]; reused [dns, outbound direct@v1, route]
INFO etemenanki_app::instance: config reloaded: built [outbound block2@v1]; drained [block@v1]; reused [dns, outbound direct@v1, route]
INFO etemenanki_app::instance: config reloaded: built [dns, outbound direct@v2]; drained [direct@v1]; reused [route]
INFO etemenanki_app::instance: config reloaded: drained [n@v1]; reused [dns, outbound direct@v1, route]

The second line is a start with a subscribe file that defines a node n. The third is the reload after an edit that only added a comment: the sources changed, so the spec was applied, and the supervisor reused everything. The sixth is a change to [dns], which rebuilds the freedom outbound as a new version as well. The last is the config dropping its [subscribe] section. At debug level, every commit also logs applied: <n> built, <n> reused, <n> swapped, <n> drained from the target etemenanki_supervisor::supervisor.

The supervisor logs two more lines of its own around these. Every listener that a start or an applied reload binds logs inbound <tag> listening on <bind> at INFO (target etemenanki_supervisor::system::listener) when Listener::start runs during the commit, so on a start these lines come before config loaded:. A TUN inbound that creates its device logs inbound <tag> owns tun device <name> when it is bound. Both are described on Listeners and the serve loop.

flowchart LR
  M["main"] --> ST["Instance::start"]
  ST --> CS["Core::start"]
  CS --> SUP["supervisor actor"]
  W["watcher task"] --> IR["Instance::reload"]
  RP["POST /v1/reload"] --> IR
  PR["PUT /v1/routes/{group}"] --> IR
  IR --> RW["Core::reload_with"]
  RW -- "apply" --> SUP
  RW -- "send_if_modified" --> AW["api watch"]
  AW --> FA["follow_api"]
  FA --> API["API server task"]
  M -- "signal" --> FA
  M -- "Instance::shutdown" --> SUP

One supervisor runs for the life of the process. Three callers reload it, all through Instance::reload and Core::reload_with, which applies to the supervisor and then publishes the running config’s API options. The API follower is the only code that starts, stops or rebinds the REST API. On a signal, main stops the follower first and the supervisor second. The sections below take these in order: Starting, The file watcher, Reloading, What a reload keeps, The API follower and Shutting down.

flowchart TB
  A["Args::parse"] --> T["init_tracing"]
  T --> Q{"--test?"}
  Q -- "yes" --> C["instance::check, print the verdict, exit 0 or 1"]
  Q -- "no" --> S["Instance::start"]
  S -- "Err" --> F1["log failed to start, exit 1"]
  S -- "Ok" --> P{"running config has an api section?"}
  P -- "yes" --> SA["start_api"]
  SA -- "Err" --> F2["log, shutdown with zero grace, exit 1"]
  SA -- "Ok" --> FA["spawn follow_api"]
  P -- "no" --> FA
  FA --> W["spawn_watcher"]
  W --> SIG["wait for SIGINT or SIGTERM"]

main runs on #[tokio::main], the multi-threaded runtime.

app/src/main.rs
fn init_tracing(config_path: &Path) {
use tracing_subscriber::EnvFilter;
let filter = EnvFilter::try_from_default_env().unwrap_or_else(|_| {
let level = config::load(config_path)
.ok()
.and_then(|c| c.log.level)
.unwrap_or_else(|| "info".to_string());
EnvFilter::new(level)
});
tracing_subscriber::fmt().with_env_filter(filter).init();
}
  1. EnvFilter::try_from_default_env() reads RUST_LOG. When it is set and parses, it wins.
  2. Otherwise, whether RUST_LOG is unset or set to something that does not parse (both are an Err), the config file is read and parsed once with config::load (the file as written, without the subscribe merge), and [log].level becomes the filter through EnvFilter::new, which takes the same directive syntax as RUST_LOG (for example info,etemenanki_supervisor=debug).
  3. When the key is absent, or the file cannot be read or does not parse, the filter is info.
  4. init() installs the global subscriber, which formats to standard output.

This runs once, before --test and before the start. Nothing later replaces the filter, so [log].level is not reloaded: a reload that changes only [log].level finds the sources changed, applies a spec in which nothing changed (config reloaded: reused [...]), and logs at the old level until the process restarts. The --test verdict on failure is a log line and obeys the same filter; the exit status is always set.

The log targets are the module paths: etemenanki_app for the lines main.rs prints (failed to start, configuration invalid, shutting down, the API and watcher lines) and etemenanki_app::instance for the lines instance.rs prints (config loaded, config reloaded, every reload line).

app/src/instance.rs
pub async fn check(path: &Path) -> Result<(), LoadError> {
check_bytes(std::fs::read(path)?).await
}
pub async fn check_bytes(config: Vec<u8>) -> Result<(), LoadError> {
let (sources, parsed) = config::sources_from(config)?;
let built = build(&config::effective(parsed, &sources)?)?;
etemenanki_supervisor::check(&built.spec).await?;
Ok(())
}

--test reads the config, reads the subscribe file it names, merges, lowers (including [api]) and runs etemenanki_supervisor::check. That builds a fresh Actor::new(SocketOptions::default()), plans against its empty running state, and runs prepare(&plan, spec, false): with bind false, every outbound, route, handler and user table is still built, and only listener::bind and Pending::pair are skipped. Because it plans against an empty state, it never reports a disruptive change. It refuses what a start would refuse, short of a port, socket path or device that cannot be bound and of the checks pairing makes, and it never binds the REST API. What it checks in detail is on the config page and in the dry-run section of Planning and applying a change.

Result Output Exit status
Accepted Configuration OK. on standard output (println!, not a log line) 0
Refused configuration invalid: <e> at ERROR, target etemenanki_app 1

Captured from the pinned binary (timestamps removed):

ERROR etemenanki_app: configuration invalid: route references unknown outbound nowhere
ERROR etemenanki_app: configuration invalid: [subscribe] needs a path naming the subscribe file
ERROR etemenanki_app: configuration invalid: [api] listen 0.0.0.0:9090: an API reachable from other hosts needs a secret; listen on a loopback address or a unix socket, or set one
ERROR etemenanki_app: configuration invalid: No such file or directory (os error 2)
ERROR etemenanki_app: configuration invalid: subscribe file missing-sub.toml: No such file or directory (os error 2)

A config file that cannot be read gives the bare io::Error, with no path: check calls std::fs::read(path)? directly, unlike a reload’s cannot read <path>: . A missing subscribe file is named by config::sources_from.

check_bytes is also what the REST API runs on an edited config before it writes the file (app/src/api.rs → ConfigFile::set_route).

Instance::start(path) calls Core::start(Supervisor::builder(), config::read_sources(&path), Arc::new(build)). Core::start:

  1. Returns the read error unchanged, if read_sources failed. Unlike a reload’s, it carries no cannot read <path>: prefix.
  2. Runs config::effective and the lowering. A merge or lowering error is returned as LoadError::Config.
  3. Calls builder.start(built.spec). That asserts the sample interval is non-zero (it panics otherwise; the app’s is the default 1 s) and runs one apply against a new, empty actor: validate, plan, prepare (binding every listener) and commit, during which each new listener logs inbound <tag> listening on <bind>. Any ApplyError is returned, a Bind error included, so a start is all or nothing. A failed prepare drops what it built, which releases every listener it had bound. Only after the apply succeeds does it spawn the sampler (and a usage sink, which the app does not have), create the command channel with mpsc::channel(16) and spawn actor.run. The builder is described on The supervisor.
  4. Logs config loaded: <summary> at INFO.
  5. Stores the supervisor and the lowering, creates the api watch with built.api, and records the sources in last with failed: None. A reload of the same files right after the start is therefore Unchanged.

main logs any error as failed to start: <e> at ERROR and exits with status 1. Captured:

ERROR etemenanki_app: failed to start: route references unknown outbound nowhere
ERROR etemenanki_app: failed to start: No such file or directory (os error 2)
ERROR etemenanki_app: failed to start: subscribe file missing-sub.toml: No such file or directory (os error 2)

A listener that cannot be bound gives failed to start: inbound <tag>: binding <bind> failed: <e>, where <bind> is the BindSpec’s Display. The app’s lowering produces four of its five forms: <host>:<port>, udp <host>:<port>, unix:<path> and tun <name> (tun auto without a name). The fifth, tun fd <n>, is a supplied descriptor (TunSource::Fd), which only the FFI produces.

main takes instance.api(), clones the current value and, when it is Some, calls start_api:

app/src/main.rs
async fn start_api(instance: Arc<Instance>, options: &ApiOptions) -> io::Result<Api>;
struct Api {
stop: oneshot::Sender<()>,
task: JoinHandle<io::Result<()>>,
}

start_api binds ApiListener::bind(&options.listen), logs API listening on <addr> at INFO (the bound TCP address from local_addr(), or the listen value for a Unix socket), takes tracker = instance.tracker(), builds etemenanki_webclient::router(ConfigFile::new(instance), tracker, options.secret.clone(), env!("CARGO_PKG_VERSION")), and spawns listener.serve(app, stopped), where stopped is the receiving half of a fresh oneshot. How ApiListener binds a TCP address or a Unix socket path, and when it removes the socket file, is on the REST API page.

When the first bind fails, main logs failed to start: API on <listen>: <e>, calls instance.shutdown(Duration::ZERO) and exits with status 1. By then the supervisor was already running and its inbounds accepting; the zero grace closes whatever they accepted at once. A config without [api] binds nothing beyond its inbounds.

main then creates the oneshot pair (stop_api, api_stopped), spawns follow_api(instance.clone(), options, running, api_stopped) (see The API follower), and calls spawn_watcher. If the watcher cannot be set up, it logs config hot-reload disabled: <e> at ERROR and carries on: the proxy runs, without reloads from the watcher. Last, it awaits wait_for_shutdown.

app/src/main.rs
fn watch_dir(path: &Path) -> PathBuf;
fn sync_watches(
watcher: &mut notify::RecommendedWatcher,
watched: &mut HashSet<PathBuf>,
instance: &Instance,
);
fn spawn_watcher(instance: Arc<Instance>) -> notify::Result<()>;

spawn_watcher creates a notify::RecommendedWatcher and points it at watch_dir of each path Instance::watched_files() returns, each with RecursiveMode::NonRecursive: it watches the directories that hold the files, not the files themselves. The watcher’s callback feeds () ticks into a channel, and one task turns those ticks into Instance::reload calls. Whether anything changed is decided by Core::reload_with, which re-reads the files by path and compares their bytes.

watch_dir(path) is the file’s parent directory, or . when the path has no parent component (-c config.toml).

spawn_watcher:

  1. Creates an unbounded tokio::sync::mpsc channel of ().
  2. Creates the RecommendedWatcher with notify::recommended_watcher, whose callback sends () ticks into the channel.
  3. Watches watch_dir(instance.config_path()). A failure here, or in step 2, returns the notify::Error, and main logs config hot-reload disabled: <e>: the config’s own directory must be watchable, “or there is no hot reload at all”.
  4. Starts watched as a HashSet holding that directory, and calls sync_watches once, which adds the subscribe file’s directory when the config names one elsewhere.
  5. Spawns the watcher task, moving the watcher, the receiver and watched into it, and returns Ok(()).

The RecommendedWatcher lives inside the task. The task owns the watcher, the watcher’s closure owns the sender, so the channel never closes and the task runs until the runtime shuts down when main returns.

flowchart TB
  N["watcher callback sends a tick"] --> R["watcher task: rx.recv()"]
  R --> D1["drain the ticks already queued"]
  D1 --> S["settle briefly"]
  S --> D2["drain again"]
  D2 --> L["instance.reload()"]
  L --> SW["sync_watches"]
  SW --> R
  • One reload per pass. After it takes a tick, the task drains the ticks already queued, sleeps for a short settle (an inline tokio::time::sleep, not a named constant), drains again, and makes one instance.reload() call for all the ticks it drained.
  • One reload at a time. The task awaits instance.reload() before it takes the next tick, so its own reloads never overlap. Reloads from the REST API are serialised against the watcher’s by the last lock in Core.
  • The result is ignored. reload_with has already logged the outcome. The watcher does not report Unchanged, which is silent.
  • Re-sync after every pass. sync_watches runs after each reload, whatever its outcome, so the set of watched directories follows the config file on disk.

sync_watches computes the wanted set from instance.watched_files(), mapped through watch_dir:

  1. For each directory wanted and not yet watched, it calls watcher.watch(dir, RecursiveMode::NonRecursive). A failure is logged as cannot watch <dir> for changes: <e> at ERROR and is not fatal.
  2. For each directory watched and no longer wanted, it calls watcher.unwatch(dir) and ignores the result.
  3. It stores the wanted set as the watched set.

The config’s own directory is always wanted, because watched_files always starts with the config path. The subscribe file’s directory is wanted while the config file on disk names it. Because watched_files reads the file on disk and not the running config, a config that starts naming a new subscribe file has that file’s directory watched from the next pass on, even while that config fails to apply.

app/src/instance.rs
impl Core {
pub async fn reload_with(&self, read: impl FnOnce() -> Read) -> Result<Reload, LoadError>;
}
impl Instance {
pub async fn reload(&self) -> Result<Reload, LoadError>;
}

Three callers reload an Instance, and all of them go through Instance::reload:

Caller Path What it does with the result
The watcher task spawn_watcher Ignores it
POST /v1/reload webclient/src/router.rs → reload → ConfigFile::reload instance.reload().await?.report(), answered as JSON or as an error status
PUT /v1/routes/{group} webclient/src/router.rs → set Holds the router’s edits lock across ConfigFile::set_route (which checks and writes the file) and the ConfigFile::reload after it, and answers with the reload’s result and the file’s new ETag

The router’s edits lock serialises a PUT with the reload that follows it, but, as ConfigControl’s doc comment says, not against reloads from elsewhere, such as the watcher; Core’s last lock does that. A reload that reads the bytes another reload last recorded returns Unchanged. The REST API answers that with 200 {"status":"unchanged"} when no failure was recorded with those bytes, and with 422 when one was. set puts the file’s new ETag on its response either way, because the file was written whether or not the reload after it succeeded.

The edit itself (ConfigFile::set_route: the If-Match check, the check_bytes dry run, the re-read that refuses with Stale when a hand edit landed during the check, and the atomic write) is on the REST API page.

sequenceDiagram
  participant C as caller
  participant K as Core::reload_with
  participant R as read closure
  participant S as Supervisor
  participant A as api watch
  C->>K: reload_with(read)
  K->>K: lock last
  K->>R: read()
  R-->>K: Sources and the parse
  alt the read failed
    K-->>C: Err(Config), logged as reload plus the error
  else sources equal last.sources
    K-->>C: Ok(Unchanged) with the recorded failure
  else the sources changed
    K->>K: config::effective, then the lowering
    alt merge or lowering failed
      K-->>C: Err(Config), the running config kept
    else built
      K->>S: apply(spec)
      alt applied
        S-->>K: ApplyReport
        K->>A: send_if_modified(built.api)
        K->>K: last becomes these sources, no failure
        K-->>C: Ok(Applied(report))
      else refused
        S-->>K: ApplyError
        K-->>C: Err(Apply), the running config kept
      end
    end
  end
  1. Lock. self.last.lock().await. The guard is held to the end of the call, across the read and across the supervisor’s apply. The read, the comparison, the apply and the update of last are therefore one step: two callers cannot both read one change and apply it twice, and a slower caller cannot apply older files after newer ones. The doc comment says it: “Load what read gives, read while no other reload runs.”
  2. Read. read() runs under the lock. For Instance::reload it is config::read_sources(&self.path), with a read error rewritten as cannot read <path>: <e>. A read error is logged as reload: <e> at ERROR (for the app, reload: cannot read <path>: <e>) and returned as LoadError::Config.
  3. Compare. When sources == last.sources, the call returns Ok(Reload::Unchanged { failed: last.failed.clone() }) without logging anything and without touching the supervisor.
  4. Merge and lower. config::effective(parsed, &sources).and_then(|cfg| (self.lowering)(&cfg)). This is where a TOML error surfaces, together with a subscribe-file error, an unknown protocol, a bad settings table and an [api] mistake. On error the call logs reload: <e>; keeping the running config at ERROR and returns LoadError::Config.
  5. Apply. self.supervisor.apply(built.spec), which is apply_with and ApplyOptions::default(): allow_disruptive is false. The supervisor’s actor validates, plans, refuses a disruptive plan, prepares and commits (see Planning and applying a change).
  6. Applied. The call logs config reloaded: <summary> at INFO, publishes the API options with send_if_modified (below), sets last to these sources with failed: None, and returns Reload::Applied(report).
  7. Refused. ApplyError::Disruptive { inbound, reason } is logged as reload refused, keeping the running config: inbound <inbound>: <reason>, which would end its live connections; restart to apply it. Every other ApplyError is logged as reload: <e>; keeping the running config. Both are at ERROR, and the error is returned as LoadError::Apply. The supervisor refuses a spec as a whole, so what runs is exactly what ran before; a_refused_spec_changes_nothing pins that for a validation error and for a bind that fails while preparing.
app/src/instance.rs
self.api.send_if_modified(|api| {
let changed = *api != built.api;
*api = built.api;
changed
});

The watch value is replaced after every applied reload, but receivers are woken only when the new Option<ApiOptions> differs from the old one: a different listen, a different secret, or the section appearing or going. A reload that is refused or fails to lower publishes nothing, so “one that is refused changes the API no more than the rest”. The binary acts on the change in follow_api, not here (see The API follower).

last holds the sources of the last recorded reload and that reload’s failure text, if it had one. The comparison in step 3 is Sources’ derived PartialEq: the config file’s bytes and, when the config names one, the subscribe file’s bytes. Only the bytes count:

  • A reload whose sources equal last.sources returns Unchanged { failed } without touching the supervisor: silent, with no parse beyond the one read_sources does, and no apply.
  • An edit to the subscribe file alone is a change, and so is an edit that only adds a comment or reorders keys. Such a reload is lowered and applied like any other (the third captured line in summary).
  • last starts as the sources Core::start loaded, with no failure.

Unchanged { failed } carries the failure text recorded with those sources, if any. The text is the error’s Display, not the log line; ApplyError::Disruptive, for example, displays as inbound <tag>: <reason>; this ends its live connections and needs allow_disruptive. The watcher ignores Unchanged; the REST API answers 422 with that text followed by (the files have not changed since this was found).

Case Log line (ERROR unless noted) reload_with returns REST API answer
The files cannot be read reload: cannot read <path>: <e> Err(LoadError::Config) 422 or 500, by the error’s kind
Sources equal to the recorded ones, no failure recorded none Ok(Unchanged { failed: None }) 200 {"status":"unchanged"}
Sources equal to the recorded ones, a failure recorded none Ok(Unchanged { failed: Some(..) }) 422 <reason> (the files have not changed since this was found)
TOML, merge or lowering error reload: <e>; keeping the running config Err(LoadError::Config) 422 for InvalidData and InvalidInput, 500 for other kinds
A disruptive change reload refused, keeping the running config: inbound <tag>: <reason>, which would end its live connections; restart to apply it Err(LoadError::Apply(Disruptive)) 422
A listener that cannot be bound reload: inbound <tag>: binding <bind> failed: <e>; keeping the running config Err(LoadError::Apply(Bind)) 500
Any other refusal reload: <e>; keeping the running config Err(LoadError::Apply(..)) 422, or 500 for Stopped
Applied INFO config reloaded: <summary> Ok(Applied(report)) 200 with the report

Refusals captured from the pinned binary (timestamps removed; a TOML error is shown by its first line only):

ERROR etemenanki_app::instance: reload: route references unknown outbound nowhere; keeping the running config
ERROR etemenanki_app::instance: reload: outbound x: unknown protocol "carrier-pigeon"; keeping the running config
ERROR etemenanki_app::instance: reload: duplicate outbound tag direct; keeping the running config
ERROR etemenanki_app::instance: reload: [api] listen 0.0.0.0:9: an API reachable from other hosts needs a secret; listen on a loopback address or a unix socket, or set one; keeping the running config
ERROR etemenanki_app::instance: reload: TOML parse error at line 5, column 1

The first and third are supervisor refusals (ApplyError::UnknownReference, ApplyError::DuplicateTag); the second and fourth are lowering errors; the last is a parse error, which config::effective reports at the lowering step.

The supervisor’s plan marks some inbound changes as disruptive (Step::Disrupt in supervisor/src/topology/spec_plan/plan.rs). Supervisor::apply passes ApplyOptions::default(), whose allow_disruptive is false, so Actor::apply refuses such a spec before preparing anything and names the first disruption the plan lists (plan.disruptions().next()). Which changes the plan marks is on Planning and applying a change.

For etemenanki-app this means:

  • a Hysteria 2 listener whose obfs changes is refused on reload, and a restart applies it. katana applies with allow_disruptive, so the same change applies there (see Node manager).
  • a TUN inbound whose device or settings change is refused on reload, and a restart applies it.

The refusal is for the whole spec, as for any other ApplyError: nothing else in it is applied either. The <reason> in the log line is one of these texts:

the hysteria2 obfuscation changed; connected clients cannot follow the new key
the tun inbound's settings changed; its runtime restarts and its flows end
the tun device changed; the old device and its flows end

So a Hysteria 2 obfuscation change reads:

ERROR etemenanki_app::instance: reload refused, keeping the running config: inbound <tag>: the hysteria2 obfuscation changed; connected clients cannot follow the new key, which would end its live connections; restart to apply it

The supervisor-level refusal is pinned by a_hysteria2_obfuscation_change_needs_allow_disruptive in supervisor/tests/hot_swap.rs.

An applied reload is an apply of a whole new spec, and the supervisor keeps everything whose spec did not change. From the app’s side:

What Across an applied reload Details
Sessions Not tied to their listener: they run under the supervisor’s root token, so stopping or swapping a listener does not cancel them. Removing an inbound or renaming its tag closes its sessions (CloseSessions). The removal policy applies to sessions bound to a removed user’s principal; the app lowers Policies::default(), whose UserRemovalPolicy is Close. Users, principals and sessions, Planning and applying a change
Stream listeners A listener whose BindSpec is unchanged keeps its socket, its accept loop and the loop’s admission limits (the semaphores run_stream_inbound creates). A changed inbound on the same bind gets a new handler for new connections (swapped). A new inbound is bound during prepare. How a changed bind is handled is on the serving page. Listeners and the serve loop
A Hysteria 2 listener on the same port Keeps its QUIC endpoint and keeps serving on its UDP port. Its circuit budget carries over only while max_circuits is unchanged (same_circuits). The rest of the swap is on the serving page. Listeners and the serve loop
Outbounds An outbound whose spec is unchanged, with an unchanged DNS spec unless it is a blackhole, is reused with the state its connector holds. A changed outbound is built as a new version, and new flows use it. Outbounds, UDP fan-out and balancers, Planning and applying a change
DNS The resolvers and their answer cache are kept while the DnsSpec is unchanged Name resolution and the DNS service
Balancers Member health and running probes are kept when the balancer and all its members are unchanged Outbounds, UDP fan-out and balancers
The route table Reused when the RouteSpec is unchanged. New connections follow the new plane; open ones keep the target they were opened on. The plane: routing each flow
[api] Follows the applied config: started, stopped or rebound The API follower
[log].level Not reloaded; init_tracing runs once Tracing
Supervisor settings Fixed for the process: the builder’s defaults from Supervisor::builder(), which are a sample interval of DEFAULT_SAMPLE_INTERVAL (1 s), no usage sink and SocketOptions::default() The supervisor
Apply options Every reload applies with ApplyOptions::default() (Supervisor::apply), and the start with the builder’s default, which is the same Disruptive changes are refused

The full table, with the plan steps behind each row, is in the “What an apply keeps” section of Planning and applying a change. Two app tests pin the session rows end to end: an_open_connection_keeps_transferring_across_a_reload (a route change sends new connections to a blackhole while an open SOCKS connection keeps echoing) and a_reload_that_removes_a_user_closes_only_their_connections (removing one SOCKS account closes that account’s open connection and refuses its next login, while another account’s connection keeps echoing).

app/src/main.rs
async fn follow_api(
instance: Arc<Instance>,
mut options: watch::Receiver<Option<ApiOptions>>,
mut running: Option<Api>,
mut stopped: oneshot::Receiver<()>,
);

follow_api keeps the REST API as the running config’s [api] says, from the API main started (running) until the stopped signal:

flowchart TB
  W["select: stop signal, or options.changed()"] -- "stop signal, or the sender is gone" --> E["stop the running API, return"]
  W -- "changed" --> B["wanted = borrow_and_update()"]
  B --> S["stop the running API, if any"]
  S --> Q{"wanted?"}
  Q -- "Some" --> ST["start_api(wanted)"]
  ST -- "Ok" --> R["running = the new API"]
  ST -- "Err" --> L1["log: stays off until [api] changes"]
  Q -- "None" --> L2["log: API stopped"]
  R --> W
  L1 --> W
  L2 --> W
  • Stop before bind. The old API is stopped before the new one binds, “so the same address can be taken again”. A change of secret alone therefore closes the listener and binds the same port again. Between the two, connections to the API are refused.
  • Latest value only. A watch keeps one value. When several applied reloads change [api] while the follower is busy, it wakes once and acts on the latest options.
  • A reload made by the API finishes first. The restart runs in this task, “not in the reload that published it, so a reload made by the API itself finishes its request before the API stops”. Core::reload_with only publishes; the request that caused the reload returns its answer, and the follower’s Api::stop then waits for it to complete. If the restart ran inside the reload, the request would wait on the stop, and the stop on the request, until the grace ran out.
  • A failed bind is not retried. When the new listener cannot bind, the follower logs API on <listen>: <e>; it stays off until [api] changes and keeps running empty. It acts again on the next change of the published options.
  • The sender outlives the loop. The watch::Sender lives in Core, and the follower holds an Arc<Instance>, so changed() returning an error cannot happen while the task runs; the break for it is defensive.

Api::stop sends on the stop channel, which completes the graceful-shutdown future passed to ApiListener::serve: axum stops accepting and lets the requests in flight finish. It then waits up to SHUTDOWN_GRACE for the server task:

Server task result Log line
Finished with Ok(()) in time none
Finished with Err(e) API: <e> at ERROR
Panicked or was cancelled (JoinError) API task failed: <e> at ERROR
Still running after 5 s task.abort() on the server task, then API: requests still in flight when it stopped were dropped at WARN

The server task is the one running ApiListener::serve, which awaits axum::serve(..).with_graceful_shutdown(..).

The follower’s other log lines are API listening on <addr> (INFO, from start_api) and API stopped: the config no longer has [api] (INFO). The second is printed whenever the published options become None, whether or not an API was running at the time: after a failed rebind, for example, none was. The API’s endpoints and authentication are on the REST API page; the operator’s view is in the user guide’s API page.

app/src/main.rs
async fn wait_for_shutdown();

On Unix, wait_for_shutdown registers a SIGTERM stream with tokio::signal::unix::signal(SignalKind::terminate()) and returns on whichever comes first, tokio::signal::ctrl_c() (SIGINT) or SIGTERM. If the SIGTERM registration fails, it waits for SIGINT alone. On other platforms it waits for Ctrl-C. Tokio keeps its handlers installed for the life of the process, so a second SIGINT or SIGTERM during the shutdown does not cut it short; SIGKILL does, and then no cleanup runs.

sequenceDiagram
  participant M as main
  participant F as follow_api task
  participant API as API server task
  participant S as supervisor actor
  M->>M: SIGINT or SIGTERM, log shutting down
  M->>F: stop_api.send(())
  F->>API: Api::stop, then wait up to 5 s
  API-->>F: finished, or aborted with a warning
  F-->>M: the task returns
  M->>S: Instance::shutdown(SHUTDOWN_GRACE)
  S->>S: stop listeners and probes, wait up to 5 s for connections
  S->>S: cancel the root, wait for every task, stop the sampler
  S-->>M: done
  M->>M: return ExitCode::SUCCESS
  1. main logs shutting down at INFO.
  2. stop_api.send(()) signals the follower, and main awaits its task, discarding the JoinHandle’s result (let _ = api.await). The follower leaves its loop at a select!: when it is in the middle of a restart, the restart completes first. tokio::select! picks at random among the branches that are ready, so when a change of the options is pending as well, one more restart can run before the loop sees the stop signal. The follower then stops the running API, if any, with up to SHUTDOWN_GRACE for requests in flight.
  3. instance.shutdown(SHUTDOWN_GRACE) runs Supervisor::shutdown(Duration::from_secs(5)). Actor::shutdown calls stop() on every listener (an accept loop that ends drops its listener, and a Unix socket file with it), cancels every balancer’s probe token, closes the task tracker, waits up to the grace with tokio::time::timeout(grace, tracker.wait()), cancels its root token so every remaining connection ends, waits for every tracked task, clears its listener handles, cancels background_stop, and awaits each background task: the sampler, and a usage sink’s push task, which the app does not have. The sequence is on The supervisor.
  4. main returns ExitCode::SUCCESS. The runtime shuts down and drops the watcher task with its notify watcher.

The API is stopped first, so a request that finishes within the API’s grace is answered before the supervisor starts shutting down. A watcher pass can still run during step 3: the supervisor’s actor handles commands in order, so an apply already sent completes before the shutdown, and one sent after it fails with ApplyError::Stopped (reload: the supervisor has shut down; keeping the running config).

How the process ends Exit status Shutdown sequence
--test, accepted 0 none; nothing was started
--test, refused 1 none
The start is refused 1 none; the refused apply released what it bound
The first API bind fails 1 Instance::shutdown(Duration::ZERO)
SIGINT or SIGTERM 0 API stop, then Instance::shutdown(SHUTDOWN_GRACE)
SIGHUP killed by the signal (129 in a shell) none
SIGKILL killed by the signal none

A signal-driven shutdown takes up to 5 s for the API’s requests plus up to 5 s for live connections, and then as long as the remaining tasks take to notice the root token’s cancellation.

Invariant Enforced by Pinned by
A start runs the whole config or nothing Core::start returns any ApplyError from SupervisorBuilder::start, which fails before it spawns the actor and the sampler; a failed prepare drops what it bound a_refused_spec_changes_nothing (for the prepare); no app-level test
Reloads of one Core never overlap last is a tokio::sync::Mutex held across the read and the apply none directly
Sources equal to the recorded ones are not applied again sources == last.sources returns Unchanged a_reload_after_the_apis_own_finds_nothing_changed (app/tests/unit/api.rs)
A reload that is not applied changes nothing that runs, the API included The supervisor’s prepare and commit; send_if_modified only after Ok a_refused_spec_changes_nothing (supervisor/tests/hot_swap.rs)
An open connection whose inbound and user are unchanged keeps relaying across a reload that changes routing Sessions run under the supervisor’s root token, not their listener’s; new connections take the new plane an_open_connection_keeps_transferring_across_a_reload; an_established_connection_survives_an_apply_that_keeps_its_inbound (supervisor level)
A protocol change on the same port keeps the socket, and a SOCKS connection accepted before a swap to HTTP keeps speaking SOCKS Listener::swap puts the new handler in the listener’s watch for new connections (swapped); the accept loop runs on a_protocol_change_on_the_same_port_keeps_the_socket_and_old_connections (supervisor level)
Removing a SOCKS account closes that account’s open connection and leaves another account’s running The removal policy the app lowers, UserRemovalPolicy::Close a_reload_that_removes_a_user_closes_only_their_connections
Removing an inbound closes its sessions Step::CloseSessions in the commit removing_an_inbound_closes_its_sessions
A changed Hysteria 2 inbound on the same port keeps serving on it The listener keeps its QUIC endpoint a_reload_keeps_the_inbound_serving_on_its_udp_port
The app never applies a disruptive change Supervisor::apply with ApplyOptions::default() a_hysteria2_obfuscation_change_needs_allow_disruptive (the supervisor’s refusal); no app-level test
The API follows the applied [api]: gone when it goes, back where it comes back, rebound when only the secret changes send_if_modified and follow_api the_api_follows_the_api_section_across_reloads
A config without [api] binds nothing beyond its inbounds main starts the API only for Some options a_config_without_api_binds_nothing_new (Linux only)
A reload made by the API answers before the API restarts The restart runs in follow_api, not in reload_with none
Tracing is set up once, from RUST_LOG or [log].level init_tracing at the top of main none
A clean shutdown releases what the inbounds own Supervisor::shutdown stops every listener socks_over_a_unix_socket_relays_and_cleans_up (the socket file is gone after SIGTERM)
Shutdown is not held up by the supervisor’s own timers or by a stalled Hysteria 2 handshake root.cancel() after the grace ends grace timers and closes a Hysteria 2 endpoint shutdown_ends_pending_grace_timers, shutdown_does_not_wait_out_a_stalled_hysteria2_handshake (supervisor level)
Failure Log line Effect
The config or subscribe file cannot be read, a TOML or lowering error, or any ApplyError failed to start: <e> Exit status 1; nothing keeps running
A listener cannot be bound failed to start: inbound <tag>: binding <bind> failed: <e> Exit status 1; listeners already bound in the same prepare are released
The REST API cannot be bound failed to start: API on <listen>: <e> Instance::shutdown(Duration::ZERO), exit status 1
The watcher cannot be created, or the config’s directory cannot be watched config hot-reload disabled: <e> The process runs without file-triggered reloads; the REST API can still reload

Every failure keeps the running config, as the outcomes table shows. A subscribe file’s directory that cannot be watched is logged as cannot watch <dir> for changes: <e> and does not stop anything else.

A rebind that fails leaves the API off with API on <listen>: <e>; it stays off until [api] changes. The proxy is not affected. A stop that overruns SHUTDOWN_GRACE calls task.abort() on the server task and logs API: requests still in flight when it stopped were dropped at WARN.

  • Only the watcher awaits reload_with to completion. An API request’s handler is dropped when its client goes away while the request is in flight (hyper sees EOF on an HTTP/1 connection, or a reset of the HTTP/2 stream), and reload_with is dropped with it at whichever await it has reached. How the actor treats a command whose caller has gone is on The supervisor.
  • The watcher task is never cancelled explicitly. It ends when the runtime shuts down after main returns.
  • The follower ends on the stopped signal from main. main sends it only after a signal, and awaits the follower before it shuts the supervisor down.
  • Api::stop bounds its wait with tokio::time::timeout(SHUTDOWN_GRACE, &mut task) and aborts the server task on timeout.
Constant Value Defined in Meaning
SHUTDOWN_GRACE 5 s app/src/main.rs Grace for the API’s requests at every API stop, and for live connections at shutdown
Grace after a failed API start Duration::ZERO app/src/main.rs → main Closes what the inbounds accepted at once
DEFAULT_SAMPLE_INTERVAL 1 s supervisor/src/track/sampler.rs The stats sampler’s tick; the app does not change it
Supervisor command queue 16 supervisor/src/supervisor.rs → SupervisorBuilder::start, mpsc::channel(16) Commands waiting for the actor, a reload’s apply among them
DEFAULT_LISTEN 127.0.0.1:9090 webclient/src/listen.rs, used by ApiConfig’s serde default default_api_listen The API’s address when [api] has no listen
Disruptions reported per refusal 1 Actor::apply → plan.disruptions().next() Only the first disruptive inbound is named

The process-wide limits of the proxy itself are collected on Limits, timeouts and memory.

The end-to-end tests run the real binary. app/tests/support/mod.rs → spawn_app(dir, "app.toml") starts CARGO_BIN_EXE_etemenanki-app with -c app.toml and the test directory as its working directory, so the config path is a bare file name and the watched directory is ., with standard output and standard error discarded. The returned Proc kills the child with SIGKILL when dropped; Proc::terminate sends SIGTERM and polls for the exit for up to 10 s, for the tests where what the process does on the way out is the thing under test. The reload tests rewrite the config file in place with std::fs::write and then poll a probe every 100 ms for up to 20 s (eventually) until a new connection shows that the new config took effect, so what the open connection does next is observed after the reload.

File Test Behaviour it pins
app/tests/integration/e2e_reload.rs an_open_connection_keeps_transferring_across_a_reload A SOCKS connection opened before a reload that adds a 127.0.0.0/8 → block rule keeps echoing, while new connections are blackholed
app/tests/integration/e2e_reload.rs a_reload_that_removes_a_user_closes_only_their_connections Removing alice from the SOCKS accounts closes her open connection within 10 s and refuses her next login; bob’s open connection keeps echoing
app/tests/integration/e2e_api.rs the_api_follows_the_api_section_across_reloads [api] removed: the port stops answering. Added on another port with secret = "one": 200 with the token, 401 without. Only the secret changed to "two": the same port answers the new token and refuses the old one
app/tests/integration/e2e_api.rs a_config_without_api_binds_nothing_new Read from /proc: a config without [api] listens on its SOCKS port only; the same config with [api] adds exactly the API’s port. Linux only
app/tests/integration/e2e_api.rs a_switch_sends_new_flows_of_the_group_through_the_new_target PUT /v1/routes/Local answers "status":"applied", so the edit’s reload applied
app/tests/integration/e2e_hysteria_inbound.rs a_reload_keeps_the_inbound_serving_on_its_udp_port A Hysteria 2 inbound rewritten from udp = false to udp = true on the same port, with a real hysteria client connected, still serves through that client after the reload. It waits 6 s after the write, then retries 20 times at 500 ms. Skips when go or the upstream build is unavailable
app/tests/integration/e2e_unix.rs socks_over_a_unix_socket_relays_and_cleans_up After Proc::terminate (SIGTERM), the inbound’s socket file is gone
app/tests/integration/e2e_tun.rs a_routed_connect_is_answered_while_the_app_runs Once the process has exited after Proc::terminate, a connect routed into the TUN device no longer succeeds: the interface and its route are gone. The kernel removes a created device when its descriptor closes, so this does not tell a clean shutdown from any other exit. Needs CAP_NET_ADMIN, and skips without it
app/tests/unit/api.rs a_reload_after_the_apis_own_finds_nothing_changed After a PUT applied an edit, POST /v1/reload answers 200 {"status":"unchanged"}: a reload of the bytes the edit’s reload recorded does not apply them again
app/tests/unit/api.rs a_switch_edits_the_file_reloads_and_returns_the_new_etag A PUT over an Instance started in-process reloads ("status":"applied") and returns the file’s new ETag
app/tests/unit/api.rs api_options_come_from_the_api_section api::options, part of build: no [api] is None; an empty [api] listens on DEFAULT_LISTEN with no secret; a value containing / is a Unix socket path; a non-local listen with a secret is accepted
app/tests/unit/api.rs an_api_open_to_other_hosts_needs_a_secret A non-local listen without a secret, an empty secret on loopback, a host name as listen and an unknown key in [api] are all refused
supervisor/tests/hot_swap.rs a_refused_spec_changes_nothing A spec refused by validation (UnknownReference) or by a bind while preparing (Bind) leaves the epoch, the routes, the users and the listeners as they were, and releases a listener bound in that prepare
supervisor/tests/hot_swap.rs a_hysteria2_obfuscation_change_needs_allow_disruptive Without allow_disruptive, an obfuscation change is ApplyError::Disruptive for hy2-in
supervisor/tests/hot_swap.rs an_established_connection_survives_an_apply_that_keeps_its_inbound An apply that adds an outbound and a rule reports the inbound as reused, neither swapped nor rebound, and an open SOCKS connection keeps echoing
supervisor/tests/hot_swap.rs a_protocol_change_on_the_same_port_keeps_the_socket_and_old_connections SOCKS to HTTP on the same port is swapped, not rebound; an open SOCKS connection keeps echoing, a connection accepted before the swap still gets the SOCKS handler, and a new one speaks HTTP
supervisor/tests/hot_swap.rs removing_an_inbound_closes_its_sessions Removing one of two SOCKS inbounds reports it as removed and closes its open connection; the other inbound’s connection keeps echoing
supervisor/tests/hot_swap.rs shutdown_ends_pending_grace_timers A removal and a drain each waiting an hour do not hold up shutdown(Duration::ZERO), and the session in its removal grace is closed
supervisor/tests/hot_swap.rs shutdown_does_not_wait_out_a_stalled_hysteria2_handshake A Hysteria 2 client stuck in its QUIC handshake does not hold up the server’s shutdown
ffi/tests/proxy.rs set_route_and_reload_switch_like_the_rest_api A Core started by the FFI with its own Lowering, not by Instance, applies a route switch through reload_with (see the FFI page)

No test covers the watcher following a newly named subscribe file, the reload log lines, Reload::report with a recorded failure, the disruptive refusal at the app level, a failed API start or rebind, or SIGHUP. A change to those paths should add one; Testing describes how the suites are laid out.