etemenanki-app: running, reloading and shutting down
Source files: 25 · checked against Etemenanki 555b7df
Etemenanki/app/src/main.rsEtemenanki/app/src/instance.rsEtemenanki/app/src/lib.rsEtemenanki/app/src/config.rsEtemenanki/app/src/api.rsEtemenanki/supervisor/src/lib.rsEtemenanki/supervisor/src/supervisor.rsEtemenanki/supervisor/src/build/apply.rsEtemenanki/supervisor/src/topology/spec_plan/plan.rsEtemenanki/supervisor/src/topology/inbound/mod.rsEtemenanki/supervisor/src/system/listener.rsEtemenanki/supervisor/src/entity/id.rsEtemenanki/supervisor/src/policy.rsEtemenanki/supervisor/src/track/sampler.rsEtemenanki/webclient/src/lib.rsEtemenanki/webclient/src/listen.rsEtemenanki/webclient/src/router.rsEtemenanki/app/tests/integration/e2e_reload.rsEtemenanki/app/tests/integration/e2e_api.rsEtemenanki/app/tests/integration/e2e_hysteria_inbound.rsEtemenanki/app/tests/integration/e2e_unix.rsEtemenanki/app/tests/integration/e2e_tun.rsEtemenanki/app/tests/support/mod.rsEtemenanki/app/tests/unit/api.rsEtemenanki/supervisor/tests/hot_swap.rs
etemenanki-app is a thin binary around one supervisor. app/src/main.rs parses the command line, sets up logging, starts an Instance, starts the REST API when the config has an [api] section, watches the config file and the subscribe file it names, and waits for a signal. app/src/instance.rs owns the load itself: it reads the files, merges and lowers them into a Spec, hands the spec to the supervisor, logs what happened and remembers the sources it last recorded. Every reload goes through the supervisor’s validate, plan, prepare and commit. Sessions are not tied to their listener, so a reload does not close them just because it ran, and a refused reload leaves everything that runs as it was.
This page follows the process from main to exit: the CLI, tracing, --test, the start, the file watcher, Core::reload_with step by step with every log line it can print, what a reload keeps, the API follower and shutdown. It is for changing main.rs or instance.rs, building another front end on Core, or explaining what a reload printed. How a config becomes a spec is on the config page, the subscribe merge on the subscribe page, and the API’s endpoints on the REST API page. What an apply does inside the supervisor is on Planning and applying a change.
Responsibilities
Section titled “Responsibilities”| Component | Owns | Does not |
|---|---|---|
app/src/main.rs → main |
Parsing Args, init_tracing, the --test branch, Instance::start, the first API start, spawning the API follower and the watcher, waiting for SIGINT or SIGTERM, and the shutdown order |
Read any config key other than [log].level |
app/src/main.rs → spawn_watcher, sync_watches, watch_dir |
The notify watcher, the set of directories it watches, and the task that turns the watcher’s ticks into Instance::reload calls |
Decide whether the files changed; Core::reload_with does |
app/src/main.rs → follow_api, start_api, Api |
Starting, stopping and rebinding the REST API as the running config’s [api] changes |
Decide when [api] changed; Core publishes that after an apply |
app/src/instance.rs → Core |
Starting a supervisor from a Read, reloading it through a Lowering, skipping sources equal to the ones it last recorded, logging each outcome, and publishing the running config’s API options |
Know where the files come from |
app/src/instance.rs → Instance |
The config path, reading the sources from disk, and the list of files worth watching | Watch anything itself |
app/src/instance.rs → check, check_bytes |
The dry run behind --test and behind every API route edit |
Bind a listener or start a task |
etemenanki_supervisor::Supervisor |
Validation, planning, prepare and commit, listeners, sessions, outbounds, shutdown | Know about files, TOML, the API or signals: the crate “knows nothing of config files, their formats, or how a reload is triggered” (supervisor/src/lib.rs) |
app/src/lib.rs makes the split explicit: the library crate is “the TOML front end of Etemenanki” (api, config, instance, lower, routes, subscribe, transport), and “the etemenanki-app binary is the thin glue on top: CLI, tracing, the file watcher that triggers instance::Instance::reload, and, with an [api] section, the REST API”.
The binary does not:
- install a
SIGHUPhandler. A reload is triggered by the file watcher, byPOST /v1/reloador byPUT /v1/routes/{group}on the REST API, and by nothing else.SIGHUPkeeps its default action and terminates the process at once, without the shutdown sequence (a shell reports exit status 129). - write logs anywhere but standard output.
tracing_subscriber::fmt()writes to stdout by default, andinit_tracingdoes not change the writer. - change any supervisor setting that outlives a spec. It starts the supervisor with
Supervisor::builder()and every default (see What a reload keeps).
Key types
Section titled “Key types”/// A minimal Xray-core-style proxy runtime.#[derive(Parser)]#[command(version, about)]struct Args { /// Path to the TOML configuration file. #[arg(short, long, default_value = "config.toml")] config: PathBuf, /// Validate the configuration and exit without binding any listener. #[arg(long)] test: bool,}The --help output of the pinned binary:
A minimal Xray-core-style proxy runtime
Usage: etemenanki-app [OPTIONS]
Options: -c, --config <CONFIG> Path to the TOML configuration file [default: config.toml] --test Validate the configuration and exit without binding any listener -h, --help Print help -V, --version Print version--version prints etemenanki-app 2.1.1: clap puts the binary’s name before the version in app/Cargo.toml. The REST API reports the same version at /v1/version, as the bare env!("CARGO_PKG_VERSION"): {"version":"2.1.1"}. The config path is kept exactly as given. A relative path is resolved against the working directory on every read, and the binary never changes its working directory. The user-facing reference is Command line.
SHUTDOWN_GRACE
Section titled “SHUTDOWN_GRACE”/// How long live connections get to finish on shutdown before they are/// closed.const SHUTDOWN_GRACE: Duration = Duration::from_secs(5);It is used twice: as the time the REST API gets to finish its requests in flight whenever it stops (Api::stop), and as the grace passed to Instance::shutdown, which the supervisor gives live connections.
LoadError
Section titled “LoadError”#[derive(Debug, thiserror::Error)]pub enum LoadError { /// The files could not be read, or what they say could not be lowered. #[error(transparent)] Config(#[from] io::Error), /// The supervisor refused the spec. #[error(transparent)] Apply(#[from] ApplyError),}
impl From<LoadError> for ControlError { /* see the table */ }Both variants are transparent, so every log line and every API error shows the inner error’s text unchanged. From<LoadError> for ControlError sorts a failure by fault for the REST API, which answers Invalid with HTTP 422 and Failed with HTTP 500:
LoadError |
ControlError |
Why |
|---|---|---|
Config(e) with e.kind() of InvalidData or InvalidInput |
Invalid(e.to_string()) |
The config’s content is at fault: a TOML error (config::parse_bytes uses InvalidData), a lowering refusal or an [api] error (InvalidInput) |
Config(e) of any other kind |
Failed(e.to_string()) |
A file that cannot be read, for example NotFound or PermissionDenied |
Apply(ApplyError::Bind { .. }) or Apply(ApplyError::Stopped) |
Failed(e.to_string()) |
Not the config’s fault: a port held elsewhere, or a supervisor that has shut down |
Apply(e), every other variant |
Invalid(e.to_string()) |
Validation, a build failure or a disruptive change |
The ApplyError variants and their texts are listed on Validation and apply errors.
Built, build and Lowering
Section titled “Built, build and Lowering”pub struct Built { pub spec: Spec<UserKey>, pub api: Option<ApiOptions>,}
pub fn build(cfg: &Config) -> io::Result<Built> { Ok(Built { spec: lower(cfg)?, api: crate::api::options(cfg)?, })}
pub type Lowering = Arc<dyn Fn(&Config) -> io::Result<Built> + Send + Sync>;A Built is everything one config produces: the spec the supervisor runs, and the REST API’s options when the config has an [api] section. build is the app’s Lowering: lower (on the config page) followed by api::options. That returns None without an [api] section. Otherwise it parses [api].listen and applies ApiOptions::new, and it wraps any OptionsError as an InvalidInput error prefixed with [api] :
| Rule | Checked by | Error text |
|---|---|---|
listen defaults to DEFAULT_LISTEN, 127.0.0.1:9090 |
ApiConfig’s serde default default_api_listen (app/src/config.rs) |
none |
A listen containing / is a Unix socket path; anything else must parse as ip:port |
Listen::from_str |
[api] listen "<s>": expected ip:port, or a path (containing a /) for a unix socket |
A secret that is set must not be empty, whatever the listen |
ApiOptions::new |
[api] secret is empty |
A listen that is neither a loopback address nor a Unix socket needs a secret |
ApiOptions::new (Listen::is_local) |
[api] listen <addr>: an API reachable from other hosts needs a secret; listen on a loopback address or a unix socket, or set one |
An [api] error is therefore a lowering error, found by --test and reported by a reload like any other config mistake. api_options_come_from_the_api_section and an_api_open_to_other_hosts_needs_a_secret (app/tests/unit/api.rs) pin these rules. Listen, ApiOptions and OptionsError are described in full on the REST API page.
Lowering is a shared closure rather than a function so that a Core can be started with a different one. The app passes Arc::new(build). A front end handed the files’ contents rather than paths, such as the mobile FFI, runs a Core of its own with its own Lowering (see the FFI page).
Read and Sources
Section titled “Read and Sources”pub type Read = io::Result<(Sources, io::Result<Config>)>;#[derive(Clone, PartialEq, Default)]pub struct Sources { pub config: Vec<u8>, pub subscribe: Option<Vec<u8>>,}
pub fn read_sources(path: &Path) -> io::Result<(Sources, io::Result<Config>)>;pub fn sources_from(config: Vec<u8>) -> io::Result<(Sources, io::Result<Config>)>;pub fn effective(parsed: io::Result<Config>, sources: &Sources) -> io::Result<Config>;A Read is one reading of the files a Core loads. The outer io::Result is the reading itself, which fails when the files cannot be gathered into Sources (the cases are on the config page). The inner one is the config’s parse, handed back with the bytes because the subscribe file’s path is known only once the config has parsed. A config that does not parse is still a set of sources; its parse error comes out of config::effective at the lowering step. Sources holds the raw bytes of the config file and of the subscribe file when the config names one, and compares them byte for byte. The three functions are described on the config page; effective merges the subscribe file (see the subscribe page).
Reload
Section titled “Reload”#[derive(Debug)]pub enum Reload { /// The files changed, and were applied. Applied(ApplyReport), /// The files are byte for byte the ones last tried, and were not tried /// again. `failed` is why that last try failed, if it did. Unchanged { failed: Option<String> },}
impl Reload { pub fn report(self) -> Result<ReloadReport, ControlError>;}report is what an API client is told:
Reload |
report() |
HTTP answer |
|---|---|---|
Applied(report) |
Ok(ReloadReport::Applied(report)) |
200, {"status":"applied", "built":[…], "swapped":[…], "drained":[…], "rebound":[…], "restarted":[…], "removed":[…], "reused":[…]} |
Unchanged { failed: None } |
Ok(ReloadReport::Unchanged) |
200, {"status":"unchanged"} |
Unchanged { failed: Some(reason) } |
Err(ControlError::Invalid(...)) with the text <reason> (the files have not changed since this was found) |
422 |
The doc comment on ReloadReport::Unchanged in webclient/src/lib.rs names the common case: “the file watcher may have reloaded them already”. Sources recorded together with a failure are refused again with that failure’s text, as they would be if tried.
Core and Last
Section titled “Core and Last”pub struct Core { supervisor: Supervisor<UserKey>, lowering: Lowering, api: tokio::sync::watch::Sender<Option<ApiOptions>>, last: tokio::sync::Mutex<Last>,}
struct Last { sources: Sources, failed: Option<String>,}
impl Core { pub async fn start( builder: SupervisorBuilder<UserKey>, read: Read, lowering: Lowering, ) -> Result<Self, LoadError>; pub fn supervisor(&self) -> &Supervisor<UserKey>; pub fn api(&self) -> tokio::sync::watch::Receiver<Option<ApiOptions>>; pub fn tracker(&self) -> Tracker; pub async fn reload_with(&self, read: impl FnOnce() -> Read) -> Result<Reload, LoadError>; pub async fn shutdown(&self, grace: Duration);}| Field | Type | Role |
|---|---|---|
supervisor |
Supervisor<UserKey> |
The running supervisor. UserKey is UserName, the name the app keys a config user by (see the config page). |
lowering |
Lowering |
How a merged Config becomes a Built, fixed for the Core’s life |
api |
tokio::sync::watch::Sender<Option<ApiOptions>> |
The API options of the config running now: None while it has no [api]. Changed only by an applied reload. |
last |
tokio::sync::Mutex<Last> |
The last recorded sources, and the text of the failure recorded with them, if any. The lock is held for a whole reload, which serialises reloads. |
Core is the load itself, wherever the files come from. It holds one Supervisor handle. Supervisor is Clone, and every clone drives the same actor over its command channel; when the last handle is dropped without shutdown, the actor sees the channel close and runs its shutdown with Duration::ZERO. A command whose send fails, or whose reply is dropped, comes back as ApplyError::Stopped (Supervisor::ask). The details are on The supervisor. api() subscribes a new receiver to the watch; the receiver starts having seen the current value, so its changed() waits for the next change. tracker() clones the supervisor’s Tracker, which the REST API reads without going through the supervisor’s actor (see Tracking flows, stats and speed limits). shutdown(grace) is Supervisor::shutdown(grace).
Instance
Section titled “Instance”pub struct Instance { path: PathBuf, core: Core,}
impl Instance { pub async fn start(path: PathBuf) -> Result<Self, LoadError>; pub fn api(&self) -> tokio::sync::watch::Receiver<Option<ApiOptions>>; pub fn config_path(&self) -> &Path; pub fn tracker(&self) -> Tracker; pub fn watched_files(&self) -> Vec<PathBuf>; pub async fn reload(&self) -> Result<Reload, LoadError>; pub async fn shutdown(&self, grace: Duration);}Instance is the app’s Core, reading its files from disk:
start(path)isCore::start(Supervisor::builder(), config::read_sources(&path), Arc::new(build)).reload()iscore.reload_withwith a read closure that callsconfig::read_sources(&self.path)and rewrites a read error asio::Error::new(e.kind(), format!("cannot read {}: {e}", ...)). The new error keeps theio::ErrorKind, soControlErrorstill sorts it by kind; the original error survives only as text.watched_files()returns the config path and, when the config file on disk parses and its[subscribe]names apath, that path. It reads the file withconfig::loadrather than asking the running config, “so a newly named subscribe file is followed even while the config naming it fails to apply”. A config that does not parse yields the config path alone.config_path()is the-cpath as given. The REST API’sConfigFileuses it to read and rewrite the file.
summary
Section titled “summary”fn summary(report: &ApplyReport) -> String;summary turns an ApplyReport into the text after config loaded: and config reloaded: . It prints each non-empty group as <label> [<first>, <second>, …], joins the groups with ; , and uses this order:
| Order | Label | Items | Printed as |
|---|---|---|---|
| 1 | built |
Resource |
dns, route, outbound <tag>@v<n>, balancer <tag>, user set <tag>, inbound <tag> |
| 2 | swapped |
inbound tag | <tag> |
| 3 | drained |
OutboundId |
<tag>@v<n> |
| 4 | rebound |
inbound tag | <tag> |
| 5 | restarted |
inbound tag | <tag> |
| 6 | removed |
inbound tag | <tag> |
| 7 | reused |
Resource |
as for built |
When every group is empty it prints nothing to run. reused comes last although it is the first field of ApplyReport, so the line leads with what changed. ApplyReport’s doc comment states that every resource of the desired spec appears in reused or built; a start has nothing to reuse, so its line lists every resource under built. Because the app never allows a disruptive change (see Disruptive changes are refused), its lines never have a restarted group. What each group means is on Planning and applying a change. Lines captured from the pinned debug binary (timestamps removed):
INFO etemenanki_app::instance: config loaded: built [dns, outbound direct@v1, route]INFO etemenanki_app::instance: config loaded: built [dns, outbound direct@v1, outbound n@v1, route]INFO etemenanki_app::instance: config reloaded: reused [dns, outbound direct@v1, route]INFO etemenanki_app::instance: config reloaded: built [outbound block@v1]; reused [dns, outbound direct@v1, route]INFO etemenanki_app::instance: config reloaded: built [outbound block2@v1]; drained [block@v1]; reused [dns, outbound direct@v1, route]INFO etemenanki_app::instance: config reloaded: built [dns, outbound direct@v2]; drained [direct@v1]; reused [route]INFO etemenanki_app::instance: config reloaded: drained [n@v1]; reused [dns, outbound direct@v1, route]The second line is a start with a subscribe file that defines a node n. The third is the reload after an edit that only added a comment: the sources changed, so the spec was applied, and the supervisor reused everything. The sixth is a change to [dns], which rebuilds the freedom outbound as a new version as well. The last is the config dropping its [subscribe] section. At debug level, every commit also logs applied: <n> built, <n> reused, <n> swapped, <n> drained from the target etemenanki_supervisor::supervisor.
The supervisor logs two more lines of its own around these. Every listener that a start or an applied reload binds logs inbound <tag> listening on <bind> at INFO (target etemenanki_supervisor::system::listener) when Listener::start runs during the commit, so on a start these lines come before config loaded:. A TUN inbound that creates its device logs inbound <tag> owns tun device <name> when it is bound. Both are described on Listeners and the serve loop.
Data flow
Section titled “Data flow”flowchart LR
M["main"] --> ST["Instance::start"]
ST --> CS["Core::start"]
CS --> SUP["supervisor actor"]
W["watcher task"] --> IR["Instance::reload"]
RP["POST /v1/reload"] --> IR
PR["PUT /v1/routes/{group}"] --> IR
IR --> RW["Core::reload_with"]
RW -- "apply" --> SUP
RW -- "send_if_modified" --> AW["api watch"]
AW --> FA["follow_api"]
FA --> API["API server task"]
M -- "signal" --> FA
M -- "Instance::shutdown" --> SUP
One supervisor runs for the life of the process. Three callers reload it, all through Instance::reload and Core::reload_with, which applies to the supervisor and then publishes the running config’s API options. The API follower is the only code that starts, stops or rebinds the REST API. On a signal, main stops the follower first and the supervisor second. The sections below take these in order: Starting, The file watcher, Reloading, What a reload keeps, The API follower and Shutting down.
Starting
Section titled “Starting”flowchart TB
A["Args::parse"] --> T["init_tracing"]
T --> Q{"--test?"}
Q -- "yes" --> C["instance::check, print the verdict, exit 0 or 1"]
Q -- "no" --> S["Instance::start"]
S -- "Err" --> F1["log failed to start, exit 1"]
S -- "Ok" --> P{"running config has an api section?"}
P -- "yes" --> SA["start_api"]
SA -- "Err" --> F2["log, shutdown with zero grace, exit 1"]
SA -- "Ok" --> FA["spawn follow_api"]
P -- "no" --> FA
FA --> W["spawn_watcher"]
W --> SIG["wait for SIGINT or SIGTERM"]
main runs on #[tokio::main], the multi-threaded runtime.
Tracing
Section titled “Tracing”fn init_tracing(config_path: &Path) { use tracing_subscriber::EnvFilter; let filter = EnvFilter::try_from_default_env().unwrap_or_else(|_| { let level = config::load(config_path) .ok() .and_then(|c| c.log.level) .unwrap_or_else(|| "info".to_string()); EnvFilter::new(level) }); tracing_subscriber::fmt().with_env_filter(filter).init();}EnvFilter::try_from_default_env()readsRUST_LOG. When it is set and parses, it wins.- Otherwise, whether
RUST_LOGis unset or set to something that does not parse (both are anErr), the config file is read and parsed once withconfig::load(the file as written, without the subscribe merge), and[log].levelbecomes the filter throughEnvFilter::new, which takes the same directive syntax asRUST_LOG(for exampleinfo,etemenanki_supervisor=debug). - When the key is absent, or the file cannot be read or does not parse, the filter is
info. init()installs the global subscriber, which formats to standard output.
This runs once, before --test and before the start. Nothing later replaces the filter, so [log].level is not reloaded: a reload that changes only [log].level finds the sources changed, applies a spec in which nothing changed (config reloaded: reused [...]), and logs at the old level until the process restarts. The --test verdict on failure is a log line and obeys the same filter; the exit status is always set.
The log targets are the module paths: etemenanki_app for the lines main.rs prints (failed to start, configuration invalid, shutting down, the API and watcher lines) and etemenanki_app::instance for the lines instance.rs prints (config loaded, config reloaded, every reload line).
--test
Section titled “--test”pub async fn check(path: &Path) -> Result<(), LoadError> { check_bytes(std::fs::read(path)?).await}
pub async fn check_bytes(config: Vec<u8>) -> Result<(), LoadError> { let (sources, parsed) = config::sources_from(config)?; let built = build(&config::effective(parsed, &sources)?)?; etemenanki_supervisor::check(&built.spec).await?; Ok(())}--test reads the config, reads the subscribe file it names, merges, lowers (including [api]) and runs etemenanki_supervisor::check. That builds a fresh Actor::new(SocketOptions::default()), plans against its empty running state, and runs prepare(&plan, spec, false): with bind false, every outbound, route, handler and user table is still built, and only listener::bind and Pending::pair are skipped. Because it plans against an empty state, it never reports a disruptive change. It refuses what a start would refuse, short of a port, socket path or device that cannot be bound and of the checks pairing makes, and it never binds the REST API. What it checks in detail is on the config page and in the dry-run section of Planning and applying a change.
| Result | Output | Exit status |
|---|---|---|
| Accepted | Configuration OK. on standard output (println!, not a log line) |
0 |
| Refused | configuration invalid: <e> at ERROR, target etemenanki_app |
1 |
Captured from the pinned binary (timestamps removed):
ERROR etemenanki_app: configuration invalid: route references unknown outbound nowhereERROR etemenanki_app: configuration invalid: [subscribe] needs a path naming the subscribe fileERROR etemenanki_app: configuration invalid: [api] listen 0.0.0.0:9090: an API reachable from other hosts needs a secret; listen on a loopback address or a unix socket, or set oneERROR etemenanki_app: configuration invalid: No such file or directory (os error 2)ERROR etemenanki_app: configuration invalid: subscribe file missing-sub.toml: No such file or directory (os error 2)A config file that cannot be read gives the bare io::Error, with no path: check calls std::fs::read(path)? directly, unlike a reload’s cannot read <path>: . A missing subscribe file is named by config::sources_from.
check_bytes is also what the REST API runs on an edited config before it writes the file (app/src/api.rs → ConfigFile::set_route).
Instance::start
Section titled “Instance::start”Instance::start(path) calls Core::start(Supervisor::builder(), config::read_sources(&path), Arc::new(build)). Core::start:
- Returns the read error unchanged, if
read_sourcesfailed. Unlike a reload’s, it carries nocannot read <path>:prefix. - Runs
config::effectiveand the lowering. A merge or lowering error is returned asLoadError::Config. - Calls
builder.start(built.spec). That asserts the sample interval is non-zero (it panics otherwise; the app’s is the default 1 s) and runs one apply against a new, empty actor: validate, plan, prepare (binding every listener) and commit, during which each new listener logsinbound <tag> listening on <bind>. AnyApplyErroris returned, aBinderror included, so a start is all or nothing. A failed prepare drops what it built, which releases every listener it had bound. Only after the apply succeeds does it spawn the sampler (and a usage sink, which the app does not have), create the command channel withmpsc::channel(16)and spawnactor.run. The builder is described on The supervisor. - Logs
config loaded: <summary>atINFO. - Stores the supervisor and the lowering, creates the
apiwatch withbuilt.api, and records the sources inlastwithfailed: None. A reload of the same files right after the start is thereforeUnchanged.
main logs any error as failed to start: <e> at ERROR and exits with status 1. Captured:
ERROR etemenanki_app: failed to start: route references unknown outbound nowhereERROR etemenanki_app: failed to start: No such file or directory (os error 2)ERROR etemenanki_app: failed to start: subscribe file missing-sub.toml: No such file or directory (os error 2)A listener that cannot be bound gives failed to start: inbound <tag>: binding <bind> failed: <e>, where <bind> is the BindSpec’s Display. The app’s lowering produces four of its five forms: <host>:<port>, udp <host>:<port>, unix:<path> and tun <name> (tun auto without a name). The fifth, tun fd <n>, is a supplied descriptor (TunSource::Fd), which only the FFI produces.
Starting the API
Section titled “Starting the API”main takes instance.api(), clones the current value and, when it is Some, calls start_api:
async fn start_api(instance: Arc<Instance>, options: &ApiOptions) -> io::Result<Api>;
struct Api { stop: oneshot::Sender<()>, task: JoinHandle<io::Result<()>>,}start_api binds ApiListener::bind(&options.listen), logs API listening on <addr> at INFO (the bound TCP address from local_addr(), or the listen value for a Unix socket), takes tracker = instance.tracker(), builds etemenanki_webclient::router(ConfigFile::new(instance), tracker, options.secret.clone(), env!("CARGO_PKG_VERSION")), and spawns listener.serve(app, stopped), where stopped is the receiving half of a fresh oneshot. How ApiListener binds a TCP address or a Unix socket path, and when it removes the socket file, is on the REST API page.
When the first bind fails, main logs failed to start: API on <listen>: <e>, calls instance.shutdown(Duration::ZERO) and exits with status 1. By then the supervisor was already running and its inbounds accepting; the zero grace closes whatever they accepted at once. A config without [api] binds nothing beyond its inbounds.
main then creates the oneshot pair (stop_api, api_stopped), spawns follow_api(instance.clone(), options, running, api_stopped) (see The API follower), and calls spawn_watcher. If the watcher cannot be set up, it logs config hot-reload disabled: <e> at ERROR and carries on: the proxy runs, without reloads from the watcher. Last, it awaits wait_for_shutdown.
The file watcher
Section titled “The file watcher”fn watch_dir(path: &Path) -> PathBuf;fn sync_watches( watcher: &mut notify::RecommendedWatcher, watched: &mut HashSet<PathBuf>, instance: &Instance,);fn spawn_watcher(instance: Arc<Instance>) -> notify::Result<()>;spawn_watcher creates a notify::RecommendedWatcher and points it at watch_dir of each path Instance::watched_files() returns, each with RecursiveMode::NonRecursive: it watches the directories that hold the files, not the files themselves. The watcher’s callback feeds () ticks into a channel, and one task turns those ticks into Instance::reload calls. Whether anything changed is decided by Core::reload_with, which re-reads the files by path and compares their bytes.
watch_dir(path) is the file’s parent directory, or . when the path has no parent component (-c config.toml).
Setting up
Section titled “Setting up”spawn_watcher:
- Creates an unbounded
tokio::sync::mpscchannel of(). - Creates the
RecommendedWatcherwithnotify::recommended_watcher, whose callback sends()ticks into the channel. - Watches
watch_dir(instance.config_path()). A failure here, or in step 2, returns thenotify::Error, andmainlogsconfig hot-reload disabled: <e>: the config’s own directory must be watchable, “or there is no hot reload at all”. - Starts
watchedas aHashSetholding that directory, and callssync_watchesonce, which adds the subscribe file’s directory when the config names one elsewhere. - Spawns the watcher task, moving the watcher, the receiver and
watchedinto it, and returnsOk(()).
The RecommendedWatcher lives inside the task. The task owns the watcher, the watcher’s closure owns the sender, so the channel never closes and the task runs until the runtime shuts down when main returns.
The watcher task
Section titled “The watcher task”flowchart TB N["watcher callback sends a tick"] --> R["watcher task: rx.recv()"] R --> D1["drain the ticks already queued"] D1 --> S["settle briefly"] S --> D2["drain again"] D2 --> L["instance.reload()"] L --> SW["sync_watches"] SW --> R
- One reload per pass. After it takes a tick, the task drains the ticks already queued, sleeps for a short settle (an inline
tokio::time::sleep, not a named constant), drains again, and makes oneinstance.reload()call for all the ticks it drained. - One reload at a time. The task awaits
instance.reload()before it takes the next tick, so its own reloads never overlap. Reloads from the REST API are serialised against the watcher’s by thelastlock inCore. - The result is ignored.
reload_withhas already logged the outcome. The watcher does not reportUnchanged, which is silent. - Re-sync after every pass.
sync_watchesruns after each reload, whatever its outcome, so the set of watched directories follows the config file on disk.
Following the subscribe file
Section titled “Following the subscribe file”sync_watches computes the wanted set from instance.watched_files(), mapped through watch_dir:
- For each directory wanted and not yet watched, it calls
watcher.watch(dir, RecursiveMode::NonRecursive). A failure is logged ascannot watch <dir> for changes: <e>atERRORand is not fatal. - For each directory watched and no longer wanted, it calls
watcher.unwatch(dir)and ignores the result. - It stores the wanted set as the watched set.
The config’s own directory is always wanted, because watched_files always starts with the config path. The subscribe file’s directory is wanted while the config file on disk names it. Because watched_files reads the file on disk and not the running config, a config that starts naming a new subscribe file has that file’s directory watched from the next pass on, even while that config fails to apply.
Reloading
Section titled “Reloading”impl Core { pub async fn reload_with(&self, read: impl FnOnce() -> Read) -> Result<Reload, LoadError>;}
impl Instance { pub async fn reload(&self) -> Result<Reload, LoadError>;}Three callers reload an Instance, and all of them go through Instance::reload:
| Caller | Path | What it does with the result |
|---|---|---|
| The watcher task | spawn_watcher |
Ignores it |
POST /v1/reload |
webclient/src/router.rs → reload → ConfigFile::reload |
instance.reload().await?.report(), answered as JSON or as an error status |
PUT /v1/routes/{group} |
webclient/src/router.rs → set |
Holds the router’s edits lock across ConfigFile::set_route (which checks and writes the file) and the ConfigFile::reload after it, and answers with the reload’s result and the file’s new ETag |
The router’s edits lock serialises a PUT with the reload that follows it, but, as ConfigControl’s doc comment says, not against reloads from elsewhere, such as the watcher; Core’s last lock does that. A reload that reads the bytes another reload last recorded returns Unchanged. The REST API answers that with 200 {"status":"unchanged"} when no failure was recorded with those bytes, and with 422 when one was. set puts the file’s new ETag on its response either way, because the file was written whether or not the reload after it succeeded.
The edit itself (ConfigFile::set_route: the If-Match check, the check_bytes dry run, the re-read that refuses with Stale when a hand edit landed during the check, and the atomic write) is on the REST API page.
Step by step
Section titled “Step by step”sequenceDiagram
participant C as caller
participant K as Core::reload_with
participant R as read closure
participant S as Supervisor
participant A as api watch
C->>K: reload_with(read)
K->>K: lock last
K->>R: read()
R-->>K: Sources and the parse
alt the read failed
K-->>C: Err(Config), logged as reload plus the error
else sources equal last.sources
K-->>C: Ok(Unchanged) with the recorded failure
else the sources changed
K->>K: config::effective, then the lowering
alt merge or lowering failed
K-->>C: Err(Config), the running config kept
else built
K->>S: apply(spec)
alt applied
S-->>K: ApplyReport
K->>A: send_if_modified(built.api)
K->>K: last becomes these sources, no failure
K-->>C: Ok(Applied(report))
else refused
S-->>K: ApplyError
K-->>C: Err(Apply), the running config kept
end
end
end
- Lock.
self.last.lock().await. The guard is held to the end of the call, across the read and across the supervisor’s apply. The read, the comparison, the apply and the update oflastare therefore one step: two callers cannot both read one change and apply it twice, and a slower caller cannot apply older files after newer ones. The doc comment says it: “Load whatreadgives, read while no other reload runs.” - Read.
read()runs under the lock. ForInstance::reloadit isconfig::read_sources(&self.path), with a read error rewritten ascannot read <path>: <e>. A read error is logged asreload: <e>atERROR(for the app,reload: cannot read <path>: <e>) and returned asLoadError::Config. - Compare. When
sources == last.sources, the call returnsOk(Reload::Unchanged { failed: last.failed.clone() })without logging anything and without touching the supervisor. - Merge and lower.
config::effective(parsed, &sources).and_then(|cfg| (self.lowering)(&cfg)). This is where a TOML error surfaces, together with a subscribe-file error, an unknown protocol, a badsettingstable and an[api]mistake. On error the call logsreload: <e>; keeping the running configatERRORand returnsLoadError::Config. - Apply.
self.supervisor.apply(built.spec), which isapply_withandApplyOptions::default():allow_disruptiveisfalse. The supervisor’s actor validates, plans, refuses a disruptive plan, prepares and commits (see Planning and applying a change). - Applied. The call logs
config reloaded: <summary>atINFO, publishes the API options withsend_if_modified(below), setslastto these sources withfailed: None, and returnsReload::Applied(report). - Refused.
ApplyError::Disruptive { inbound, reason }is logged asreload refused, keeping the running config: inbound <inbound>: <reason>, which would end its live connections; restart to apply it. Every otherApplyErroris logged asreload: <e>; keeping the running config. Both are atERROR, and the error is returned asLoadError::Apply. The supervisor refuses a spec as a whole, so what runs is exactly what ran before;a_refused_spec_changes_nothingpins that for a validation error and for a bind that fails while preparing.
Publishing the API options
Section titled “Publishing the API options”self.api.send_if_modified(|api| { let changed = *api != built.api; *api = built.api; changed});The watch value is replaced after every applied reload, but receivers are woken only when the new Option<ApiOptions> differs from the old one: a different listen, a different secret, or the section appearing or going. A reload that is refused or fails to lower publishes nothing, so “one that is refused changes the API no more than the rest”. The binary acts on the change in follow_api, not here (see The API follower).
Identical sources
Section titled “Identical sources”last holds the sources of the last recorded reload and that reload’s failure text, if it had one. The comparison in step 3 is Sources’ derived PartialEq: the config file’s bytes and, when the config names one, the subscribe file’s bytes. Only the bytes count:
- A reload whose sources equal
last.sourcesreturnsUnchanged { failed }without touching the supervisor: silent, with no parse beyond the oneread_sourcesdoes, and no apply. - An edit to the subscribe file alone is a change, and so is an edit that only adds a comment or reorders keys. Such a reload is lowered and applied like any other (the third captured line in
summary). laststarts as the sourcesCore::startloaded, with no failure.
Unchanged { failed } carries the failure text recorded with those sources, if any. The text is the error’s Display, not the log line; ApplyError::Disruptive, for example, displays as inbound <tag>: <reason>; this ends its live connections and needs allow_disruptive. The watcher ignores Unchanged; the REST API answers 422 with that text followed by (the files have not changed since this was found).
Outcomes and log lines
Section titled “Outcomes and log lines”| Case | Log line (ERROR unless noted) |
reload_with returns |
REST API answer |
|---|---|---|---|
| The files cannot be read | reload: cannot read <path>: <e> |
Err(LoadError::Config) |
422 or 500, by the error’s kind |
| Sources equal to the recorded ones, no failure recorded | none | Ok(Unchanged { failed: None }) |
200 {"status":"unchanged"} |
| Sources equal to the recorded ones, a failure recorded | none | Ok(Unchanged { failed: Some(..) }) |
422 <reason> (the files have not changed since this was found) |
| TOML, merge or lowering error | reload: <e>; keeping the running config |
Err(LoadError::Config) |
422 for InvalidData and InvalidInput, 500 for other kinds |
| A disruptive change | reload refused, keeping the running config: inbound <tag>: <reason>, which would end its live connections; restart to apply it |
Err(LoadError::Apply(Disruptive)) |
422 |
| A listener that cannot be bound | reload: inbound <tag>: binding <bind> failed: <e>; keeping the running config |
Err(LoadError::Apply(Bind)) |
500 |
| Any other refusal | reload: <e>; keeping the running config |
Err(LoadError::Apply(..)) |
422, or 500 for Stopped |
| Applied | INFO config reloaded: <summary> |
Ok(Applied(report)) |
200 with the report |
Refusals captured from the pinned binary (timestamps removed; a TOML error is shown by its first line only):
ERROR etemenanki_app::instance: reload: route references unknown outbound nowhere; keeping the running configERROR etemenanki_app::instance: reload: outbound x: unknown protocol "carrier-pigeon"; keeping the running configERROR etemenanki_app::instance: reload: duplicate outbound tag direct; keeping the running configERROR etemenanki_app::instance: reload: [api] listen 0.0.0.0:9: an API reachable from other hosts needs a secret; listen on a loopback address or a unix socket, or set one; keeping the running configERROR etemenanki_app::instance: reload: TOML parse error at line 5, column 1The first and third are supervisor refusals (ApplyError::UnknownReference, ApplyError::DuplicateTag); the second and fourth are lowering errors; the last is a parse error, which config::effective reports at the lowering step.
Disruptive changes are refused
Section titled “Disruptive changes are refused”The supervisor’s plan marks some inbound changes as disruptive (Step::Disrupt in supervisor/src/topology/spec_plan/plan.rs). Supervisor::apply passes ApplyOptions::default(), whose allow_disruptive is false, so Actor::apply refuses such a spec before preparing anything and names the first disruption the plan lists (plan.disruptions().next()). Which changes the plan marks is on Planning and applying a change.
For etemenanki-app this means:
- a Hysteria 2 listener whose
obfschanges is refused on reload, and a restart applies it. katana applies withallow_disruptive, so the same change applies there (see Node manager). - a TUN inbound whose device or settings change is refused on reload, and a restart applies it.
The refusal is for the whole spec, as for any other ApplyError: nothing else in it is applied either. The <reason> in the log line is one of these texts:
the hysteria2 obfuscation changed; connected clients cannot follow the new keythe tun inbound's settings changed; its runtime restarts and its flows endthe tun device changed; the old device and its flows endSo a Hysteria 2 obfuscation change reads:
ERROR etemenanki_app::instance: reload refused, keeping the running config: inbound <tag>: the hysteria2 obfuscation changed; connected clients cannot follow the new key, which would end its live connections; restart to apply itThe supervisor-level refusal is pinned by a_hysteria2_obfuscation_change_needs_allow_disruptive in supervisor/tests/hot_swap.rs.
What a reload keeps
Section titled “What a reload keeps”An applied reload is an apply of a whole new spec, and the supervisor keeps everything whose spec did not change. From the app’s side:
| What | Across an applied reload | Details |
|---|---|---|
| Sessions | Not tied to their listener: they run under the supervisor’s root token, so stopping or swapping a listener does not cancel them. Removing an inbound or renaming its tag closes its sessions (CloseSessions). The removal policy applies to sessions bound to a removed user’s principal; the app lowers Policies::default(), whose UserRemovalPolicy is Close. |
Users, principals and sessions, Planning and applying a change |
| Stream listeners | A listener whose BindSpec is unchanged keeps its socket, its accept loop and the loop’s admission limits (the semaphores run_stream_inbound creates). A changed inbound on the same bind gets a new handler for new connections (swapped). A new inbound is bound during prepare. How a changed bind is handled is on the serving page. |
Listeners and the serve loop |
| A Hysteria 2 listener on the same port | Keeps its QUIC endpoint and keeps serving on its UDP port. Its circuit budget carries over only while max_circuits is unchanged (same_circuits). The rest of the swap is on the serving page. |
Listeners and the serve loop |
| Outbounds | An outbound whose spec is unchanged, with an unchanged DNS spec unless it is a blackhole, is reused with the state its connector holds. A changed outbound is built as a new version, and new flows use it. | Outbounds, UDP fan-out and balancers, Planning and applying a change |
| DNS | The resolvers and their answer cache are kept while the DnsSpec is unchanged |
Name resolution and the DNS service |
| Balancers | Member health and running probes are kept when the balancer and all its members are unchanged | Outbounds, UDP fan-out and balancers |
| The route table | Reused when the RouteSpec is unchanged. New connections follow the new plane; open ones keep the target they were opened on. |
The plane: routing each flow |
[api] |
Follows the applied config: started, stopped or rebound | The API follower |
[log].level |
Not reloaded; init_tracing runs once |
Tracing |
| Supervisor settings | Fixed for the process: the builder’s defaults from Supervisor::builder(), which are a sample interval of DEFAULT_SAMPLE_INTERVAL (1 s), no usage sink and SocketOptions::default() |
The supervisor |
| Apply options | Every reload applies with ApplyOptions::default() (Supervisor::apply), and the start with the builder’s default, which is the same |
Disruptive changes are refused |
The full table, with the plan steps behind each row, is in the “What an apply keeps” section of Planning and applying a change. Two app tests pin the session rows end to end: an_open_connection_keeps_transferring_across_a_reload (a route change sends new connections to a blackhole while an open SOCKS connection keeps echoing) and a_reload_that_removes_a_user_closes_only_their_connections (removing one SOCKS account closes that account’s open connection and refuses its next login, while another account’s connection keeps echoing).
The API follower
Section titled “The API follower”async fn follow_api( instance: Arc<Instance>, mut options: watch::Receiver<Option<ApiOptions>>, mut running: Option<Api>, mut stopped: oneshot::Receiver<()>,);follow_api keeps the REST API as the running config’s [api] says, from the API main started (running) until the stopped signal:
flowchart TB
W["select: stop signal, or options.changed()"] -- "stop signal, or the sender is gone" --> E["stop the running API, return"]
W -- "changed" --> B["wanted = borrow_and_update()"]
B --> S["stop the running API, if any"]
S --> Q{"wanted?"}
Q -- "Some" --> ST["start_api(wanted)"]
ST -- "Ok" --> R["running = the new API"]
ST -- "Err" --> L1["log: stays off until [api] changes"]
Q -- "None" --> L2["log: API stopped"]
R --> W
L1 --> W
L2 --> W
- Stop before bind. The old API is stopped before the new one binds, “so the same address can be taken again”. A change of
secretalone therefore closes the listener and binds the same port again. Between the two, connections to the API are refused. - Latest value only. A watch keeps one value. When several applied reloads change
[api]while the follower is busy, it wakes once and acts on the latest options. - A reload made by the API finishes first. The restart runs in this task, “not in the reload that published it, so a reload made by the API itself finishes its request before the API stops”.
Core::reload_withonly publishes; the request that caused the reload returns its answer, and the follower’sApi::stopthen waits for it to complete. If the restart ran inside the reload, the request would wait on the stop, and the stop on the request, until the grace ran out. - A failed bind is not retried. When the new listener cannot bind, the follower logs
API on <listen>: <e>; it stays off until [api] changesand keepsrunningempty. It acts again on the next change of the published options. - The sender outlives the loop. The
watch::Senderlives inCore, and the follower holds anArc<Instance>, sochanged()returning an error cannot happen while the task runs; thebreakfor it is defensive.
Api::stop sends on the stop channel, which completes the graceful-shutdown future passed to ApiListener::serve: axum stops accepting and lets the requests in flight finish. It then waits up to SHUTDOWN_GRACE for the server task:
| Server task result | Log line |
|---|---|
Finished with Ok(()) in time |
none |
Finished with Err(e) |
API: <e> at ERROR |
Panicked or was cancelled (JoinError) |
API task failed: <e> at ERROR |
| Still running after 5 s | task.abort() on the server task, then API: requests still in flight when it stopped were dropped at WARN |
The server task is the one running ApiListener::serve, which awaits axum::serve(..).with_graceful_shutdown(..).
The follower’s other log lines are API listening on <addr> (INFO, from start_api) and API stopped: the config no longer has [api] (INFO). The second is printed whenever the published options become None, whether or not an API was running at the time: after a failed rebind, for example, none was. The API’s endpoints and authentication are on the REST API page; the operator’s view is in the user guide’s API page.
Shutting down
Section titled “Shutting down”async fn wait_for_shutdown();On Unix, wait_for_shutdown registers a SIGTERM stream with tokio::signal::unix::signal(SignalKind::terminate()) and returns on whichever comes first, tokio::signal::ctrl_c() (SIGINT) or SIGTERM. If the SIGTERM registration fails, it waits for SIGINT alone. On other platforms it waits for Ctrl-C. Tokio keeps its handlers installed for the life of the process, so a second SIGINT or SIGTERM during the shutdown does not cut it short; SIGKILL does, and then no cleanup runs.
sequenceDiagram participant M as main participant F as follow_api task participant API as API server task participant S as supervisor actor M->>M: SIGINT or SIGTERM, log shutting down M->>F: stop_api.send(()) F->>API: Api::stop, then wait up to 5 s API-->>F: finished, or aborted with a warning F-->>M: the task returns M->>S: Instance::shutdown(SHUTDOWN_GRACE) S->>S: stop listeners and probes, wait up to 5 s for connections S->>S: cancel the root, wait for every task, stop the sampler S-->>M: done M->>M: return ExitCode::SUCCESS
mainlogsshutting downatINFO.stop_api.send(())signals the follower, andmainawaits its task, discarding theJoinHandle’s result (let _ = api.await). The follower leaves its loop at aselect!: when it is in the middle of a restart, the restart completes first.tokio::select!picks at random among the branches that are ready, so when a change of the options is pending as well, one more restart can run before the loop sees the stop signal. The follower then stops the running API, if any, with up toSHUTDOWN_GRACEfor requests in flight.instance.shutdown(SHUTDOWN_GRACE)runsSupervisor::shutdown(Duration::from_secs(5)).Actor::shutdowncallsstop()on every listener (an accept loop that ends drops its listener, and a Unix socket file with it), cancels every balancer’s probe token, closes the task tracker, waits up to the grace withtokio::time::timeout(grace, tracker.wait()), cancels its root token so every remaining connection ends, waits for every tracked task, clears its listener handles, cancelsbackground_stop, and awaits each background task: the sampler, and a usage sink’s push task, which the app does not have. The sequence is on The supervisor.mainreturnsExitCode::SUCCESS. The runtime shuts down and drops the watcher task with itsnotifywatcher.
The API is stopped first, so a request that finishes within the API’s grace is answered before the supervisor starts shutting down. A watcher pass can still run during step 3: the supervisor’s actor handles commands in order, so an apply already sent completes before the shutdown, and one sent after it fails with ApplyError::Stopped (reload: the supervisor has shut down; keeping the running config).
| How the process ends | Exit status | Shutdown sequence |
|---|---|---|
--test, accepted |
0 | none; nothing was started |
--test, refused |
1 | none |
| The start is refused | 1 | none; the refused apply released what it bound |
| The first API bind fails | 1 | Instance::shutdown(Duration::ZERO) |
SIGINT or SIGTERM |
0 | API stop, then Instance::shutdown(SHUTDOWN_GRACE) |
SIGHUP |
killed by the signal (129 in a shell) | none |
SIGKILL |
killed by the signal | none |
A signal-driven shutdown takes up to 5 s for the API’s requests plus up to 5 s for live connections, and then as long as the remaining tasks take to notice the root token’s cancellation.
Invariants
Section titled “Invariants”| Invariant | Enforced by | Pinned by |
|---|---|---|
| A start runs the whole config or nothing | Core::start returns any ApplyError from SupervisorBuilder::start, which fails before it spawns the actor and the sampler; a failed prepare drops what it bound |
a_refused_spec_changes_nothing (for the prepare); no app-level test |
Reloads of one Core never overlap |
last is a tokio::sync::Mutex held across the read and the apply |
none directly |
| Sources equal to the recorded ones are not applied again | sources == last.sources returns Unchanged |
a_reload_after_the_apis_own_finds_nothing_changed (app/tests/unit/api.rs) |
| A reload that is not applied changes nothing that runs, the API included | The supervisor’s prepare and commit; send_if_modified only after Ok |
a_refused_spec_changes_nothing (supervisor/tests/hot_swap.rs) |
| An open connection whose inbound and user are unchanged keeps relaying across a reload that changes routing | Sessions run under the supervisor’s root token, not their listener’s; new connections take the new plane | an_open_connection_keeps_transferring_across_a_reload; an_established_connection_survives_an_apply_that_keeps_its_inbound (supervisor level) |
| A protocol change on the same port keeps the socket, and a SOCKS connection accepted before a swap to HTTP keeps speaking SOCKS | Listener::swap puts the new handler in the listener’s watch for new connections (swapped); the accept loop runs on |
a_protocol_change_on_the_same_port_keeps_the_socket_and_old_connections (supervisor level) |
| Removing a SOCKS account closes that account’s open connection and leaves another account’s running | The removal policy the app lowers, UserRemovalPolicy::Close |
a_reload_that_removes_a_user_closes_only_their_connections |
| Removing an inbound closes its sessions | Step::CloseSessions in the commit |
removing_an_inbound_closes_its_sessions |
| A changed Hysteria 2 inbound on the same port keeps serving on it | The listener keeps its QUIC endpoint | a_reload_keeps_the_inbound_serving_on_its_udp_port |
| The app never applies a disruptive change | Supervisor::apply with ApplyOptions::default() |
a_hysteria2_obfuscation_change_needs_allow_disruptive (the supervisor’s refusal); no app-level test |
The API follows the applied [api]: gone when it goes, back where it comes back, rebound when only the secret changes |
send_if_modified and follow_api |
the_api_follows_the_api_section_across_reloads |
A config without [api] binds nothing beyond its inbounds |
main starts the API only for Some options |
a_config_without_api_binds_nothing_new (Linux only) |
| A reload made by the API answers before the API restarts | The restart runs in follow_api, not in reload_with |
none |
Tracing is set up once, from RUST_LOG or [log].level |
init_tracing at the top of main |
none |
| A clean shutdown releases what the inbounds own | Supervisor::shutdown stops every listener |
socks_over_a_unix_socket_relays_and_cleans_up (the socket file is gone after SIGTERM) |
| Shutdown is not held up by the supervisor’s own timers or by a stalled Hysteria 2 handshake | root.cancel() after the grace ends grace timers and closes a Hysteria 2 endpoint |
shutdown_ends_pending_grace_timers, shutdown_does_not_wait_out_a_stalled_hysteria2_handshake (supervisor level) |
Failure paths and cancellation
Section titled “Failure paths and cancellation”At start
Section titled “At start”| Failure | Log line | Effect |
|---|---|---|
The config or subscribe file cannot be read, a TOML or lowering error, or any ApplyError |
failed to start: <e> |
Exit status 1; nothing keeps running |
| A listener cannot be bound | failed to start: inbound <tag>: binding <bind> failed: <e> |
Exit status 1; listeners already bound in the same prepare are released |
| The REST API cannot be bound | failed to start: API on <listen>: <e> |
Instance::shutdown(Duration::ZERO), exit status 1 |
| The watcher cannot be created, or the config’s directory cannot be watched | config hot-reload disabled: <e> |
The process runs without file-triggered reloads; the REST API can still reload |
During a reload
Section titled “During a reload”Every failure keeps the running config, as the outcomes table shows. A subscribe file’s directory that cannot be watched is logged as cannot watch <dir> for changes: <e> and does not stop anything else.
In the API follower
Section titled “In the API follower”A rebind that fails leaves the API off with API on <listen>: <e>; it stays off until [api] changes. The proxy is not affected. A stop that overruns SHUTDOWN_GRACE calls task.abort() on the server task and logs API: requests still in flight when it stopped were dropped at WARN.
Cancellation
Section titled “Cancellation”- Only the watcher awaits
reload_withto completion. An API request’s handler is dropped when its client goes away while the request is in flight (hyper sees EOF on an HTTP/1 connection, or a reset of the HTTP/2 stream), andreload_withis dropped with it at whichever await it has reached. How the actor treats a command whose caller has gone is on The supervisor. - The watcher task is never cancelled explicitly. It ends when the runtime shuts down after
mainreturns. - The follower ends on the
stoppedsignal frommain.mainsends it only after a signal, and awaits the follower before it shuts the supervisor down. Api::stopbounds its wait withtokio::time::timeout(SHUTDOWN_GRACE, &mut task)and aborts the server task on timeout.
Limits
Section titled “Limits”| Constant | Value | Defined in | Meaning |
|---|---|---|---|
SHUTDOWN_GRACE |
5 s | app/src/main.rs |
Grace for the API’s requests at every API stop, and for live connections at shutdown |
| Grace after a failed API start | Duration::ZERO |
app/src/main.rs → main |
Closes what the inbounds accepted at once |
DEFAULT_SAMPLE_INTERVAL |
1 s | supervisor/src/track/sampler.rs |
The stats sampler’s tick; the app does not change it |
| Supervisor command queue | 16 | supervisor/src/supervisor.rs → SupervisorBuilder::start, mpsc::channel(16) |
Commands waiting for the actor, a reload’s apply among them |
DEFAULT_LISTEN |
127.0.0.1:9090 |
webclient/src/listen.rs, used by ApiConfig’s serde default default_api_listen |
The API’s address when [api] has no listen |
| Disruptions reported per refusal | 1 | Actor::apply → plan.disruptions().next() |
Only the first disruptive inbound is named |
The process-wide limits of the proxy itself are collected on Limits, timeouts and memory.
The end-to-end tests run the real binary. app/tests/support/mod.rs → spawn_app(dir, "app.toml") starts CARGO_BIN_EXE_etemenanki-app with -c app.toml and the test directory as its working directory, so the config path is a bare file name and the watched directory is ., with standard output and standard error discarded. The returned Proc kills the child with SIGKILL when dropped; Proc::terminate sends SIGTERM and polls for the exit for up to 10 s, for the tests where what the process does on the way out is the thing under test. The reload tests rewrite the config file in place with std::fs::write and then poll a probe every 100 ms for up to 20 s (eventually) until a new connection shows that the new config took effect, so what the open connection does next is observed after the reload.
| File | Test | Behaviour it pins |
|---|---|---|
app/tests/integration/e2e_reload.rs |
an_open_connection_keeps_transferring_across_a_reload |
A SOCKS connection opened before a reload that adds a 127.0.0.0/8 → block rule keeps echoing, while new connections are blackholed |
app/tests/integration/e2e_reload.rs |
a_reload_that_removes_a_user_closes_only_their_connections |
Removing alice from the SOCKS accounts closes her open connection within 10 s and refuses her next login; bob’s open connection keeps echoing |
app/tests/integration/e2e_api.rs |
the_api_follows_the_api_section_across_reloads |
[api] removed: the port stops answering. Added on another port with secret = "one": 200 with the token, 401 without. Only the secret changed to "two": the same port answers the new token and refuses the old one |
app/tests/integration/e2e_api.rs |
a_config_without_api_binds_nothing_new |
Read from /proc: a config without [api] listens on its SOCKS port only; the same config with [api] adds exactly the API’s port. Linux only |
app/tests/integration/e2e_api.rs |
a_switch_sends_new_flows_of_the_group_through_the_new_target |
PUT /v1/routes/Local answers "status":"applied", so the edit’s reload applied |
app/tests/integration/e2e_hysteria_inbound.rs |
a_reload_keeps_the_inbound_serving_on_its_udp_port |
A Hysteria 2 inbound rewritten from udp = false to udp = true on the same port, with a real hysteria client connected, still serves through that client after the reload. It waits 6 s after the write, then retries 20 times at 500 ms. Skips when go or the upstream build is unavailable |
app/tests/integration/e2e_unix.rs |
socks_over_a_unix_socket_relays_and_cleans_up |
After Proc::terminate (SIGTERM), the inbound’s socket file is gone |
app/tests/integration/e2e_tun.rs |
a_routed_connect_is_answered_while_the_app_runs |
Once the process has exited after Proc::terminate, a connect routed into the TUN device no longer succeeds: the interface and its route are gone. The kernel removes a created device when its descriptor closes, so this does not tell a clean shutdown from any other exit. Needs CAP_NET_ADMIN, and skips without it |
app/tests/unit/api.rs |
a_reload_after_the_apis_own_finds_nothing_changed |
After a PUT applied an edit, POST /v1/reload answers 200 {"status":"unchanged"}: a reload of the bytes the edit’s reload recorded does not apply them again |
app/tests/unit/api.rs |
a_switch_edits_the_file_reloads_and_returns_the_new_etag |
A PUT over an Instance started in-process reloads ("status":"applied") and returns the file’s new ETag |
app/tests/unit/api.rs |
api_options_come_from_the_api_section |
api::options, part of build: no [api] is None; an empty [api] listens on DEFAULT_LISTEN with no secret; a value containing / is a Unix socket path; a non-local listen with a secret is accepted |
app/tests/unit/api.rs |
an_api_open_to_other_hosts_needs_a_secret |
A non-local listen without a secret, an empty secret on loopback, a host name as listen and an unknown key in [api] are all refused |
supervisor/tests/hot_swap.rs |
a_refused_spec_changes_nothing |
A spec refused by validation (UnknownReference) or by a bind while preparing (Bind) leaves the epoch, the routes, the users and the listeners as they were, and releases a listener bound in that prepare |
supervisor/tests/hot_swap.rs |
a_hysteria2_obfuscation_change_needs_allow_disruptive |
Without allow_disruptive, an obfuscation change is ApplyError::Disruptive for hy2-in |
supervisor/tests/hot_swap.rs |
an_established_connection_survives_an_apply_that_keeps_its_inbound |
An apply that adds an outbound and a rule reports the inbound as reused, neither swapped nor rebound, and an open SOCKS connection keeps echoing |
supervisor/tests/hot_swap.rs |
a_protocol_change_on_the_same_port_keeps_the_socket_and_old_connections |
SOCKS to HTTP on the same port is swapped, not rebound; an open SOCKS connection keeps echoing, a connection accepted before the swap still gets the SOCKS handler, and a new one speaks HTTP |
supervisor/tests/hot_swap.rs |
removing_an_inbound_closes_its_sessions |
Removing one of two SOCKS inbounds reports it as removed and closes its open connection; the other inbound’s connection keeps echoing |
supervisor/tests/hot_swap.rs |
shutdown_ends_pending_grace_timers |
A removal and a drain each waiting an hour do not hold up shutdown(Duration::ZERO), and the session in its removal grace is closed |
supervisor/tests/hot_swap.rs |
shutdown_does_not_wait_out_a_stalled_hysteria2_handshake |
A Hysteria 2 client stuck in its QUIC handshake does not hold up the server’s shutdown |
ffi/tests/proxy.rs |
set_route_and_reload_switch_like_the_rest_api |
A Core started by the FFI with its own Lowering, not by Instance, applies a route switch through reload_with (see the FFI page) |
No test covers the watcher following a newly named subscribe file, the reload log lines, Reload::report with a recorded failure, the disruptive refusal at the app level, a failed API start or rebind, or SIGHUP. A change to those paths should add one; Testing describes how the suites are laid out.