跳转到内容

etemenanki-app:运行、重载与关停

源码文件:25 个 · 核对版本 Etemenanki 555b7df
  • Etemenanki/app/src/main.rs
  • Etemenanki/app/src/instance.rs
  • Etemenanki/app/src/lib.rs
  • Etemenanki/app/src/config.rs
  • Etemenanki/app/src/api.rs
  • Etemenanki/supervisor/src/lib.rs
  • Etemenanki/supervisor/src/supervisor.rs
  • Etemenanki/supervisor/src/build/apply.rs
  • Etemenanki/supervisor/src/topology/spec_plan/plan.rs
  • Etemenanki/supervisor/src/topology/inbound/mod.rs
  • Etemenanki/supervisor/src/system/listener.rs
  • Etemenanki/supervisor/src/entity/id.rs
  • Etemenanki/supervisor/src/policy.rs
  • Etemenanki/supervisor/src/track/sampler.rs
  • Etemenanki/webclient/src/lib.rs
  • Etemenanki/webclient/src/listen.rs
  • Etemenanki/webclient/src/router.rs
  • Etemenanki/app/tests/integration/e2e_reload.rs
  • Etemenanki/app/tests/integration/e2e_api.rs
  • Etemenanki/app/tests/integration/e2e_hysteria_inbound.rs
  • Etemenanki/app/tests/integration/e2e_unix.rs
  • Etemenanki/app/tests/integration/e2e_tun.rs
  • Etemenanki/app/tests/support/mod.rs
  • Etemenanki/app/tests/unit/api.rs
  • Etemenanki/supervisor/tests/hot_swap.rs

etemenanki-app 是包在一个 supervisor(监管器)外面的一层薄薄的二进制程序。app/src/main.rs 解析命令行,设置日志,启动一个 Instance,在配置有 [api] 段时启动 REST API,监视配置文件及其指定的订阅文件,然后等待信号。app/src/instance.rs 负责加载本身:它读取文件,把它们合并后降为一个 Spec,把这个 spec(期望状态)交给 supervisor,记录发生了什么,并记住最近一次记录的 sources(配置文件与订阅文件的原始字节)。每次重载都要经过 supervisor 的校验、规划、准备(prepare)和提交(commit)。会话不绑定在它的监听器上,所以重载不会仅仅因为运行过就关闭会话;被拒绝的重载让正在运行的一切保持原样。

本页从 main 一直跟到进程退出:命令行、tracing、--test、启动、文件监视器、逐步展开的 Core::reload_with 及其可能打印的每一行日志、一次重载保留什么、API 跟随器(follow_api),以及关停。本页面向要修改 main.rs 或 instance.rs、要在 Core 之上构建另一个前端程序,或者要解释某次重载打印了什么的读者。配置如何变成 spec 见配置页,订阅合并见订阅页,API 的端点见 REST API 页。一次应用(apply)在 supervisor 内部做什么,见规划并应用变更。

组件 负责 不负责
app/src/main.rs → main 解析 Args、init_tracing、--test 分支、Instance::start、第一次启动 API、spawn API 跟随器和文件监视器、等待 SIGINT 或 SIGTERM,以及关停顺序 读取 [log].level 以外的任何配置键
app/src/main.rs → spawn_watcher、sync_watches、watch_dir notify 监视器、它监视的目录集合,以及把监视器的 tick 变成 Instance::reload 调用的任务 判断文件是否变化;这由 Core::reload_with 决定
app/src/main.rs → follow_api、start_api、Api 随正在运行的配置的 [api] 变化,启动、停止和重新绑定 REST API 判断 [api] 何时变化;Core 在一次应用之后发布这一点
app/src/instance.rs → Core 从一个 Read 启动 supervisor,通过一个 Lowering 重载它,跳过与上次记录相同的 sources,记录每种结果,并发布正在运行的配置的 API 选项 知道文件从哪里来
app/src/instance.rs → Instance 配置路径、从磁盘读取 sources,以及值得监视的文件列表 自己监视任何东西
app/src/instance.rs → check、check_bytes --test 背后的试运行,以及每次 API 路由编辑背后的试运行 绑定监听器或启动任务
etemenanki_supervisor::Supervisor 校验、规划、准备和提交,监听器、会话、出站、关停 了解文件、TOML、API 或信号:这个 crate “对配置文件、配置格式以及重载如何触发一无所知”(supervisor/src/lib.rs)

app/src/lib.rs 把这种分工写得很明确:库 crate 是“Etemenanki 的 TOML 前端程序”(api、config、instance、lower、routes、subscribe、transport),而“etemenanki-app 二进制是上面一层薄薄的胶水:CLI、tracing、触发 instance::Instance::reload 的文件监视器,以及有 [api] 段时的 REST API”。

这个二进制程序不会:

  • 安装 SIGHUP 信号处理函数。触发重载的只有文件监视器,以及 REST API 上的 POST /v1/reload 和 PUT /v1/routes/{group},别无其他。SIGHUP 保持默认动作,会立即终止进程,不走关停流程(shell 报告的退出状态为 129)。
  • 把日志写到标准输出以外的任何地方。tracing_subscriber::fmt() 默认写到 stdout,init_tracing 不改变 writer。
  • 修改任何跨 spec 存续的 supervisor 设置。它用 Supervisor::builder() 及其全部默认值启动 supervisor(见一次重载保留什么)。
app/src/main.rs
/// A minimal Xray-core-style proxy runtime.
#[derive(Parser)]
#[command(version, about)]
struct Args {
/// Path to the TOML configuration file.
#[arg(short, long, default_value = "config.toml")]
config: PathBuf,
/// Validate the configuration and exit without binding any listener.
#[arg(long)]
test: bool,
}

锁定版本二进制的 --help 输出:

A minimal Xray-core-style proxy runtime
Usage: etemenanki-app [OPTIONS]
Options:
-c, --config <CONFIG> Path to the TOML configuration file [default: config.toml]
--test Validate the configuration and exit without binding any listener
-h, --help Print help
-V, --version Print version

--version 打印 etemenanki-app 2.1.1:clap 在 app/Cargo.toml 中的版本号前面加上二进制名。REST API 在 /v1/version 报告同一个版本,内容是不加修饰的 env!("CARGO_PKG_VERSION"):{"version":"2.1.1"}。配置路径按给定的原样保存。相对路径在每次读取时都相对于工作目录解析,而这个二进制从不改变自己的工作目录。面向用户的参考见命令行。

app/src/main.rs
/// How long live connections get to finish on shutdown before they are
/// closed.
const SHUTDOWN_GRACE: Duration = Duration::from_secs(5);

它用在两处:一是 REST API 每次停止时(Api::stop)留给进行中请求完成的时间,二是传给 Instance::shutdown 的宽限期,supervisor 用它等待存活的连接。

app/src/instance.rs
#[derive(Debug, thiserror::Error)]
pub enum LoadError {
/// The files could not be read, or what they say could not be lowered.
#[error(transparent)]
Config(#[from] io::Error),
/// The supervisor refused the spec.
#[error(transparent)]
Apply(#[from] ApplyError),
}
impl From<LoadError> for ControlError { /* 见下表 */ }

两个变体都是 transparent,所以每一行日志和每一个 API 错误都原样显示内层错误的文本。From<LoadError> for ControlError 为 REST API 按过错归属给失败分类:Invalid 应答 HTTP 422,Failed 应答 HTTP 500。

LoadError ControlError 原因
Config(e),且 e.kind() 为 InvalidData 或 InvalidInput Invalid(e.to_string()) 过错在配置内容:TOML 错误(config::parse_bytes 使用 InvalidData)、lowering(降为 spec)时的拒绝,或 [api] 错误(InvalidInput)
其他任何 kind 的 Config(e) Failed(e.to_string()) 无法读取的文件,例如 NotFound 或 PermissionDenied
Apply(ApplyError::Bind { .. }) 或 Apply(ApplyError::Stopped) Failed(e.to_string()) 不是配置的过错:端口被别处占用,或 supervisor 已经关停
Apply(e),其他所有变体 Invalid(e.to_string()) 校验失败、构建失败或中断性变更

ApplyError 的各个变体及其文本列在校验与应用错误中。

app/src/instance.rs
pub struct Built {
pub spec: Spec<UserKey>,
pub api: Option<ApiOptions>,
}
pub fn build(cfg: &Config) -> io::Result<Built> {
Ok(Built {
spec: lower(cfg)?,
api: crate::api::options(cfg)?,
})
}
pub type Lowering = Arc<dyn Fn(&Config) -> io::Result<Built> + Send + Sync>;

一个 Built 就是一份配置产出的全部内容:supervisor 运行的 spec,以及配置有 [api] 段时 REST API 的选项。build 是 app 的 Lowering:先 lower(见配置页),再 api::options。后者在没有 [api] 段时返回 None;否则解析 [api].listen 并应用 ApiOptions::new,并把任何 OptionsError 包装成带 [api] 前缀的 InvalidInput 错误:

规则 检查者 错误文本
listen 默认为 DEFAULT_LISTEN,即 127.0.0.1:9090 ApiConfig 的 serde 默认值 default_api_listen(app/src/config.rs) 无
含有 / 的 listen 是 Unix socket 路径;其他值必须能解析为 ip:port Listen::from_str [api] listen "<s>": expected ip:port, or a path (containing a /) for a unix socket
设置了的 secret 不能为空,无论 listen 是什么 ApiOptions::new [api] secret is empty
既不是回环地址、也不是 Unix socket 的 listen 需要 secret ApiOptions::new(Listen::is_local) [api] listen <addr>: an API reachable from other hosts needs a secret; listen on a loopback address or a unix socket, or set one

因此 [api] 错误属于 lowering 错误,和其他配置错误一样,由 --test 发现、由重载报告。api_options_come_from_the_api_section 和 an_api_open_to_other_hosts_needs_a_secret(app/tests/unit/api.rs)固定了这些规则。Listen、ApiOptions 和 OptionsError 的完整说明见 REST API 页。

Lowering 是一个共享闭包而不是函数,这样 Core 就能以另一个 lowering 启动。app 传入的是 Arc::new(build)。拿到的是文件内容而不是路径的前端程序(例如移动端 FFI)运行自己的 Core,并使用自己的 Lowering(见 FFI 页)。

app/src/instance.rs
pub type Read = io::Result<(Sources, io::Result<Config>)>;
app/src/config.rs
#[derive(Clone, PartialEq, Default)]
pub struct Sources {
pub config: Vec<u8>,
pub subscribe: Option<Vec<u8>>,
}
pub fn read_sources(path: &Path) -> io::Result<(Sources, io::Result<Config>)>;
pub fn sources_from(config: Vec<u8>) -> io::Result<(Sources, io::Result<Config>)>;
pub fn effective(parsed: io::Result<Config>, sources: &Sources) -> io::Result<Config>;

一个 Read 是对 Core 所加载文件的一次读取。外层的 io::Result 是读取本身,文件无法汇集成 Sources 时失败(具体情形见配置页)。内层的是配置的解析结果,它和字节一起返回,因为订阅文件的路径要等配置解析之后才知道。解析失败的配置仍然是一组 sources;它的解析错误要到 lowering 这一步才由 config::effective 给出。Sources 保存配置文件的原始字节,以及配置指定了订阅文件时订阅文件的原始字节,并逐字节比较。这三个函数见配置页;effective 负责合并订阅文件(见订阅页)。

app/src/instance.rs
#[derive(Debug)]
pub enum Reload {
/// The files changed, and were applied.
Applied(ApplyReport),
/// The files are byte for byte the ones last tried, and were not tried
/// again. `failed` is why that last try failed, if it did.
Unchanged { failed: Option<String> },
}
impl Reload {
pub fn report(self) -> Result<ReloadReport, ControlError>;
}

report 决定 API 客户端得到什么答复:

Reload report() HTTP 应答
Applied(report) Ok(ReloadReport::Applied(report)) 200,{"status":"applied", "built":[…], "swapped":[…], "drained":[…], "rebound":[…], "restarted":[…], "removed":[…], "reused":[…]}
Unchanged { failed: None } Ok(ReloadReport::Unchanged) 200,{"status":"unchanged"}
Unchanged { failed: Some(reason) } Err(ControlError::Invalid(...)),文本为 <reason> (the files have not changed since this was found) 422

webclient/src/lib.rs 中 ReloadReport::Unchanged 的文档注释点出了常见情形:“文件监视器可能已经重载过它们了”。与失败一起记录的 sources 会以那次失败的文本再次被拒绝,和真的再试一次的结果一样。

app/src/instance.rs
pub struct Core {
supervisor: Supervisor<UserKey>,
lowering: Lowering,
api: tokio::sync::watch::Sender<Option<ApiOptions>>,
last: tokio::sync::Mutex<Last>,
}
struct Last {
sources: Sources,
failed: Option<String>,
}
impl Core {
pub async fn start(
builder: SupervisorBuilder<UserKey>,
read: Read,
lowering: Lowering,
) -> Result<Self, LoadError>;
pub fn supervisor(&self) -> &Supervisor<UserKey>;
pub fn api(&self) -> tokio::sync::watch::Receiver<Option<ApiOptions>>;
pub fn tracker(&self) -> Tracker;
pub async fn reload_with(&self, read: impl FnOnce() -> Read) -> Result<Reload, LoadError>;
pub async fn shutdown(&self, grace: Duration);
}
字段 类型 作用
supervisor Supervisor<UserKey> 正在运行的 supervisor。UserKey 就是 UserName,即 app 用来标识配置用户的名字(见配置页)。
lowering Lowering 合并后的 Config 如何变成 Built,在 Core 的整个生命周期内固定
api tokio::sync::watch::Sender<Option<ApiOptions>> 当前正在运行的配置的 API 选项:没有 [api] 时为 None。只有已应用的重载才会改变它。
last tokio::sync::Mutex<Last> 最近一次记录的 sources,以及与之一起记录的失败文本(如果有)。这把锁在整次重载期间都被持有,从而把重载串行化。

Core 就是加载本身,不管文件从哪里来。它持有一个 Supervisor 句柄。Supervisor 实现了 Clone,每个克隆都通过同一个命令通道驱动同一个 actor;最后一个句柄在没有调用 shutdown 的情况下被丢弃时,actor 会看到通道关闭,并以 Duration::ZERO 执行关停。发送失败或回复被丢弃的命令,会以 ApplyError::Stopped 返回(Supervisor::ask)。细节见 supervisor 概览。api() 为这个 watch 订阅一个新的 receiver;receiver 创建时就视为已看到当前值,所以它的 changed() 会等待下一次变化。tracker() 克隆 supervisor 的 Tracker,REST API 读取它时不经过 supervisor 的 actor(见流跟踪、统计与限速)。shutdown(grace) 就是 Supervisor::shutdown(grace)。

app/src/instance.rs
pub struct Instance {
path: PathBuf,
core: Core,
}
impl Instance {
pub async fn start(path: PathBuf) -> Result<Self, LoadError>;
pub fn api(&self) -> tokio::sync::watch::Receiver<Option<ApiOptions>>;
pub fn config_path(&self) -> &Path;
pub fn tracker(&self) -> Tracker;
pub fn watched_files(&self) -> Vec<PathBuf>;
pub async fn reload(&self) -> Result<Reload, LoadError>;
pub async fn shutdown(&self, grace: Duration);
}

Instance 是 app 的 Core,它从磁盘读取文件:

  • start(path) 即 Core::start(Supervisor::builder(), config::read_sources(&path), Arc::new(build))。
  • reload() 即 core.reload_with,所用的读取闭包调用 config::read_sources(&self.path),并把读取错误改写为 io::Error::new(e.kind(), format!("cannot read {}: {e}", ...))。新错误保留原来的 io::ErrorKind,所以 ControlError 仍能按 kind 给它分类;原错误只以文本形式保留下来。
  • watched_files() 返回配置路径;当磁盘上的配置文件能解析、且它的 [subscribe] 指定了 path 时,还返回该路径。它用 config::load 读取文件,而不是询问正在运行的配置,“这样即使指定新订阅文件的配置应用失败,也会跟随这个新订阅文件”。无法解析的配置只返回配置路径。
  • config_path() 是按原样给出的 -c 路径。REST API 的 ConfigFile 用它读取和重写文件。
app/src/instance.rs
fn summary(report: &ApplyReport) -> String;

summary 把一个 ApplyReport 变成 config loaded: 和 config reloaded: 后面的文本。它把每个非空分组打印为 <label> [<first>, <second>, …],用 ; 连接各组,顺序如下:

顺序 标签 条目 打印为
1 built Resource dns、route、outbound <tag>@v<n>、balancer <tag>、user set <tag>、inbound <tag>
2 swapped 入站 tag <tag>
3 drained OutboundId <tag>@v<n>
4 rebound 入站 tag <tag>
5 restarted 入站 tag <tag>
6 removed 入站 tag <tag>
7 reused Resource 同 built

所有分组都为空时,它打印 nothing to run。reused 虽然是 ApplyReport 的第一个字段,却排在最后,这样一行日志先写发生了什么变化。ApplyReport 的文档注释说明,期望 spec 中的每个资源都会出现在 reused 或 built 中;启动时没有可复用的东西,所以启动那一行把所有资源都列在 built 下。由于 app 从不允许中断性变更(见中断性变更会被拒绝),它的日志行里永远没有 restarted 分组。每个分组的含义见规划并应用变更。以下日志行取自锁定版本的 debug 二进制(已去掉时间戳):

INFO etemenanki_app::instance: config loaded: built [dns, outbound direct@v1, route]
INFO etemenanki_app::instance: config loaded: built [dns, outbound direct@v1, outbound n@v1, route]
INFO etemenanki_app::instance: config reloaded: reused [dns, outbound direct@v1, route]
INFO etemenanki_app::instance: config reloaded: built [outbound block@v1]; reused [dns, outbound direct@v1, route]
INFO etemenanki_app::instance: config reloaded: built [outbound block2@v1]; drained [block@v1]; reused [dns, outbound direct@v1, route]
INFO etemenanki_app::instance: config reloaded: built [dns, outbound direct@v2]; drained [direct@v1]; reused [route]
INFO etemenanki_app::instance: config reloaded: drained [n@v1]; reused [dns, outbound direct@v1, route]

第二行是一次带订阅文件的启动,订阅文件定义了一个节点 n。第三行是一次只加了注释的编辑之后的重载:sources 变了,所以 spec 被应用,而 supervisor 复用了一切。第六行是对 [dns] 的修改,它连带把 freedom 出站重建为一个新版本。最后一行是配置去掉了 [subscribe] 段。在 debug 级别,每次提交还会从 target etemenanki_supervisor::supervisor 记录 applied: <n> built, <n> reused, <n> swapped, <n> drained。

在这些日志前后,supervisor 自己还会记录两种日志。启动或已应用的重载所绑定的每个监听器,在提交阶段运行 Listener::start 时以 INFO 级别记录 inbound <tag> listening on <bind>(target 为 etemenanki_supervisor::system::listener),所以启动时这些行出现在 config loaded: 之前。创建自己设备的 TUN 入站在绑定时记录 inbound <tag> owns tun device <name>。两者都在监听器与服务循环中说明。

flowchart LR
  M["main"] --> ST["Instance::start"]
  ST --> CS["Core::start"]
  CS --> SUP["supervisor actor"]
  W["监视任务"] --> IR["Instance::reload"]
  RP["POST /v1/reload"] --> IR
  PR["PUT /v1/routes/{group}"] --> IR
  IR --> RW["Core::reload_with"]
  RW -- "应用" --> SUP
  RW -- "send_if_modified" --> AW["api watch"]
  AW --> FA["follow_api"]
  FA --> API["API 服务器任务"]
  M -- "信号" --> FA
  M -- "Instance::shutdown" --> SUP

整个进程生命周期内只运行一个 supervisor。三个调用方会重载它,全都经过 Instance::reload 和 Core::reload_with;后者先把变更应用到 supervisor,再发布正在运行的配置的 API 选项。API 跟随器是唯一会启动、停止或重新绑定 REST API 的代码。收到信号时,main 先停止跟随器,再停止 supervisor。下面各节按顺序介绍:启动、文件监视器、重载、一次重载保留什么、API 跟随器和关停。

flowchart TB
  A["Args::parse"] --> T["init_tracing"]
  T --> Q{"--test?"}
  Q -- "是" --> C["instance::check,打印结论,以 0 或 1 退出"]
  Q -- "否" --> S["Instance::start"]
  S -- "Err" --> F1["记录 failed to start,以 1 退出"]
  S -- "Ok" --> P{"正在运行的配置有 api 段吗?"}
  P -- "有" --> SA["start_api"]
  SA -- "Err" --> F2["记录日志,以零宽限期关停,以 1 退出"]
  SA -- "Ok" --> FA["spawn follow_api"]
  P -- "没有" --> FA
  FA --> W["spawn_watcher"]
  W --> SIG["等待 SIGINT 或 SIGTERM"]

main 运行在 #[tokio::main] 上,即多线程运行时。

app/src/main.rs
fn init_tracing(config_path: &Path) {
use tracing_subscriber::EnvFilter;
let filter = EnvFilter::try_from_default_env().unwrap_or_else(|_| {
let level = config::load(config_path)
.ok()
.and_then(|c| c.log.level)
.unwrap_or_else(|| "info".to_string());
EnvFilter::new(level)
});
tracing_subscriber::fmt().with_env_filter(filter).init();
}
  1. EnvFilter::try_from_default_env() 读取 RUST_LOG。它已设置且能解析时,以它为准。
  2. 否则,无论 RUST_LOG 是未设置,还是设置了无法解析的值(两者都是 Err),都会用 config::load 读取并解析一次配置文件(按文件原样,不做订阅合并),[log].level 通过 EnvFilter::new 成为过滤器;EnvFilter::new 接受与 RUST_LOG 相同的指令语法(例如 info,etemenanki_supervisor=debug)。
  3. 没有这个键,或者文件无法读取或无法解析时,过滤器为 info。
  4. init() 安装全局 subscriber,它把格式化后的日志写到标准输出。

这一步只运行一次,在 --test 和启动之前。之后没有任何代码替换过滤器,所以 [log].level 不会被重载:只改了 [log].level 的重载会发现 sources 变了,应用一个什么都没变的 spec(config reloaded: reused [...]),然后继续以旧级别记录日志,直到进程重启。--test 失败时的结论是一行日志,同样受这个过滤器约束;退出状态则总会设置。

日志 target 就是模块路径:main.rs 打印的行(failed to start、configuration invalid、shutting down,以及 API 和文件监视器的行)为 etemenanki_app,instance.rs 打印的行(config loaded、config reloaded 以及每一条 reload 行)为 etemenanki_app::instance。

app/src/instance.rs
pub async fn check(path: &Path) -> Result<(), LoadError> {
check_bytes(std::fs::read(path)?).await
}
pub async fn check_bytes(config: Vec<u8>) -> Result<(), LoadError> {
let (sources, parsed) = config::sources_from(config)?;
let built = build(&config::effective(parsed, &sources)?)?;
etemenanki_supervisor::check(&built.spec).await?;
Ok(())
}

--test 读取配置和它指定的订阅文件,合并,降为 spec(包括 [api]),然后运行 etemenanki_supervisor::check。后者新建一个 Actor::new(SocketOptions::default()),以它空的运行状态为基准做规划,再运行 prepare(&plan, spec, false):bind 为 false 时,每个出站、路由、handler(处理器)和用户表仍会构建,只跳过 listener::bind 和 Pending::pair。由于以空状态为基准规划,它永远不会报告中断性变更。启动会拒绝的东西它都会拒绝,只有无法绑定的端口、socket 路径或设备,以及配对(pairing)所做的检查除外;它也从不绑定 REST API。它具体检查什么,见配置页,以及规划并应用变更中讲试运行的那一节。

结果 输出 退出状态
通过 标准输出上的 Configuration OK.(println!,不是日志行) 0
拒绝 ERROR 级别的 configuration invalid: <e>,target 为 etemenanki_app 1

取自锁定版本的二进制(已去掉时间戳):

ERROR etemenanki_app: configuration invalid: route references unknown outbound nowhere
ERROR etemenanki_app: configuration invalid: [subscribe] needs a path naming the subscribe file
ERROR etemenanki_app: configuration invalid: [api] listen 0.0.0.0:9090: an API reachable from other hosts needs a secret; listen on a loopback address or a unix socket, or set one
ERROR etemenanki_app: configuration invalid: No such file or directory (os error 2)
ERROR etemenanki_app: configuration invalid: subscribe file missing-sub.toml: No such file or directory (os error 2)

无法读取的配置文件给出的是不带路径的裸 io::Error:check 直接调用 std::fs::read(path)?,不像重载那样加上 cannot read <path>: 。缺失的订阅文件则由 config::sources_from 在错误里写明文件名。

REST API 在写入编辑后的配置之前,运行的也是 check_bytes(app/src/api.rs → ConfigFile::set_route)。

Instance::start(path) 调用 Core::start(Supervisor::builder(), config::read_sources(&path), Arc::new(build))。Core::start:

  1. 如果 read_sources 失败,原样返回读取错误。与重载不同,它不带 cannot read <path>: 前缀。
  2. 运行 config::effective 和 lowering。合并或 lowering 错误以 LoadError::Config 返回。
  3. 调用 builder.start(built.spec)。它断言采样间隔不为零(否则 panic;app 用的是默认的 1 秒),并在一个新的空 actor 上执行一次应用:校验、规划、准备(绑定每个监听器)和提交,其间每个新监听器记录 inbound <tag> listening on <bind>。任何 ApplyError 都会返回,包括 Bind 错误,所以启动要么全部成功,要么什么都不运行。失败的准备阶段会丢弃它构建的东西,从而释放它已绑定的每个监听器。只有在应用成功之后,它才 spawn 采样器(以及 usage sink,app 没有),用 mpsc::channel(16) 创建命令通道,并 spawn actor.run。builder 见 supervisor 概览。
  4. 以 INFO 级别记录 config loaded: <summary>。
  5. 保存 supervisor 和 lowering,用 built.api 创建 api watch,并把 sources 记录进 last,failed: None。因此,启动之后紧接着对相同文件的重载会得到 Unchanged。

main 把任何错误以 ERROR 级别记录为 failed to start: <e>,并以状态 1 退出。捕获的输出:

ERROR etemenanki_app: failed to start: route references unknown outbound nowhere
ERROR etemenanki_app: failed to start: No such file or directory (os error 2)
ERROR etemenanki_app: failed to start: subscribe file missing-sub.toml: No such file or directory (os error 2)

无法绑定的监听器给出 failed to start: inbound <tag>: binding <bind> failed: <e>,其中 <bind> 是 BindSpec 的 Display。它有五种形式,app 的 lowering 会产生其中四种:<host>:<port>、udp <host>:<port>、unix:<path> 和 tun <name>(没有名字时为 tun auto)。第五种 tun fd <n> 是外部提供的描述符(TunSource::Fd),只有 FFI 会产生。

main 取 instance.api(),克隆当前值,值为 Some 时调用 start_api:

app/src/main.rs
async fn start_api(instance: Arc<Instance>, options: &ApiOptions) -> io::Result<Api>;
struct Api {
stop: oneshot::Sender<()>,
task: JoinHandle<io::Result<()>>,
}

start_api 用 ApiListener::bind(&options.listen) 绑定,以 INFO 级别记录 API listening on <addr>(TCP 时是 local_addr() 给出的实际绑定地址,Unix socket 时是 listen 的值),取 tracker = instance.tracker(),构建 etemenanki_webclient::router(ConfigFile::new(instance), tracker, options.secret.clone(), env!("CARGO_PKG_VERSION")),然后 spawn listener.serve(app, stopped),其中 stopped 是一个新 oneshot 的接收端。ApiListener 如何绑定 TCP 地址或 Unix socket 路径、何时删除 socket 文件,见 REST API 页。

第一次绑定失败时,main 记录 failed to start: API on <listen>: <e>,调用 instance.shutdown(Duration::ZERO),并以状态 1 退出。此时 supervisor 已经在运行,入站也已在接受连接;零宽限期会立即关闭它们已接受的一切。没有 [api] 的配置除了入站之外不绑定任何东西。

接着 main 创建 oneshot 对 (stop_api, api_stopped),spawn follow_api(instance.clone(), options, running, api_stopped)(见 API 跟随器),并调用 spawn_watcher。文件监视器无法建立时,它以 ERROR 级别记录 config hot-reload disabled: <e> 并继续运行:代理照常工作,只是没有来自文件监视器的重载。最后它 await wait_for_shutdown。

app/src/main.rs
fn watch_dir(path: &Path) -> PathBuf;
fn sync_watches(
watcher: &mut notify::RecommendedWatcher,
watched: &mut HashSet<PathBuf>,
instance: &Instance,
);
fn spawn_watcher(instance: Arc<Instance>) -> notify::Result<()>;

spawn_watcher 创建一个 notify::RecommendedWatcher,让它监视 Instance::watched_files() 返回的每个路径的 watch_dir,每个都用 RecursiveMode::NonRecursive:它监视的是存放文件的目录,而不是文件本身。监视器的回调把 () tick 送进一个通道,由一个任务把这些 tick 变成 Instance::reload 调用。是否有变化由 Core::reload_with 判断,它按路径重新读取文件并比较字节。

watch_dir(path) 是文件的父目录;路径没有父目录部分时(-c config.toml)为 .。

spawn_watcher:

  1. 创建一个元素为 () 的无界 tokio::sync::mpsc 通道。
  2. 用 notify::recommended_watcher 创建 RecommendedWatcher,其回调把 () tick 送进通道。
  3. 监视 watch_dir(instance.config_path())。这一步或第 2 步失败时返回 notify::Error,main 记录 config hot-reload disabled: <e>:配置自己所在的目录必须能被监视,“否则就完全没有热重载”。
  4. 以一个包含该目录的 HashSet 作为 watched 的初值,并调用一次 sync_watches;配置在别处指定了订阅文件时,它会把订阅文件所在的目录加进来。
  5. spawn 监视任务,把监视器、receiver 和 watched move 进去,返回 Ok(())。

RecommendedWatcher 住在任务内部。任务拥有监视器,监视器的闭包拥有 sender,所以通道永远不会关闭,任务一直运行到 main 返回、运行时关闭为止。

flowchart TB
  N["监视器回调发送一个 tick"] --> R["监视任务:rx.recv()"]
  R --> D1["取走已排队的 tick"]
  D1 --> S["短暂等待稳定"]
  S --> D2["再取走一次"]
  D2 --> L["instance.reload()"]
  L --> SW["sync_watches"]
  SW --> R
  • 每轮一次重载。 拿到一个 tick 之后,任务先取走已经排队的 tick,再做一次短暂的稳定等待(内联的 tokio::time::sleep,不是具名常量),然后再取走一次,最后为这一轮取走的所有 tick 只调用一次 instance.reload()。
  • 一次只有一个重载。 任务在取下一个 tick 之前会 await instance.reload(),所以它自己的重载从不重叠。来自 REST API 的重载与监视器的重载之间,由 Core 中的 last 锁串行化。
  • 忽略结果。 reload_with 已经记录了结果。监视器不报告 Unchanged,这种情况是静默的。
  • 每轮之后重新同步。 每次重载之后,不论结果如何,都运行 sync_watches,所以被监视的目录集合跟随磁盘上的配置文件。

sync_watches 把 instance.watched_files() 经 watch_dir 映射,算出想要监视的目录集合:

  1. 对每个想要但尚未监视的目录,调用 watcher.watch(dir, RecursiveMode::NonRecursive)。失败以 ERROR 级别记录为 cannot watch <dir> for changes: <e>,但不是致命错误。
  2. 对每个正在监视但不再需要的目录,调用 watcher.unwatch(dir) 并忽略结果。
  3. 把想要的集合存为已监视集合。

配置自己所在的目录总是需要监视的,因为 watched_files 总以配置路径开头。只要磁盘上的配置文件指定了订阅文件,订阅文件所在的目录就需要监视。由于 watched_files 读的是磁盘上的文件而不是正在运行的配置,一份开始指定新订阅文件的配置,即使应用失败,从下一轮起也会让那个文件所在的目录被监视。

app/src/instance.rs
impl Core {
pub async fn reload_with(&self, read: impl FnOnce() -> Read) -> Result<Reload, LoadError>;
}
impl Instance {
pub async fn reload(&self) -> Result<Reload, LoadError>;
}

有三个调用方会重载 Instance,它们都经过 Instance::reload:

调用方 路径 如何处理结果
监视任务 spawn_watcher 忽略
POST /v1/reload webclient/src/router.rs → reload → ConfigFile::reload instance.reload().await?.report(),以 JSON 或错误状态码应答
PUT /v1/routes/{group} webclient/src/router.rs → set 在 ConfigFile::set_route(检查并写入文件)以及其后的 ConfigFile::reload 期间一直持有 router 的 edits 锁,并以重载结果和文件的新 ETag 应答

router 的 edits 锁把一次 PUT 与它之后的重载串行化,但正如 ConfigControl 的文档注释所说,它管不到来自别处(例如文件监视器)的重载;那是 Core 的 last 锁的事。读到的字节恰好是另一次重载最近记录的字节时,重载返回 Unchanged。如果这些字节没有与失败一起记录,REST API 以 200 {"status":"unchanged"} 应答;记录了失败则应答 422。无论哪种情况,set 都会在应答上带上文件的新 ETag,因为不管其后的重载是否成功,文件都已经写入了。

编辑本身(ConfigFile::set_route:If-Match 检查、check_bytes 试运行、在检查期间有人手工编辑时以 Stale 拒绝的重新读取,以及原子写入)见 REST API 页。

sequenceDiagram
  participant C as 调用方
  participant K as Core::reload_with
  participant R as read 闭包
  participant S as Supervisor
  participant A as api watch
  C->>K: reload_with(read)
  K->>K: 锁住 last
  K->>R: read()
  R-->>K: Sources 与解析结果
  alt 读取失败
    K-->>C: Err(Config),日志为 reload 加上错误
  else sources 等于 last.sources
    K-->>C: Ok(Unchanged),带上记录的失败
  else sources 有变化
    K->>K: config::effective,然后 lowering
    alt 合并或 lowering 失败
      K-->>C: Err(Config),保留正在运行的配置
    else 构建成功
      K->>S: apply(spec)
      alt 已应用
        S-->>K: ApplyReport
        K->>A: send_if_modified(built.api)
        K->>K: last 更新为这些 sources,无失败
        K-->>C: Ok(Applied(report))
      else 被拒绝
        S-->>K: ApplyError
        K-->>C: Err(Apply),保留正在运行的配置
      end
    end
  end
  1. 加锁。 self.last.lock().await。guard 一直持有到调用结束,跨越读取和 supervisor 的应用。因此读取、比较、应用和更新 last 是一个整体步骤:两个调用方不可能读到同一个变化并应用两次,较慢的调用方也不可能在较新的文件之后应用较旧的文件。文档注释这样写道:“加载 read 给出的内容;读取时不会有其他重载在运行。”
  2. 读取。 read() 在锁内运行。对 Instance::reload 来说它是 config::read_sources(&self.path),读取错误被改写为 cannot read <path>: <e>。读取错误以 ERROR 级别记录为 reload: <e>(对 app 来说是 reload: cannot read <path>: <e>),并以 LoadError::Config 返回。
  3. 比较。 sources == last.sources 时,调用返回 Ok(Reload::Unchanged { failed: last.failed.clone() }),不记录任何日志,也不碰 supervisor。
  4. 合并并降为 spec。 config::effective(parsed, &sources).and_then(|cfg| (self.lowering)(&cfg))。TOML 错误就在这里出现,同时出现的还有订阅文件错误、未知协议、错误的 settings 表和 [api] 错误。出错时调用以 ERROR 级别记录 reload: <e>; keeping the running config,并返回 LoadError::Config。
  5. 应用。 self.supervisor.apply(built.spec),即 apply_with 加上 ApplyOptions::default():allow_disruptive 为 false。supervisor 的 actor 依次校验、规划、拒绝中断性计划、准备和提交(见规划并应用变更)。
  6. 已应用。 调用以 INFO 级别记录 config reloaded: <summary>,用 send_if_modified 发布 API 选项(见下文),把 last 设为这些 sources 且 failed: None,并返回 Reload::Applied(report)。
  7. 被拒绝。 ApplyError::Disruptive { inbound, reason } 记录为 reload refused, keeping the running config: inbound <inbound>: <reason>, which would end its live connections; restart to apply it。其他任何 ApplyError 记录为 reload: <e>; keeping the running config。两者都是 ERROR 级别,错误以 LoadError::Apply 返回。supervisor 整体拒绝一个 spec,所以正在运行的内容与之前完全相同;a_refused_spec_changes_nothing 针对校验错误和准备阶段的绑定失败固定了这一点。
app/src/instance.rs
self.api.send_if_modified(|api| {
let changed = *api != built.api;
*api = built.api;
changed
});

每次已应用的重载之后,watch 的值都会被替换,但只有新的 Option<ApiOptions> 与旧值不同时才会唤醒 receiver:listen 不同、secret 不同,或者这个段出现或消失。被拒绝或 lowering 失败的重载什么也不发布,所以“被拒绝的重载不会改变 API,正如它不会改变其他任何东西”。二进制在 follow_api 中而不是在这里对变化做出反应(见 API 跟随器)。

last 保存最近一次记录的重载的 sources,以及那次重载的失败文本(如果有)。第 3 步的比较使用 Sources 派生的 PartialEq:比较配置文件的字节,以及配置指定了订阅文件时订阅文件的字节。只看字节:

  • sources 等于 last.sources 的重载返回 Unchanged { failed },不碰 supervisor:静默,除了 read_sources 做的那一次解析之外不再解析,也不应用。
  • 只改订阅文件也算变化,只加注释或调整键顺序的编辑同样算。这样的重载和其他重载一样被降为 spec 并应用(见 summary 中捕获的第三行)。
  • last 的初值是 Core::start 加载的 sources,没有失败。

Unchanged { failed } 携带与这些 sources 一起记录的失败文本(如果有)。这段文本是错误的 Display,而不是日志行;例如 ApplyError::Disruptive 显示为 inbound <tag>: <reason>; this ends its live connections and needs allow_disruptive。文件监视器忽略 Unchanged;REST API 以 422 应答,内容是这段文本后接 (the files have not changed since this was found)。

情形 日志行(除注明外均为 ERROR) reload_with 返回 REST API 应答
文件无法读取 reload: cannot read <path>: <e> Err(LoadError::Config) 按错误的 kind 为 422 或 500
sources 与记录的相同,没有记录失败 无 Ok(Unchanged { failed: None }) 200 {"status":"unchanged"}
sources 与记录的相同,记录了失败 无 Ok(Unchanged { failed: Some(..) }) 422 <reason> (the files have not changed since this was found)
TOML、合并或 lowering 错误 reload: <e>; keeping the running config Err(LoadError::Config) InvalidData 和 InvalidInput 为 422,其他 kind 为 500
中断性变更 reload refused, keeping the running config: inbound <tag>: <reason>, which would end its live connections; restart to apply it Err(LoadError::Apply(Disruptive)) 422
监听器无法绑定 reload: inbound <tag>: binding <bind> failed: <e>; keeping the running config Err(LoadError::Apply(Bind)) 500
其他任何拒绝 reload: <e>; keeping the running config Err(LoadError::Apply(..)) 422,Stopped 为 500
已应用 INFO config reloaded: <summary> Ok(Applied(report)) 200,附带报告

取自锁定版本二进制的拒绝日志(已去掉时间戳;TOML 错误只显示第一行):

ERROR etemenanki_app::instance: reload: route references unknown outbound nowhere; keeping the running config
ERROR etemenanki_app::instance: reload: outbound x: unknown protocol "carrier-pigeon"; keeping the running config
ERROR etemenanki_app::instance: reload: duplicate outbound tag direct; keeping the running config
ERROR etemenanki_app::instance: reload: [api] listen 0.0.0.0:9: an API reachable from other hosts needs a secret; listen on a loopback address or a unix socket, or set one; keeping the running config
ERROR etemenanki_app::instance: reload: TOML parse error at line 5, column 1

第一行和第三行是 supervisor 的拒绝(ApplyError::UnknownReference、ApplyError::DuplicateTag);第二行和第四行是 lowering 错误;最后一行是解析错误,由 config::effective 在 lowering 这一步报告。

supervisor 的计划会把某些入站变更标记为中断性的(supervisor/src/topology/spec_plan/plan.rs 中的 Step::Disrupt)。Supervisor::apply 传入 ApplyOptions::default(),其 allow_disruptive 为 false,所以 Actor::apply 在准备任何东西之前就拒绝这样的 spec,并指出计划列出的第一个中断(plan.disruptions().next())。计划会标记哪些变更,见规划并应用变更。

对 etemenanki-app 来说,这意味着:

  • obfs 有变化的 Hysteria 2 监听器在重载时被拒绝,重启后生效。katana 以 allow_disruptive 应用,所以同样的变更在 katana 上能够应用(见节点管理器)。
  • 设备或设置有变化的 TUN 入站在重载时被拒绝,重启后生效。

和其他任何 ApplyError 一样,拒绝针对的是整个 spec:其中的其他内容也都不会被应用。日志行中的 <reason> 是以下文本之一:

the hysteria2 obfuscation changed; connected clients cannot follow the new key
the tun inbound's settings changed; its runtime restarts and its flows end
the tun device changed; the old device and its flows end

所以 Hysteria 2 混淆变更的日志是:

ERROR etemenanki_app::instance: reload refused, keeping the running config: inbound <tag>: the hysteria2 obfuscation changed; connected clients cannot follow the new key, which would end its live connections; restart to apply it

supervisor 层面的拒绝由 supervisor/tests/hot_swap.rs 中的 a_hysteria2_obfuscation_change_needs_allow_disruptive 固定。

一次已应用的重载就是对一份完整新 spec 的应用,supervisor 会保留 spec 没有变化的一切。从 app 的角度看:

对象 在已应用的重载之后 详见
会话 不绑定在监听器上:它们运行在 supervisor 的根 token 之下,所以停止或替换监听器不会取消它们。删除入站或修改其 tag 会关闭它的会话(CloseSessions)。删除策略作用于绑定在被删除用户的 principal(身份主体)上的会话;app 降为 spec 时使用 Policies::default(),其 UserRemovalPolicy 为 Close。 用户、principal 与会话、规划并应用变更
stream 监听器 BindSpec 未变的监听器保留它的 socket、accept 循环以及该循环的准入限制(run_stream_inbound 创建的信号量)。同一绑定上变化了的入站为新连接换上新的 handler(swapped)。新入站在准备阶段绑定。绑定变化时如何处理,见右侧页面。 监听器与服务循环
同一端口上的 Hysteria 2 监听器 保留它的 QUIC endpoint,继续在它的 UDP 端口上服务。只有 max_circuits 不变时(same_circuits),它的 circuit 预算才会延续。替换的其余部分见右侧页面。 监听器与服务循环
出站 spec 未变、且 DNS spec 也未变(blackhole 除外)的出站被复用,连同它的 connector 持有的状态。变化了的出站被构建为一个新版本,新的流使用它。 出站、UDP fan-out 与负载均衡器、规划并应用变更
DNS DnsSpec 不变时,保留解析器及其应答缓存 名称解析与 DNS 服务
负载均衡器 负载均衡器及其所有成员都未变时,保留成员健康状态和正在运行的探测 出站、UDP fan-out 与负载均衡器
路由表 RouteSpec 不变时被复用。新连接遵循新的 plane(数据平面);已打开的连接保留它们打开时的目标。 plane:为每个流选路
[api] 跟随已应用的配置:启动、停止或重新绑定 API 跟随器
[log].level 不重载;init_tracing 只运行一次 日志初始化
supervisor 设置 在进程生命周期内固定:Supervisor::builder() 的默认值,即采样间隔 DEFAULT_SAMPLE_INTERVAL(1 秒)、没有 usage sink,以及 SocketOptions::default() supervisor 概览
应用选项 每次重载都以 ApplyOptions::default() 应用(Supervisor::apply),启动用的是 builder 的默认值,两者相同 中断性变更会被拒绝

完整的表格以及每一行背后的计划步骤,见规划并应用变更中“一次应用保留什么”一节。两个 app 测试端到端地固定了会话这几行:an_open_connection_keeps_transferring_across_a_reload(路由变更把新连接送进 blackhole,而一条已打开的 SOCKS 连接继续回显)和 a_reload_that_removes_a_user_closes_only_their_connections(删除一个 SOCKS 账户会关闭该账户已打开的连接并拒绝它的下一次登录,而另一个账户的连接继续回显)。

app/src/main.rs
async fn follow_api(
instance: Arc<Instance>,
mut options: watch::Receiver<Option<ApiOptions>>,
mut running: Option<Api>,
mut stopped: oneshot::Receiver<()>,
);

follow_api 让 REST API 始终符合正在运行的配置的 [api],从 main 启动的那个 API(running)开始,直到收到 stopped 信号:

flowchart TB
  W["select:停止信号,或 options.changed()"] -- "停止信号,或 sender 已不在" --> E["停止正在运行的 API,返回"]
  W -- "changed" --> B["wanted = borrow_and_update()"]
  B --> S["停止正在运行的 API(如果有)"]
  S --> Q{"wanted 是什么?"}
  Q -- "Some" --> ST["start_api(wanted)"]
  ST -- "Ok" --> R["running = 新的 API"]
  ST -- "Err" --> L1["记录日志:在 [api] 变化之前保持关闭"]
  Q -- "None" --> L2["记录日志:API stopped"]
  R --> W
  L1 --> W
  L2 --> W
  • 先停再绑。 旧 API 在新 API 绑定之前停止,“这样同一个地址可以再次被占用”。因此只改 secret 也会关闭监听器,再重新绑定同一个端口。在这两步之间,到 API 的连接会被拒绝。
  • 只看最新值。 watch 只保存一个值。跟随器忙碌期间如果有多次已应用的重载改变了 [api],它只醒来一次,并按最新的选项行事。
  • 由 API 发起的重载先完成。 重启在这个任务里运行,“而不是在发布它的那次重载里,这样由 API 自己发起的重载能在 API 停止之前完成它的请求”。Core::reload_with 只负责发布;引发重载的请求返回它的应答,然后跟随器的 Api::stop 等待它完成。如果重启在重载内部运行,请求会等待停止,停止又会等待请求,直到宽限期耗尽。
  • 绑定失败不重试。 新监听器无法绑定时,跟随器记录 API on <listen>: <e>; it stays off until [api] changes,并让 running 保持为空。它要等发布的选项下一次变化时才再次行动。
  • sender 比循环活得久。 watch::Sender 位于 Core 中,而跟随器持有一个 Arc<Instance>,所以任务运行期间 changed() 不可能返回错误;针对这种情况的 break 只是防御性的。

Api::stop 在 stop 通道上发送,这会让传给 ApiListener::serve 的优雅关停 future 完成:axum 停止接受新连接,让进行中的请求完成。然后它最多等待 SHUTDOWN_GRACE,让服务器任务结束:

服务器任务的结果 日志行
及时以 Ok(()) 结束 无
以 Err(e) 结束 ERROR 级别的 API: <e>
panic 或被取消(JoinError) ERROR 级别的 API task failed: <e>
5 秒后仍在运行 对服务器任务调用 task.abort(),然后以 WARN 级别记录 API: requests still in flight when it stopped were dropped

服务器任务就是运行 ApiListener::serve 的那个任务,它 await axum::serve(..).with_graceful_shutdown(..)。

跟随器的另外两种日志是 API listening on <addr>(INFO,来自 start_api)和 API stopped: the config no longer has [api](INFO)。只要发布的选项变成 None,就会打印后者,不管当时是否有 API 在运行:例如在一次重新绑定失败之后,就没有 API 在运行。API 的端点和认证见 REST API 页;运维视角见用户指南的 API 页。

app/src/main.rs
async fn wait_for_shutdown();

在 Unix 上,wait_for_shutdown 用 tokio::signal::unix::signal(SignalKind::terminate()) 注册一个 SIGTERM 流,在 tokio::signal::ctrl_c()(SIGINT)和 SIGTERM 中先到的那个到来时返回。SIGTERM 注册失败时,它只等 SIGINT。在其他平台上它等待 Ctrl-C。Tokio 在进程的整个生命周期内保持它的信号处理函数处于安装状态,所以关停期间再来一个 SIGINT 或 SIGTERM 不会打断关停;SIGKILL 会,而且那时不会做任何清理。

sequenceDiagram
  participant M as main
  participant F as follow_api 任务
  participant API as API 服务器任务
  participant S as supervisor actor
  M->>M: SIGINT 或 SIGTERM,记录 shutting down
  M->>F: stop_api.send(())
  F->>API: Api::stop,然后最多等 5 秒
  API-->>F: 结束,或被 abort 并记录警告
  F-->>M: 任务返回
  M->>S: Instance::shutdown(SHUTDOWN_GRACE)
  S->>S: 停止监听器和探测,最多等 5 秒让连接结束
  S->>S: 取消根 token,等待每个任务,停止采样器
  S-->>M: 完成
  M->>M: 返回 ExitCode::SUCCESS
  1. main 以 INFO 级别记录 shutting down。
  2. stop_api.send(()) 通知跟随器,main await 它的任务,并丢弃 JoinHandle 的结果(let _ = api.await)。跟随器在一个 select! 处离开循环:如果它正处于一次重启之中,会先完成这次重启。tokio::select! 在已就绪的分支中随机挑选,所以当同时还有一次选项变化待处理时,循环在看到停止信号之前可能还会再做一次重启。随后跟随器停止正在运行的 API(如果有),给进行中的请求最多 SHUTDOWN_GRACE 的时间。
  3. instance.shutdown(SHUTDOWN_GRACE) 运行 Supervisor::shutdown(Duration::from_secs(5))。Actor::shutdown 对每个监听器调用 stop()(结束的 accept 循环会丢弃它的监听器,Unix socket 文件也随之删除),取消每个负载均衡器的探测 token,关闭任务 tracker,用 tokio::time::timeout(grace, tracker.wait()) 最多等待一个宽限期,取消根 token 以结束所有剩余连接,等待每个被跟踪的任务,清空监听器句柄,取消 background_stop,并 await 每个后台任务:采样器,以及 usage sink 的推送任务(app 没有)。完整顺序见 supervisor 概览。
  4. main 返回 ExitCode::SUCCESS。运行时关闭,并丢弃监视任务及其 notify 监视器。

API 先停止,所以在 API 宽限期内完成的请求会在 supervisor 开始关停之前得到应答。第 3 步期间,文件监视器仍可能跑完一轮:supervisor 的 actor 按顺序处理命令,所以已经发出的应用会在关停之前完成,之后发出的则以 ApplyError::Stopped 失败(reload: the supervisor has shut down; keeping the running config)。

进程如何结束 退出状态 关停流程
--test,通过 0 无;什么都没有启动
--test,拒绝 1 无
启动被拒绝 1 无;被拒绝的应用已释放它绑定的东西
第一次 API 绑定失败 1 Instance::shutdown(Duration::ZERO)
SIGINT 或 SIGTERM 0 先停止 API,再 Instance::shutdown(SHUTDOWN_GRACE)
SIGHUP 被信号杀死(在 shell 中为 129) 无
SIGKILL 被信号杀死 无

由信号触发的关停,最多用 5 秒等 API 的请求,再最多用 5 秒等存活的连接,之后再加上剩余任务察觉根 token 被取消所需的时间。

不变量 由谁保证 由谁固定
启动要么运行整份配置,要么什么都不运行 Core::start 返回 SupervisorBuilder::start 给出的任何 ApplyError,后者在 spawn actor 和采样器之前就失败;失败的准备阶段会丢弃它绑定的东西 a_refused_spec_changes_nothing(针对准备阶段);没有 app 层面的测试
同一个 Core 的重载从不重叠 last 是一个在读取和应用期间一直持有的 tokio::sync::Mutex 没有直接的测试
与记录相同的 sources 不会被再次应用 sources == last.sources 返回 Unchanged a_reload_after_the_apis_own_finds_nothing_changed(app/tests/unit/api.rs)
没有被应用的重载不改变任何正在运行的东西,包括 API supervisor 的准备和提交;只有 Ok 之后才 send_if_modified a_refused_spec_changes_nothing(supervisor/tests/hot_swap.rs)
入站和用户都没变的已打开连接,在改变路由的重载之后继续中继 会话运行在 supervisor 的根 token 之下,而不是监听器的 token 之下;新连接使用新的 plane an_open_connection_keeps_transferring_across_a_reload;an_established_connection_survives_an_apply_that_keeps_its_inbound(supervisor 层面)
同一端口上的协议变更保留 socket,替换为 HTTP 之前接受的 SOCKS 连接继续讲 SOCKS Listener::swap 把新的 handler 放进监听器的 watch,供新连接使用(swapped);accept 循环继续运行 a_protocol_change_on_the_same_port_keeps_the_socket_and_old_connections(supervisor 层面)
删除一个 SOCKS 账户会关闭该账户已打开的连接,另一个账户的连接继续运行 app 降为 spec 时使用的删除策略 UserRemovalPolicy::Close a_reload_that_removes_a_user_closes_only_their_connections
删除入站会关闭它的会话 提交阶段的 Step::CloseSessions removing_an_inbound_closes_its_sessions
同一端口上变化了的 Hysteria 2 入站继续在该端口上服务 监听器保留它的 QUIC endpoint a_reload_keeps_the_inbound_serving_on_its_udp_port
app 从不应用中断性变更 Supervisor::apply 使用 ApplyOptions::default() a_hysteria2_obfuscation_change_needs_allow_disruptive(supervisor 的拒绝);没有 app 层面的测试
API 跟随已应用的 [api]:段消失时 API 消失,段回来时 API 在对应地址回来,只改 secret 时重新绑定 send_if_modified 和 follow_api the_api_follows_the_api_section_across_reloads
没有 [api] 的配置除了入站不绑定任何东西 main 只在选项为 Some 时启动 API a_config_without_api_binds_nothing_new(仅 Linux)
由 API 发起的重载在 API 重启之前应答 重启在 follow_api 中运行,而不是在 reload_with 中 无
tracing 只根据 RUST_LOG 或 [log].level 初始化一次 main 开头的 init_tracing 无
干净的关停会释放入站拥有的资源 Supervisor::shutdown 停止每个监听器 socks_over_a_unix_socket_relays_and_cleans_up(SIGTERM 之后 socket 文件消失)
关停不会被 supervisor 自己的定时器或卡住的 Hysteria 2 握手拖住 宽限期结束后的 root.cancel() 结束宽限定时器,并关闭 Hysteria 2 endpoint shutdown_ends_pending_grace_timers、shutdown_does_not_wait_out_a_stalled_hysteria2_handshake(supervisor 层面)
失败 日志行 后果
配置或订阅文件无法读取、TOML 或 lowering 错误,或任何 ApplyError failed to start: <e> 退出状态 1;什么都不会继续运行
监听器无法绑定 failed to start: inbound <tag>: binding <bind> failed: <e> 退出状态 1;同一次准备中已绑定的监听器被释放
REST API 无法绑定 failed to start: API on <listen>: <e> Instance::shutdown(Duration::ZERO),退出状态 1
无法创建监视器,或无法监视配置所在的目录 config hot-reload disabled: <e> 进程照常运行,但没有由文件触发的重载;REST API 仍然可以重载

如结果表所示,每种失败都保留正在运行的配置。订阅文件所在目录无法监视时,记录为 cannot watch <dir> for changes: <e>,不会让其他任何东西停下来。

重新绑定失败时,API 保持关闭,并记录 API on <listen>: <e>; it stays off until [api] changes。代理不受影响。超过 SHUTDOWN_GRACE 的停止会对服务器任务调用 task.abort(),并以 WARN 级别记录 API: requests still in flight when it stopped were dropped。

  • 只有文件监视器会把 reload_with await 到完成。API 请求进行中如果客户端离开,请求的 handler 会被丢弃(hyper 在 HTTP/1 连接上看到 EOF,或者 HTTP/2 stream 被 reset),reload_with 也随之在它已到达的那个 await 点被丢弃。actor 如何对待调用方已经离开的命令,见 supervisor 概览。
  • 监视任务从不被显式取消。它在 main 返回、运行时关闭时结束。
  • 跟随器在收到 main 发来的 stopped 信号时结束。main 只在收到信号之后才发送它,并在关停 supervisor 之前 await 跟随器。
  • Api::stop 用 tokio::time::timeout(SHUTDOWN_GRACE, &mut task) 限制等待时间,超时就 abort 服务器任务。
常量 值 定义位置 含义
SHUTDOWN_GRACE 5 秒 app/src/main.rs 每次停止 API 时给 API 请求的宽限期,以及关停时给存活连接的宽限期
API 启动失败后的宽限期 Duration::ZERO app/src/main.rs → main 立即关闭入站已接受的一切
DEFAULT_SAMPLE_INTERVAL 1 秒 supervisor/src/track/sampler.rs 统计采样器的 tick 间隔;app 不修改它
supervisor 命令队列 16 supervisor/src/supervisor.rs → SupervisorBuilder::start,mpsc::channel(16) 等待 actor 处理的命令,重载的应用也在其中
DEFAULT_LISTEN 127.0.0.1:9090 webclient/src/listen.rs,由 ApiConfig 的 serde 默认值 default_api_listen 使用 [api] 没有 listen 时 API 的地址
每次拒绝报告的中断数 1 Actor::apply → plan.disruptions().next() 只指出第一个中断性入站

代理本身的进程级上限汇总在限制、超时与内存一页中。

端到端测试运行真实的二进制。app/tests/support/mod.rs → spawn_app(dir, "app.toml") 以 -c app.toml 启动 CARGO_BIN_EXE_etemenanki-app,工作目录为测试目录,所以配置路径是一个裸文件名,被监视的目录是 .;标准输出和标准错误被丢弃。返回的 Proc 在被丢弃时用 SIGKILL 杀死子进程;Proc::terminate 发送 SIGTERM 并轮询等待退出,最多 10 秒,用于那些要测的正是进程退出时行为的测试。重载测试用 std::fs::write 原地重写配置文件,然后每 100 毫秒探测一次,最多 20 秒(eventually),直到一条新连接表明新配置已生效,这样已打开的连接接下来做什么,是在重载之后观察到的。

文件 测试 固定的行为
app/tests/integration/e2e_reload.rs an_open_connection_keeps_transferring_across_a_reload 在一次加入 127.0.0.0/8 → block 规则的重载之前打开的 SOCKS 连接继续回显,而新连接被 blackhole
app/tests/integration/e2e_reload.rs a_reload_that_removes_a_user_closes_only_their_connections 从 SOCKS 账户中删除 alice,会在 10 秒内关闭她已打开的连接,并拒绝她的下一次登录;bob 已打开的连接继续回显
app/tests/integration/e2e_api.rs the_api_follows_the_api_section_across_reloads 删除 [api]:端口不再应答。在另一个端口上以 secret = "one" 加回来:带 token 为 200,不带为 401。只把 secret 改为 "two":同一个端口接受新 token、拒绝旧 token
app/tests/integration/e2e_api.rs a_config_without_api_binds_nothing_new 从 /proc 读取:没有 [api] 的配置只监听它的 SOCKS 端口;同一配置加上 [api] 后恰好多出 API 的端口。仅 Linux
app/tests/integration/e2e_api.rs a_switch_sends_new_flows_of_the_group_through_the_new_target PUT /v1/routes/Local 应答 "status":"applied",即编辑之后的重载已应用
app/tests/integration/e2e_hysteria_inbound.rs a_reload_keeps_the_inbound_serving_on_its_udp_port 一个 Hysteria 2 入站在同一端口上从 udp = false 改写为 udp = true,连着一个真实的 hysteria 客户端,重载后仍然能通过该客户端服务。它在写入后等待 6 秒,然后以 500 毫秒间隔重试 20 次。go 或上游构建不可用时跳过
app/tests/integration/e2e_unix.rs socks_over_a_unix_socket_relays_and_cleans_up Proc::terminate(SIGTERM)之后,入站的 socket 文件消失
app/tests/integration/e2e_tun.rs a_routed_connect_is_answered_while_the_app_runs 进程在 Proc::terminate 之后退出,此后路由进 TUN 设备的连接不再成功:接口及其路由都已消失。内核会在描述符关闭时删除创建的设备,所以这个测试无法区分干净的关停和其他任何退出。需要 CAP_NET_ADMIN,没有时跳过
app/tests/unit/api.rs a_reload_after_the_apis_own_finds_nothing_changed 一次 PUT 应用编辑之后,POST /v1/reload 应答 200 {"status":"unchanged"}:对编辑后那次重载所记录字节的重载不会再次应用它们
app/tests/unit/api.rs a_switch_edits_the_file_reloads_and_returns_the_new_etag 在进程内启动的 Instance 上执行 PUT 会触发重载("status":"applied"),并返回文件的新 ETag
app/tests/unit/api.rs api_options_come_from_the_api_section api::options(build 的一部分):没有 [api] 为 None;空的 [api] 在 DEFAULT_LISTEN 上监听,没有 secret;含有 / 的值是 Unix socket 路径;带 secret 的非本地 listen 被接受
app/tests/unit/api.rs an_api_open_to_other_hosts_needs_a_secret 不带 secret 的非本地 listen、回环地址上的空 secret、用主机名作为 listen,以及 [api] 中的未知键,全部被拒绝
supervisor/tests/hot_swap.rs a_refused_spec_changes_nothing 被校验(UnknownReference)或准备阶段的绑定(Bind)拒绝的 spec,让 epoch、路由、用户和监听器保持原样,并释放在那次准备中绑定的监听器
supervisor/tests/hot_swap.rs a_hysteria2_obfuscation_change_needs_allow_disruptive 没有 allow_disruptive 时,混淆变更对 hy2-in 是 ApplyError::Disruptive
supervisor/tests/hot_swap.rs an_established_connection_survives_an_apply_that_keeps_its_inbound 一次加入一个出站和一条规则的应用把入站报告为 reused,既不是 swapped 也不是 rebound,已打开的 SOCKS 连接继续回显
supervisor/tests/hot_swap.rs a_protocol_change_on_the_same_port_keeps_the_socket_and_old_connections 同一端口上从 SOCKS 改为 HTTP 是 swapped 而不是 rebound;已打开的 SOCKS 连接继续回显,替换之前接受的连接仍然得到 SOCKS handler,新连接讲 HTTP
supervisor/tests/hot_swap.rs removing_an_inbound_closes_its_sessions 删除两个 SOCKS 入站中的一个,会把它报告为 removed 并关闭它已打开的连接;另一个入站的连接继续回显
supervisor/tests/hot_swap.rs shutdown_ends_pending_grace_timers 各等待一小时的一次删除和一次排空不会拖住 shutdown(Duration::ZERO),处于删除宽限期内的会话被关闭
supervisor/tests/hot_swap.rs shutdown_does_not_wait_out_a_stalled_hysteria2_handshake 卡在 QUIC 握手中的 Hysteria 2 客户端不会拖住服务器的关停
ffi/tests/proxy.rs set_route_and_reload_switch_like_the_rest_api 由 FFI 用自己的 Lowering 启动(而不是由 Instance 启动)的 Core,通过 reload_with 应用一次路由切换(见 FFI 页)

没有测试覆盖以下内容:文件监视器跟随新指定的订阅文件、reload 日志行、带有已记录失败的 Reload::report、app 层面的中断性拒绝、API 启动或重新绑定失败,以及 SIGHUP。修改这些路径时应该补上测试;各测试套件的组织方式见测试。