Node manager
Source files: 19 · checked against katana v3.0.1
katana/src/manager/node.rskatana/src/manager/mod.rskatana/src/manager/transport.rskatana/src/manager/proxy.rskatana/src/main.rskatana/src/runtime.rskatana/src/serve.rskatana/src/api/mod.rskatana/src/api/newv2board.rskatana/src/api/sspanel.rskatana/src/config.rskatana/src/traffic.rskatana/src/rule.rskatana/src/router.rskatana/tests/unit/e2e.rskatana/tests/unit/runtime.rskatana/tests/unit/traffic.rskatana/tests/unit/rule.rskatana/tests/integration/sniff.rs
Every [[node]] entry in a katana config becomes one NodeManager running in one tokio task. The manager is the node’s control loop: it asks the panel (or, for a locally described Hysteria 2 node, the config file) what the node should look like, keeps a listener that matches, and sends the panel the node’s traffic and audit hits. It also receives the config-file updates the process root fans out on a hot reload.
This page is for contributors who change src/manager/. It covers the manager’s state and locks, the run loop from the bootstrap retries to exit, the reconcile ladder that decides how much of a node to rebuild, the static-update path, the report paths and the small helpers in src/manager/mod.rs. The listener subtree it owns (TransportManager and ProxyManager) is introduced here only as far as the manager drives it; serving, admission and traffic accounting have their own pages.
Responsibilities
Section titled “Responsibilities”NodeManager owns:
- Desired state. The last node description and user list it applied (
Cur), which every later change is diffed against. - The listener subtree. At most one
TransportManagerat a time, built bybring_upand dropped bytear_down. - The node’s shared registries. One
NodeTraffic(per-user counters) and oneRuleManager(audit rules and hits). Both outlive every listener rebuild, which is what keeps byte counts continuous across a rebuild. - The compiled router. The node’s
[node.route]compiled against the global outbound pool, recompiled when either changes. - Reporting. One traffic report and one illegal-access report per poll cycle, and a final flush on exit.
It deliberately does not own:
- the panel HTTP protocol (the
PanelClientit holds, see the panel clients page); it only swaps in a new client when a config edit changes the settings the client was built from; - node identity and respawn on a reload (
src/runtime.rsdecides that before it sends an update); - per-connection serving, handshake deadlines and admission (
src/serve.rs,src/connector.rs).
flowchart TB runtime["runtime::spawn_built"] nm["NodeManager (one task per node)"] api["PanelClient"] tm["TransportManager (one listener generation)"] pm["ProxyManager (user tables, pre-auth permits)"] traffic["NodeTraffic (Arc, shared)"] rules["RuleManager (Arc, shared)"] runtime -->|"spawns run(), holds static_tx"| nm nm -->|"node_info, user_list, node_rule, reports"| api nm -->|"bring_up / tear_down"| tm tm --> pm nm --- traffic nm --- rules pm --- traffic
Key types
Section titled “Key types”NodeManager
Section titled “NodeManager”pub struct NodeManager { api: Mutex<Arc<PanelClient>>, cfg: Mutex<NodeConfig>, pool: Mutex<Arc<HashMap<CompactString, Arc<Outbound>>>>, router: Mutex<Arc<Router<Outbound>>>, traffic: Arc<NodeTraffic>, rules: Arc<RuleManager>, transport: tokio::sync::Mutex<Option<TransportManager>>, cur: tokio::sync::Mutex<Cur>,}
impl NodeManager { pub fn new( api: PanelClient, cfg: NodeConfig, pool: Arc<HashMap<CompactString, Arc<Outbound>>>, ) -> io::Result<Arc<Self>>;
fn api(&self) -> Arc<PanelClient>;
pub async fn run( self: Arc<Self>, shutdown: CancellationToken, mut static_rx: mpsc::Receiver<StaticUpdate>, );}new compiles the initial router with build_router(&cfg.route, &pool, &CompactString::from("direct")), so any router build error fails the constructor: a route or default that names an unknown outbound tag, an invalid cidr or port entry, or geodata that cannot be loaded. At start-up, runtime::spawn_node logs that as node <id>: build router: <error> and skips the node; the process still starts if at least one other node spawned. On a reload, the runtime builds every node the reload adds, router included, before it touches any running node, and refuses the whole reload if one does not build.
Mutex here is parking_lot::Mutex. The fields fall into two lock families:
| Field | Lock | Written by | Read by |
|---|---|---|---|
api |
parking_lot::Mutex |
new, apply_static (StaticUpdate::Config) |
node_info, try_bootstrap, poll_cycle, refresh_rules, reports, all through api() |
cfg |
parking_lot::Mutex |
apply_static (StaticUpdate::Config) |
almost every method, for one field at a time |
pool |
parking_lot::Mutex |
apply_static (StaticUpdate::Outbounds) |
rebuild_router |
router |
parking_lot::Mutex |
apply_static (StaticUpdate::Config), rebuild_router |
bring_up (cloned into the new Dispatcher) |
traffic |
none (Arc, interior locks) |
bring_up, reconcile, ProxyManager::refresh |
report_traffic |
rules |
none (Arc, interior locks) |
refresh_rules |
report_illegal, every flow’s audit check |
transport |
tokio::sync::Mutex |
bring_up, tear_down |
reconcile (listener check, user refresh) |
cur |
tokio::sync::Mutex |
set_cur |
poll_cycle, reconcile, apply_static, refresh_rules, report_illegal |
The poll loop and the static updates run in the same task, one after the other, so none of these locks is ever contended. They exist only to carry mutable state across &self methods on the shared Arc<NodeManager>. Two rules keep that true if you add code:
- A
parking_lotguard is never held across an.await. Every use is a short expression such asself.cfg.lock().controller.listen_ip.clone(), or a block that copies the fields it needs and drops the guard.api()clones theArc<PanelClient>out of its lock, so a panel request runs without the guard, and a request already in flight whenapply_staticswaps the client finishes on the old one. The compiler backs this up:tokio::spawnrequires therunfuture to beSend, and aparking_lotguard is notSendunlesslock_api’ssend_guardfeature is on, which nothing in katana’s dependency graph enables. - The two tokio mutexes are never held at the same time. Each method takes one, copies what it needs out, and releases it before taking the other.
#[derive(Default)]struct Cur { node: Option<NodeInfo>, users: Vec<UserInfo>, node_tag: CompactString,}Cur is the last applied desired state, not the last fetched one. node is None only before bootstrap finishes; apply_static reads that to tell whether a node is up yet. node_tag is the audit and report tag of that state, the key under which refresh_rules stores rules and report_illegal drains hits.
StaticUpdate
Section titled “StaticUpdate”pub enum StaticUpdate { Config(Box<NodeConfig>), Outbounds(Arc<HashMap<CompactString, Arc<Outbound>>>),}The runtime sends these over a bounded mpsc channel of capacity 16, created when it spawns the node (runtime::spawn_built). Config carries this node’s new [[node]] block and is sent only when the block changed and the node’s identity did not. Identity is the tuple (panel_type lowercased, api.host, api.node_id, api.key, panel node type); a change to any of those makes the runtime cancel the node and spawn a new one instead. The panel node type is the node_type a newV2board request names (vless for a V2ray-family node, api.node_type V2ray, Vmess or Vless, with enable_vless, otherwise the lowercased api.node_type), because UniProxy finds a node by its id and that type. It is empty for SSPanel, which finds a node by id alone. A Config may therefore change [node.api] fields the panel client was built from, such as timeout, speed_limit or rule_list_path; the node rebuilds its client for those (see StaticUpdate::Config). Outbounds carries the rebuilt global pool and is sent to every node the runtime holds a handle for whenever the [[outbound]] list changed.
The subtree it drives
Section titled “The subtree it drives”impl TransportManager { pub async fn start( node: &NodeInfo, cert: &CertConfig, listen_ip: &str, enable_vless: bool, sniff: bool, users: &[UserInfo], traffic: Arc<NodeTraffic>, rules: Arc<RuleManager>, router: Arc<Router<Outbound>>, node_tag: CompactString, hysteria: &HysteriaConfig, ) -> io::Result<Self>;
pub fn proxy(&self) -> &Arc<ProxyManager>;
pub async fn shutdown(self);}impl ProxyManager { pub fn refresh(&self, node: &NodeInfo, users: &[UserInfo], enable_vless: bool); pub fn retire_all(&self);}TransportManager::start builds everything that can fail (transport, protocol tables or Hysteria authenticator and server) and binds the socket before it commits the staged user set, so a bad description never half-binds. ProxyManager::refresh replaces only the user tables of a live listener. The manager calls nothing else on the subtree.
Lifecycle of run
Section titled “Lifecycle of run”if self.bootstrap(&shutdown, &mut static_rx).await { self.serve(&shutdown, &mut static_rx).await;}self.tear_down().await;self.report_traffic().await;self.report_illegal().await;run has three phases. bootstrap brings the node up, retrying until it is up or the token is cancelled, and returns whether it came up. serve is the steady-state loop and runs only after a successful bootstrap. The exit steps run in both cases, including for a node that was cancelled before it ever came up.
sequenceDiagram
participant RT as runtime
participant NM as NodeManager::run
participant P as PanelClient
participant TM as TransportManager
RT->>NM: tokio::spawn(nm.run(shutdown, static_rx))
loop bootstrap, until up or shutdown
NM->>P: forget_etags(), node_info(), user_list()
NM->>TM: bring_up(node, users)
NM->>NM: set_cur(Some(node), users, tag)
NM->>P: node_rule() unless disable_get_rule
Note over NM: on failure, wait and apply static updates
end
loop serve, until shutdown
alt interval tick
NM->>NM: poll_cycle(false)
else static update
NM->>NM: apply_static(update), maybe reset interval
end
end
NM->>TM: tear_down()
NM->>P: report_user_traffic(), report_illegal()
Bootstrap
Section titled “Bootstrap”bootstrap calls try_bootstrap until one attempt succeeds. One attempt runs these steps and stops at the first that fails:
- Fresh reads.
api().forget_etags()clears the client’s ETag cache. Nothing a failed attempt read was applied, and a panel answers a request it has answered before with 304 Not Modified, which would leave the attempt with nothing to start from. - Node description.
node_inforeturns the panel’s answer, except for a Hysteria 2 node whose[node.hysteria].portis non-zero (next section). - Port check. A description with
port == 0is refused. - Users.
api().user_list()must return a list. An empty list is accepted. - Tag.
node_tag(&node, &listen_ip). - Listener.
bring_up(&node, &users). With an empty user list it commits an empty registry and binds nothing, so the node comes up dark and the first poll that sees users builds the listener. - State.
set_cur(Some(node), users, tag). - Rules.
refresh_rules(), unlesscontroller.disable_get_ruleis set.
An attempt fails with one of these reasons:
| Condition | Reason |
|---|---|
node_info returns Ok(None) |
panel returned no node info |
node_info returns Err |
node_info failed: <error> |
| the description has port 0 | panel returned port 0 |
user_list returns Ok(None) |
panel returned no user list |
user_list returns Err |
user_list failed: <error> |
bring_up fails (build or bind) |
initial start failed: <error> |
The newV2board client rejects a server_port of 0 itself, with the error newV2board: server port must be > 0, so on that panel a zero port fails the attempt as node_info failed. The port-0 row is reached when the client passes a zero port through, for example SSPanel with an offset_port_node of "0".
After a failure, bootstrap logs node <id>: <reason>; retrying in <n>s at error and waits. The wait starts at BOOTSTRAP_RETRY_MIN (1 s) and doubles after each failure up to BOOTSTRAP_RETRY_MAX (60 s), and it is never longer than the poll period (poll_period()), so a node that syncs every few seconds also retries every few seconds. With the default update_periodic of 60 s the waits are 1, 2, 4, 8, 16, 32 and then 60 seconds.
Each wait is one tokio::time::sleep, created once, so updates that arrive during the wait cannot push the retry back. While it waits, the node applies static updates as they arrive, which also keeps the runtime’s bounded channel from filling:
- An
Outboundsupdate stores the pool and recompiles the router. There is no listener to rebuild yet. - A
Configupdate is applied or refused (seeStaticUpdate::Config) and ends the wait at once, because the edit may be the fix. With the node not up, it only stores the new config, client and router, and the next attempt reads with the new client and builds from the new config.
Both the attempt and the wait sit in a biased select! whose first branch is the cancellation token, so a node that is cancelled while it is still coming up stops at once, dropping any panel request in flight.
While a node is in this loop it serves nothing on its port. The process and the other nodes are not affected.
The local Hysteria 2 description
Section titled “The local Hysteria 2 description”node_info reads api.node_type, [node.hysteria] and api.speed_limit from the current cfg on every call. When the node type parses as NodeType::Hysteria2 and hysteria.port != 0, it builds the NodeInfo itself and never asks the panel:
NodeInfo field |
Value |
|---|---|
node_type |
NodeType::Hysteria2 |
port |
hysteria.port |
speed_limit |
api.speed_limit in Mbps, times 1_000_000.0 / 8.0, cast to u64 (the cast saturates, so a negative value becomes 0, unlimited) |
transport, enable_tls |
Transport::Tcp, true: a QUIC listener has no stream transport and its TLS is part of the handshake |
obfs_type, obfs_password |
hysteria.obfs, hysteria.obfs_password, empty when unset |
| every other field | empty or false |
With hysteria.port == 0 a Hysteria 2 node asks the panel like any other node type. The settings a panel cannot express (credential source, UDP relay, masquerade) come from [node.hysteria] in both cases, read by bring_up.
Because this path cannot fail, a locally described Hysteria 2 node never falls back to the last applied description in poll_cycle.
The poll interval
Section titled “The poll interval”fn poll_period(&self) -> Duration { Duration::from_secs(self.cfg.lock().controller.update_periodic.max(1))}update_periodic defaults to 60 seconds (default_update_periodic in src/config.rs) and is clamped to at least 1 second. serve builds a tokio::time::interval from it and immediately awaits the first tick, which an interval completes at once. The first poll_cycle therefore runs one full period after bootstrap, not straight after it. The interval uses tokio’s default MissedTickBehavior::Burst: when a cycle takes longer than the period, for example because several panel requests each ran into their timeout, the missed ticks fire back to back.
The select loop
Section titled “The select loop”loop { tokio::select! { _ = shutdown.cancelled() => break, _ = interval.tick() => self.poll_cycle(false).await, Some(update) = static_rx.recv() => { self.apply_static(update).await; let new_period = self.poll_period(); if new_period != period { period = new_period; interval = tokio::time::interval(period); interval.tick().await; } } }}- Polls and static updates are strictly serialized. A static update that arrives during a poll waits for the poll to finish, and the other way round.
- After each static update the loop recomputes the period. When
update_periodicchanged, it replaces the interval and again consumes the immediate tick, so the next poll comes one new period after the update. - If every sender is dropped,
recv()yieldsNone, theSome(update)pattern fails andselect!disables that branch for that iteration. The loop keeps polling. - Cancellation is observed only between iterations. A cancel that arrives during
poll_cycleorapply_static(which may itself run a poll cycle) takes effect when that call returns. When several branches are ready at once,tokio::select!picks one at random, so a static update already queued when the token is cancelled may or may not be applied before the loop exits. The panel requests inside a cycle are bounded by the HTTP client timeout,api.timeoutseconds or 5 when it is 0 (ApiConfig::timeout_secs).
The listener goes down first, so no relay is still adding bytes when the counters are read, and the final report carries everything up to the moment of shutdown instead of dropping up to one poll period of traffic. The flush is a single attempt: if the panel rejects it, the failure is logged and the bytes go with the task. For a node cancelled during bootstrap there is usually nothing to tear down or report; the exit steps still run so that a listener generation an attempt had already bound is shut down in order (see rebuild, bring_up and tear_down). The runtime awaits each node’s task after cancelling it, both on process shutdown and when a reload removes a node, so the flush completes before the process exits.
The poll cycle
Section titled “The poll cycle”async fn poll_cycle(&self, rebuild: bool)The timer calls it with rebuild = false. apply_static calls it once, straight after a config edit, with rebuild = true when the edit changes something the listener is built from (see StaticUpdate::Config). Each call runs these steps in order, whatever the earlier ones returned:
| Step | Call | On Ok(None) |
On Err |
|---|---|---|---|
| 1 | node_info() |
reuse cur.node |
log node <id>: node_info: <error> at warn, reuse cur.node |
| 2 | api().user_list() |
reuse cur.users |
log node <id>: user_list: <error> at warn, reuse cur.users |
| 3 | reconcile(node, users, rebuild) |
||
| 4 | refresh_rules(), unless disable_get_rule |
keep the rules | log at warn, keep the rules |
| 5 | report_traffic() |
restore residuals, log at warn |
|
| 6 | report_illegal() |
log at warn |
Ok(None) is the panel clients’ “not modified” answer (HTTP 304 against a cached ETag). Treating it and a failed request the same way, as “nothing new”, is what lets a flaky panel leave a running node alone: the fallback description is exactly the one already applied, so reconcile finds nothing to change.
A description whose port is 0 is also treated as a miss, because a listener could not serve it. The cycle logs node <id>: refreshed port is 0, keeping the last one at error, keeps cur.node, and still reconciles the user list the panel returned, so user changes keep applying while the panel’s description is unusable. As at bootstrap, a newV2board panel that returns server_port 0 does not reach this check: its client returns Err, so the cycle takes the step 1 fallback instead.
The fallback reads cur.node with .expect("node_info present after bootstrap"). It holds because serve runs only after bootstrap returned true, which happens only after try_bootstrap ran set_cur(Some(node), …), and every later set_cur also passes Some. apply_static runs a poll cycle only when cur.node is already Some.
Reconcile
Section titled “Reconcile”async fn reconcile(&self, node: NodeInfo, users: Vec<UserInfo>, rebuild: bool)reconcile takes the first branch that applies, in the order that drops the fewest connections for the change at hand. rebuild forces the full-rebuild branch, for a local edit that the panel’s answer cannot show:
stateDiagram-v2 [*] --> UsersCheck UsersCheck --> TearDown: users is empty UsersCheck --> ListenerCheck: users present ListenerCheck --> FullRebuild: rebuild forced, or no transport ListenerCheck --> FullRebuild: transport_eq or protocol_eq is false ListenerCheck --> UserCheck: same listener and protocol UserCheck --> Refresh: user_set_differs or speed_limit changed UserCheck --> Unchanged: nothing differs TearDown --> SetCur FullRebuild --> SetCur Refresh --> SetCur Unchanged --> SetCur SetCur --> [*]
| Branch | Trigger | Action | Effect on connections |
|---|---|---|---|
| User-less | users.is_empty() |
tear_down(), then traffic.commit(traffic.prepare(Vec::new())) |
every connection drops; every counter moves to draining |
| Full rebuild | rebuild is true, or no TransportManager, or !node.transport_eq(prev), or !node.protocol_eq(prev) |
rebuild(&node, &users) |
every connection drops |
| In-place refresh | user_set_differs(&cur.users, &users), or node.speed_limit differs from the previous node’s |
transport.proxy().refresh(&node, &users, enable_vless) |
unchanged users keep their connections; departed and rebound users are retired |
| No change | none of the above | nothing | none |
Every branch ends with set_cur(Some(node), users, tag), where tag is computed from the new description and the current controller.listen_ip.
A few details matter when you change this code:
- “No transport” is the retry path. When a rebuild fails,
transportstaysNone. The next cycle takes the full-rebuild branch again because of!has_transport, which is how a node that went dark after a failed rebuild, or after a user-less period, comes back without any explicit retry loop. enable_vlesscomes from the config. The in-place refresh passescfg.api.enable_vless, notnode.enable_vless, toProxyManager::refresh, matching whatbring_uppasses toTransportManager::start.speed_limitis a user-bucket change. Neither equality below includes it, because the node limit only feeds each user’s rate (determine_rate). A change creates fresh counters for the users whose effective rate changed; their live connections stay open, and flows they open afterwards run at the new rate.- The refresh is transactional.
ProxyManager::refreshbuilds the replacement tables before publishing anything. If that build fails it logsproxy refresh build failed, keeping current: <error>and leaves both the tables and the registry untouched. When it succeeds it commits the registry first and swaps the tables second, so a newly added user is never authenticated by the new table while the registry still refuses their flows.
What forces a full rebuild
Section titled “What forces a full rebuild”NodeInfo::transport_eq and NodeInfo::protocol_eq in src/api/mod.rs split the description into the two layers a rebuild cares about:
NodeInfo field |
Compared by | A change means |
|---|---|---|
port, transport, host, path, service_name, authority, enable_tls, header, headers, enable_reality, accept_proxy_protocol, obfs_type, obfs_password |
transport_eq |
full rebuild |
node_type, enable_vless, vless_flow, cypher_method, server_key |
protocol_eq |
full rebuild |
speed_limit |
neither | in-place refresh |
The obfuscation fields are in transport_eq because they are part of the wire: without the rebuild the listener would keep the old key and lock out every client that took the new one.
rebuild, bring_up and tear_down
Section titled “rebuild, bring_up and tear_down”async fn rebuild(&self, node: &NodeInfo, users: &[UserInfo]);async fn bring_up(&self, node: &NodeInfo, users: &[UserInfo]) -> io::Result<()>;async fn tear_down(&self);rebuildistear_downfollowed bybring_up. Abring_uperror is logged asnode <id>: rebuild failed: <error>and the node stays dark until the next cycle retries.bring_upwith no users commits an empty registry and returnsOk(())without binding. Otherwise it copiescontroller.cert,controller.listen_ip,api.enable_vless,!controller.disable_sniffingand[node.hysteria]out ofcfg, clones the current router, takes thetransportlock, callsTransportManager::start, stores the result and logsnode <id>: listening on <listen_ip>:<port>. Because the lock is taken beforestart, no.awaitlies between a finishedstartand the store. A bootstrap attempt that shutdown cuts short after its bind has therefore already stored its generation, andrun’stear_downshuts it down in order, including the wait for a QUIC endpoint to release its UDP socket.tear_downtakes theTransportManagerout of its slot and awaitsTransportManager::shutdown, which cancels the accept scope and waits for every task under it, retires every user lease, and for a Hysteria 2 node awaits the QUIC endpoint until it has released its UDP socket. The last step is what letsrebuildbind the same UDP port immediately afterwards.
TransportManager also implements Drop as a backstop: it cancels the scope and retires every user if a generation is ever dropped without shutdown. It cannot wait for the tasks to finish, so every path in the manager uses shutdown.
The bind happens in a fresh TransportManager after the old one is gone, so a full rebuild has a short window in which the port is closed. A rebuild whose bind fails (for example because another process took the port in that window) leaves the node dark until a later cycle succeeds.
Listener state across cycles
Section titled “Listener state across cycles”stateDiagram-v2 [*] --> Bootstrap Bootstrap --> Bootstrap: attempt failed, wait and retry Bootstrap --> Dark: no users Bootstrap --> Serving: bring_up succeeded Bootstrap --> Stopped: shutdown Serving --> Serving: in-place refresh or no change Serving --> Serving: full rebuild succeeded Serving --> Dark: users empty, or rebuild failed Dark --> Serving: next cycle rebuilds with users Serving --> Stopped: shutdown Dark --> Stopped: shutdown Stopped --> [*]
Dark means the task is alive but holds no TransportManager. Bootstrap loops until the node is up or stopped; the task never ends while its token is live.
Static updates
Section titled “Static updates”async fn apply_static(&self, u: StaticUpdate)StaticUpdate::Config
Section titled “StaticUpdate::Config”A config edit is taken whole or not at all. apply_static first builds what the edit needs, then stores it, then applies it with one poll cycle:
flowchart TB
start["StaticUpdate::Config(new)"]
client{"panel_type or api changed?"}
newclient["PanelClient::new(new)"]
route{"route changed?"}
router["build_router(new route, pool)"]
refused["refused: keep the running config"]
store["store cfg, router, client"]
resync{"bootstrapped, and rebuild or new client?"}
poll["poll_cycle(rebuild)"]
stored["stored only, read later"]
start --> client
client -->|yes| newclient
client -->|no| route
newclient -->|Err| refused
newclient -->|Ok| route
route -->|yes| router
route -->|no| store
router -->|Err| refused
router -->|Ok| store
store --> resync
resync -->|yes| poll
resync -->|no| stored
- Client. When
panel_typeor any[node.api]field differs,PanelClient::newbuilds a client from the new block. The client copies settings at construction (the node type, VLESS, the speed override, the rule file, the timeout), so without a new one the node would keep answering to the old values. The new client callsinherit_routeson the old one: a newV2board client takes over the routes the old one last read, so the audit rules derived from theirblockentries hold until the new client reads the node config itself. - Router. When
routediffers,build_routercompiles the new[node.route]against the current pool with"direct"as the fallback tag. - Refusal. If either build fails, the manager logs
node <id>: config edit refused, keeping the running one: <error>aterrorand returns. Nothing from the edit is stored, including fields that needed no build. - Store. The new
cfgreplaces the old one. A new router replacesrouterand logsnode <id>: router rebuilt; a new client replacesapi. - Apply. The edit needs a listener rebuild when any of these changed:
route,controller.listen_ip,controller.cert,controller.disable_sniffing,api.enable_vlessor anything in[node.hysteria]. The listener is built from these local settings, and the panel’s answer does not change when they do. If the node is bootstrapped (cur.nodeisSome) and the edit needs a rebuild or brought a new client,apply_staticruns onepoll_cycle(rebuild)at once.
That poll is a full cycle. A new client holds no ETags, so the panel answers in full and the node and users are read as the edited settings read them. reconcile then takes the full-rebuild branch when rebuild is set, and otherwise the narrowest branch the fresh answer needs. The poll also refreshes the rules and sends the traffic and illegal-access reports, as a timer poll does. It does not reset the poll interval.
For a node that is still in bootstrap, nothing is polled: the edit is stored, or refused as in step 3 if it does not build. Either way the wait ends at once and the next attempt reads with the new client and builds from the new config.
| Changed field | What happens | Takes effect |
|---|---|---|
route (the whole [node.route] table) |
new router, then poll_cycle(true): full rebuild |
now |
controller.listen_ip, controller.cert, controller.disable_sniffing |
poll_cycle(true): full rebuild |
now |
any [node.hysteria] field |
poll_cycle(true): full rebuild |
now |
api.enable_vless |
new client, then poll_cycle(true): full rebuild. On newV2board with a V2ray or Vmess node this changes the identity instead, and the runtime respawns the node |
now |
api.speed_limit, api.rule_list_path |
new client, then poll_cycle(false): an in-place refresh, unless the fresh answer changes the transport or protocol; the same poll refreshes the rules |
now |
api.timeout and the other [node.api] fields outside the identity (vless_flow, device_limit, disable_custom_config, node_type on SSPanel, a case-only node_type change on newV2board), and a case-only panel_type change |
new client, then poll_cycle(false): the reconcile ladder decides from the fresh answer |
now |
controller.update_periodic |
stored; the select loop replaces the interval | the next poll, one new period later |
controller.disable_get_rule, controller.disable_upload_traffic |
stored | the next poll cycle |
A routing edit rebuilds the whole transport rather than cancelling authenticated connections, because the requirement is that a routing edit drops every connection on the node, including ones still in their handshake. Cancelling user leases alone would miss those. The rebuild keeps byte continuity: NodeTraffic persists, so prepare reuses the counters of users whose key, uid and rate did not change.
StaticUpdate::Outbounds
Section titled “StaticUpdate::Outbounds”The new pool replaces pool, rebuild_router recompiles the router, and apply_route_change rebuilds the transport from cur.node and cur.users. Every node receives this update when the pool changed, so an [[outbound]] edit drops every connection on every node that is up. A node still in bootstrap has no cur.node yet, so it only keeps the new router for its next attempt.
runtime::apply_reload sends Outbounds to every running node before it removes, reconfigures or adds any node. Two consequences follow:
- A node that the same reload removes can still rebuild once against the new pool. The runtime cancels its token right after queuing the update, and when both are ready the
serveloop picks one at random, so whether that rebuild happens depends on timing. - A reload that changes both
[[outbound]]and a node’s[[node]]block (in a way that triggers a rebuild) rebuilds that node twice in a row: once forOutbounds, then once forConfigthrough itspoll_cycle(true). The select loop handles them one after the other, possibly with a timer poll in between.
rebuild_router
Section titled “rebuild_router”fn rebuild_router(&self)Only the Outbounds path calls it. It compiles the current cfg.route against the new pool with "direct" as the fallback tag. On success it replaces router and logs node <id>: router rebuilt. On failure it logs node <id>: route rebuild failed, keeping current: <error> and keeps the old router. apply_route_change still runs, so the fresh listener binds with the router that was current before the pool changed.
Most router errors are caught before this point. The runtime builds the new pool before it sends Outbounds and rejects the whole reload if that fails, so a bad [[outbound]] never reaches a node. It also builds every node the reload adds, router included, before it touches a running node. apply_static refuses a route edit that does not compile. What is left for rebuild_router is the unchanged route of a node the reload keeps: a route or default that names a tag the new pool no longer has (for example after an [[outbound]] was removed), or geodata that can no longer be loaded.
Rules and reports
Section titled “Rules and reports”refresh_rules
Section titled “refresh_rules”async fn refresh_rules(&self)api().node_rule() returning Ok(Some(rules)) replaces the rules stored under cur.node_tag through RuleManager::update, which skips the write when the new list has the same ids and pattern strings in the same order. Ok(None) keeps the current rules, and Err is logged as node <id>: node_rule: <error> at warn. Only the SSPanel client makes a request here and can answer Ok(None) or Err. The newV2board client always returns Ok(Some): it builds the list from api.rule_list_path and the block routes cached by its last node_info that returned a body (or inherited from the client it replaced), without contacting the panel.
report_traffic
Section titled “report_traffic”async fn report_traffic(&self)With controller.disable_upload_traffic set, it calls traffic.clear_residuals() and traffic.prune_draining() and returns. Nothing will ever be sent, so neither set may grow for the life of the process.
Otherwise:
traffic.snapshot()returns one row per live counter, one per draining counter, and one per parked residual. A draining counter with no writer left (Arc::strong_count == 1) is emitted as a counterless row and removed from the draining set.- Rows with
up == 0 && down == 0are skipped. - The rest are summed per
uidwith saturating adds. A residual or draining row can coexist with a live counter for the same user, and the newV2board payload is keyed by uid, so two rows would let one overwrite the other. - Rows with a counter go to a commit list, rows without one to a residual list.
- If no uid has bytes, nothing is sent.
- One
api().report_user_traffic(&reports)call carries every uid. - On success, every counter in the commit list gets
commit_reported(up, down), which subtracts exactly the reported amounts so bytes added during the request stay for the next report. - On failure,
traffic.restore_residuals(residuals)parks the counterless rows again, and the log readsnode <id>: report traffic: <error>. Live counters were not touched, so they still hold their bytes.
This is what makes a failed report lose nothing and a retry bill nothing twice. The registry side (prepare, commit, draining and residuals) is described on the traffic accounting page.
report_illegal
Section titled “report_illegal”async fn report_illegal(&self)It drains the hits recorded under cur.node_tag (RuleManager::drain) and, when there are any, sends them with api().report_illegal. The hits are drained before the request, so a failed report is logged as node <id>: report illegal: <error> and those hits are not sent again. The newV2board client’s report_illegal is a no-op, and the SSPanel client leaves out hits of local rules (rule id -1, from rule_list_path) and sends nothing when no panel rule was hit.
Helpers in src/manager/mod.rs
Section titled “Helpers in src/manager/mod.rs”pub(crate) fn user_key(by_email: bool, u: &UserInfo) -> Option<AuthKey>;pub(crate) fn user_tag(by_email: bool, u: &UserInfo) -> Option<Arc<UserTag>>;pub(crate) fn node_tag(node: &NodeInfo, listen_ip: &str) -> CompactString;pub(crate) fn build_user_entries( node: &NodeInfo, users: &[UserInfo],) -> (Vec<UserEntry>, Vec<UserInfo>);pub(crate) fn user_set_differs(a: &[UserInfo], b: &[UserInfo]) -> bool;| Helper | Behaviour |
|---|---|
user_key |
With by_email, AuthKey::Name(traffic_email(u)): the panel email, or the uid as a string when the email is empty. Otherwise AuthKey::Uuid of u.uuid, or None when it does not parse. by_email comes from NodeType::keys_by_email, which is false for V2ray and true for Trojan, Shadowsocks and Hysteria2. |
user_tag |
user_key wrapped as Arc<UserTag> with the user’s uid: the payload a protocol’s user table carries for each credential. |
node_tag |
Type_listenip_port, where Type is V2ray, Trojan, Shadowsocks or Hysteria2. For example V2ray_0.0.0.0_443. Used as the audit rule key, the dispatcher tag and the handshake-failure log tag. |
build_user_entries |
For every user with a key, one UserEntry { key, uid, rate } with rate = determine_rate(node.speed_limit, u.speed_limit), plus the list of those users. A user without a key, which can only happen on a V2ray node whose uuid does not parse, is skipped with skipping user <uid>: uuid is not a valid UUID at warn; the log names the uid and never the credential. |
user_set_differs |
Order-independent set comparison on the whole UserInfo (it derives Eq and Hash). Different lengths differ; otherwise it checks that every entry of b is in a HashSet of a. Any field change, a password or speed limit included, counts as a difference. |
determine_rate in src/traffic.rs returns the smaller of the two limits, where 0 means unlimited on that side: (0, 0) is 0, (0, u) is u, (n, 0) is n, (n, u) is n.min(u).
Both TransportManager::start and ProxyManager::refresh go through build_user_entries, and both pass the same by_email to user_tag. That is the single place the registry key and the protocol table’s key are derived from, so the two cannot disagree about how a node type keys its users.
Invariants
Section titled “Invariants”| Invariant | Enforced by | Pinned by |
|---|---|---|
| A user-set change keeps unchanged users’ live connections and drops departed users’ connections. | reconcile routes user changes to ProxyManager::refresh; NodeTraffic::prepare keeps counters with the same key, uid and rate and lists departed keys for retirement. |
unchanged_user_survives_user_refresh in tests/unit/e2e.rs |
| A Hysteria 2 user refresh never rebinds the UDP socket. | Tables::Hysteria replaces only the authenticator (Hy2Inbound::set_authenticator); admission refuses a retired user’s new flows. |
a_retired_user_stops_while_the_rest_keep_their_connections, repeated_user_refreshes_never_disturb_a_live_connection in tests/unit/e2e.rs |
| A routing edit drops every connection on the node. | apply_static builds the new router and runs poll_cycle(true), whose reconcile takes the full-rebuild branch. |
route_change_drops_connections in tests/unit/e2e.rs |
| A node that cannot come up keeps trying and comes up once the cause clears. | bootstrap retries try_bootstrap with a capped, doubling wait; each attempt calls forget_etags so a 304 cannot leave it without a node or users. |
a_node_comes_up_once_the_panel_answers, a_node_whose_port_is_taken_comes_up_once_it_is_free in tests/unit/e2e.rs |
| A node stops promptly while it is still retrying its bootstrap. | bootstrap watches the token first in a biased select!, during each attempt and each wait. |
a_node_that_never_bootstraps_still_stops in tests/unit/e2e.rs |
| A config edit to a field the panel client was built from takes effect in the running node. | apply_static builds a new PanelClient and runs poll_cycle, which reads without ETags and reconciles. |
an_sspanel_api_edit_takes_effect_in_place, a_client_edit_takes_effect_without_dropping_connections in tests/unit/runtime.rs |
| A reload with a node that does not build leaves the running node untouched. | runtime::apply_reload builds added nodes and checks every changed node’s panel client before it applies anything; apply_static refuses an edit whose client or router does not build. |
a_reload_with_a_node_that_does_not_build_changes_nothing in tests/unit/runtime.rs |
| Every relayed byte is attributed to the authenticated uid and reported. | bring_up shares one NodeTraffic with the dispatcher; report_traffic sums per uid. |
vmess_traffic_is_metered_and_reported, proxy_outbound_relays_and_meters, a_hysteria_node_relays_and_meters in tests/unit/e2e.rs |
| A panel-described Hysteria 2 node takes its port and obfuscation from the panel. | node_info asks the panel when hysteria.port == 0; obfs_type and obfs_password are part of transport_eq. |
a_panel_described_hysteria_node_serves_obfuscated_traffic in tests/unit/e2e.rs |
| A credential the panel never issued does not proxy. | The authenticator is built from build_user_entries’ valid users only. |
a_hysteria_node_refuses_an_unknown_credential in tests/unit/e2e.rs |
| A failed report loses no bytes and a retry bills none twice. | commit_reported runs only on success and subtracts the reported amounts; residual rows are restored on failure. |
restored_residuals_are_retried, commit_reported_preserves_concurrent, rate_change_drains_old_counter_and_reports_once, a_departed_users_late_bytes_are_still_reported, rebound_credential_reports_the_old_uid_separately in tests/unit/traffic.rs |
| The effective user rate is the smaller non-zero limit. | determine_rate |
determine_rate_min_nonzero in tests/unit/traffic.rs |
| An identical rule list does not replace the stored one; drained hits are cleared. | RuleManager::update compares ids and pattern strings; RuleManager::drain empties the set. |
update_skips_identical_ruleset, detect_records_and_drains in tests/unit/rule.rs |
disable_sniffing reaches the listener. |
bring_up passes !controller.disable_sniffing to TransportManager::start. |
a_sniffed_host_reaches_a_domain_rule, disable_sniffing_stops_the_domain_rule_matching in tests/integration/sniff.rs |
cur.node is Some whenever poll_cycle runs. |
serve runs only after bootstrap returned true, which requires try_bootstrap’s set_cur(Some(node), …); apply_static polls only when cur.node is Some; every other set_cur passes Some. |
no test; poll_cycle would panic the node task otherwise |
| No listener generation is leaked. | Every path goes through tear_down, which awaits TransportManager::shutdown; Drop cancels as a backstop. |
no dedicated test |
Failure paths and cancellation
Section titled “Failure paths and cancellation”| Failure | Where | Result |
|---|---|---|
| Router build fails at spawn (unknown outbound tag, bad matcher, geodata) | NodeManager::new |
at start-up the node is not spawned, and the process exits only if no node spawned at all; on a reload that adds the node, the whole reload is refused |
Bootstrap node_info or user_list error or Ok(None), port 0, or bring_up error |
try_bootstrap |
attempt fails and is logged; retried after 1 s, doubling to 60 s and capped at the poll period; a Config edit retries at once |
| Bootstrap user list is empty | try_bootstrap |
node comes up dark; the first poll with users builds the listener |
Poll node_info or user_list error |
poll_cycle |
last applied value reused; nothing changes |
| Poll description with port 0 | poll_cycle |
last applied description kept; the panel’s user list is still reconciled |
| Config edit whose client or router does not build | apply_static |
edit refused whole; the node keeps running on its current config |
| Full rebuild fails | rebuild |
node dark; retried every cycle through !has_transport |
| In-place refresh build fails | ProxyManager::refresh |
tables and registry untouched |
Router rebuild after an Outbounds update fails |
rebuild_router |
old router kept; the transport is still rebuilt |
| Traffic report fails | report_traffic |
residuals restored, counters untouched; retried next cycle |
| Illegal report fails | report_illegal |
hits dropped |
| Final flush fails | run exit |
bytes lost with the task |
Cancellation comes from the node’s CancellationToken, a child of the process root token. The runtime cancels it on SIGINT or SIGTERM, and when a reload removes the node or changes its identity. Bootstrap watches the token with a biased select!, both around each attempt and during each wait, so a node cancelled before it came up stops at once and drops any panel request in flight. Inside the serve loop, cancellation takes effect between iterations. The runtime then awaits the task, which covers the listener teardown and both final reports.
Limits
Section titled “Limits”| Name | Value | Where |
|---|---|---|
update_periodic default |
60 s | default_update_periodic in src/config.rs |
| Minimum poll period | 1 s | NodeManager::poll_period (update_periodic.max(1)) |
| Static update channel | 16 updates | mpsc::channel(16) in runtime::spawn_built; the runtime’s send().await waits when it is full; the node drains it in both bootstrap and serve |
| First bootstrap retry wait | 1 s | BOOTSTRAP_RETRY_MIN in src/manager/node.rs |
| Longest bootstrap retry wait | 60 s, or the poll period if shorter | BOOTSTRAP_RETRY_MAX in src/manager/node.rs; NodeManager::bootstrap caps each wait at poll_period() |
| Panel request timeout | api.timeout s, 5 s when 0 |
ApiConfig::timeout_secs in src/config.rs |
| Router fallback tag | "direct" |
NodeManager::new, apply_static, rebuild_router |
| Pre-auth streams per listener generation (stream nodes) | 512 | MAX_PREAUTH_STREAMS_PER_NODE in src/manager/proxy.rs |
| Handshake-failure warning threshold | more than 10 per second | HANDSHAKE_FAILURE_ALERT_PER_SEC in src/manager/proxy.rs |
The last two belong to the ProxyManager a bring_up creates, so a full rebuild resets them; an in-place refresh does not.
The manager has no unit tests of its own. It is exercised end to end by tests/unit/e2e.rs, which src/main.rs compiles into the binary’s test harness as the e2e module, and by the reload tests in tests/unit/runtime.rs, which src/runtime.rs compiles as its tests module.
Each e2e test starts a hand-rolled fake newV2board panel on a loopback port that serves /UniProxy/config and /UniProxy/user and records every /UniProxy/push body. fake_panel, fake_panel_dynamic and fake_panel_with_config are thin wrappers over fake_panel_quirky, which takes a Quirks value and also returns how many times the node config was requested:
Quirks field |
Effect |
|---|---|
config_failures |
answer that many /UniProxy/config requests with 500 Internal Server Error before the first real answer, like a panel that is not up yet |
etags |
tag every answer with an ETag, and answer a request that sends the tag back with 304 Not Modified |
fake_sspanel serves an SSPanel mod_mu panel with one V2ray node, described by the legacy server string, and one user. test_node_cfg sets update_periodic = 1, and the e2e spawn_node helper builds a real newV2board PanelClient and NodeManager from the config it is given, spawns run, and returns the static-update sender so the channel stays open. The runtime tests instead hold a node the way runtime::run does and pass edited configs through apply_reload.
| Test | What it proves about the manager |
|---|---|
vmess_traffic_is_metered_and_reported |
bootstrap binds a VMess node; report_traffic delivers the relayed bytes under uid 1001 |
unchanged_user_survives_user_refresh |
adding a user keeps user A’s connection; removing A drops it; A’s bytes from both phases are reported |
proxy_outbound_relays_and_meters |
the router built in new reaches a proxy outbound and metering still attributes the bytes |
route_change_drops_connections |
StaticUpdate::Config with a different route drops a live connection |
a_hysteria_node_relays_and_meters |
the local Hysteria 2 description (hysteria.port != 0) binds, relays and reports |
a_hysteria_node_refuses_an_unknown_credential |
the Hysteria 2 authenticator holds only panel users |
a_retired_user_stops_while_the_rest_keep_their_connections |
a Hysteria 2 refresh retires user B while A’s QUIC connection keeps working |
a_panel_described_hysteria_node_serves_obfuscated_traffic |
with hysteria.port == 0 the port and Salamander key come from the panel |
repeated_user_refreshes_never_disturb_a_live_connection |
several consecutive refreshes leave one QUIC connection intact |
a_node_comes_up_once_the_panel_answers |
with the first two config requests failing, bootstrap retries and the node binds and relays after the third |
a_node_whose_port_is_taken_comes_up_once_it_is_free |
a failed bind is retried, and against a panel that answers 304 via ETags each attempt still reads the node and users in full |
a_node_that_never_bootstraps_still_stops |
a node retrying against a panel that always fails stops within 2 seconds of cancellation |
The reload tests in tests/unit/runtime.rs that reach the manager:
| Test | What it proves about the manager |
|---|---|
an_sspanel_node_is_its_panel_node_id, a_newv2board_node_is_also_the_type_it_asks_for |
which [[node]] edits change the identity (a respawn) and which reach the node as a Config update |
a_newv2board_type_edit_respawns_the_node |
turning enable_vless off on a newV2board V2ray node respawns it, and the new node serves VMess |
a_reload_with_a_node_that_does_not_build_changes_nothing |
an unknown node_type together with a route edit leaves the running node and its live connection untouched |
an_sspanel_api_edit_takes_effect_in_place |
on SSPanel, turning enable_vless off reaches the running node, which rebuilds its client and listener and serves VMess |
a_client_edit_takes_effect_without_dropping_connections |
a rule_list_path edit gets a new client whose rules refuse new flows to the target at once, while an open connection keeps relaying |
Run them from the katana checkout:
cargo test --locked e2e::cargo test --locked runtime::tests::No test covers the port-0 checks, the Outbounds update, the per-uid merge in report_traffic or the exit flush in isolation. If you change any of them, add a test in tests/unit/e2e.rs. The port-0 checks in try_bootstrap and poll_cycle need a description whose port is 0 after parsing. The newV2board client checks server_port for 0 as an i64 and only then casts it to u16, so through the fake panel only a value that truncates to 0, such as 65536, reaches them.