TUN
Source files: 22 · checked against Etemenanki 596916d
Etemenanki/protocols/src/tun/mod.rsEtemenanki/protocols/src/tun/config.rsEtemenanki/protocols/src/tun/device.rsEtemenanki/protocols/src/tun/inbound.rsEtemenanki/protocols/src/tun/tracked.rsEtemenanki/protocols/src/tun/udp.rsEtemenanki/protocols/src/core/mod.rsEtemenanki/protocols/src/sniff/mod.rsEtemenanki/protocols/src/lib.rsEtemenanki/protocols/Cargo.tomlEtemenanki/app/src/inbound/tun.rsEtemenanki/app/src/inbound/mod.rsEtemenanki/app/src/transport.rsEtemenanki/app/src/config.rsEtemenanki/app/src/instance.rsEtemenanki/app/src/connector.rsEtemenanki/app/src/main.rsEtemenanki/protocols/tests/pipeline/tun.rsEtemenanki/protocols/tests/unit/tun/udp.rsEtemenanki/protocols/tests/unit/core/mod.rsEtemenanki/app/tests/unit/inbound.rsEtemenanki/app/tests/integration/e2e_tun.rs
The TUN inbound turns a layer-3 network interface into a source of proxied flows. It creates the interface, assigns its addresses, installs the routes the operator listed, and hands the descriptor to a userspace IP stack (the ipstack crate). The stack terminates every TCP connection and UDP flow the host routes into the interface; each TCP connection then runs as its own ProxyServerRuntime over a PassthroughCore, and all UDP flows from one client source (address and port) run as one runtime over a TunUdpCore.
This page is for contributors who change protocols/src/tun/ or its glue in app/src/inbound/tun.rs. It assumes the core and runtime model from Server cores and Server runtime. For the operator’s view (settings, examples, routing the host into the device) see the TUN user guide.
Responsibilities
Section titled “Responsibilities”Why TUN is not a transport under a core
Section titled “Why TUN is not a transport under a core”The stream inbounds pair a byte-stream transport (TCP, TLS, WebSocket, gRPC) with a protocol core that decodes one client’s stream. A TUN device does not fit that shape: it is one file descriptor that carries every host the OS routes into it, one IP packet at a time. No single stream in it could stand for a connection.
So, like the Hysteria 2 listener, TunInbound takes the whole device for the lifetime of the generation. The userspace stack demultiplexes packets into TCP streams and UDP flows, and the inbound spawns one runtime per TCP stream and one per UDP client source. There is no protocol header and no credential: the IP stack already knows each flow’s destination, so the core has nothing to parse.
The inbound owns the interface end to end. open creates it, assigns its addresses and installs routes; nothing removes them explicitly. The kernel destroys the interface, and the device-scoped routes with it, when the last descriptor closes, which is when the generation ends.
What the inbound drops
Section titled “What the inbound drops”| Traffic | What happens | Where |
|---|---|---|
| ICMP and every other non-TCP/UDP protocol | Discarded; the stack has nowhere to send it. Logged at trace as tun: dropping a packet of an unsupported protocol. |
protocols/src/tun/inbound.rs → TunInbound::run, the UnknownTransport / UnknownNetwork arm |
UDP, when udp = false |
The IpStackUdpStream is dropped at once; its Drop retires the stack’s session. |
TunInbound::run |
| A UDP reply from a peer the client never addressed, or for a flow the link has already retired | Dropped: ipstack can only originate a flow from a packet it has seen, so the reply has no flow to ride. Logged at debug as tun: dropping a reply from <addr>: the client never addressed it. |
protocols/src/tun/udp.rs → TunUdpLink::poll_send_to |
| A UDP reply whose source is a domain | Dropped: an IP packet cannot carry a domain as its source. Logged at debug. |
TunUdpCore::handle, Event::Datagram |
A new TCP connection, or a UDP flow that needs a new association, while max_flows permits are exhausted |
Dropped. Logged at debug as tun: dropping a flow; the flow limit is reached. |
TunInbound::run |
| A new UDP flow while its association’s queue is full | Dropped; the client retransmits. | TunInbound::run, TrySendError::Full |
Because replies can only come back on a flow the client opened, the module docs describe the client’s view as an address-restricted cone: one association (and, with a single outbound, one outbound socket) serves every peer the client talks to, but nothing reaches the client from a peer it has not addressed. TunUdpLink keys flows by the far end’s full SocketAddr, so in practice the filter is stricter than that name: a reply is delivered only from the exact address and port the client has sent to, and a reply from a different port on the same host is dropped.
Keep the outbound path off the device
Section titled “Keep the outbound path off the device”Every flow is dialled by an outbound on the host’s real uplink. If the operator routes the default route (or the outbound’s server) into the TUN device, those dials are steered back into the device and the proxy relays its own traffic into itself. The inbound does not detect this loop; the module docs and the user guide tell operators to route only what should be proxied.
Where the code is compiled
Section titled “Where the code is compiled”The module is gated twice:
protocols/Cargo.toml→ featuretun = ["dep:ipstack", "dep:tun-rs", "dep:rtnetlink"], off by default so the classic protocols do not inherit an interface-management stack.rtnetlinkis a Linux-only target dependency.protocols/src/lib.rs→#[cfg(all(feature = "tun", unix))] pub mod tun;.
etemenanki-app enables the feature. katana does not, so it has no TUN inbound.
Key types
Section titled “Key types”Configuration
Section titled “Configuration”protocols/src/tun/config.rs holds what the inbound needs after the app has validated it:
pub const DEFAULT_MTU: u16 = 1500;pub const DEFAULT_UDP_IDLE_TIMEOUT: Duration = Duration::from_secs(60);pub const DEFAULT_MAX_FLOWS: usize = 65_536;
pub struct TunConfig<T> { pub user: Arc<T>, pub mtu: u16, pub udp: bool, pub udp_idle_timeout: Duration, pub max_flows: usize,}user is what every flow is attributed to, since there is no credential on the wire. The app passes Arc::new(()), and each Flow carries a NetworkUser whose authorization is an empty UserAuthorization::UsernamePassword.
The app parses [inbound.settings] into app/src/config.rs → TunInboundSettings (deny_unknown_fields) and splits it between the device and the inbound in app/src/inbound/tun.rs → build_tun_inbound:
TunInboundSettings field |
Goes to | Default and validation |
|---|---|---|
name: Option<String> |
DeviceSpec::name |
None: the kernel picks a name. |
mtu: Option<u16> |
DeviceSpec::mtu and TunConfig::mtu |
DEFAULT_MTU (1500). Below 1280 fails with tun mtu must be at least 1280. |
address: Vec<String> |
DeviceSpec::addresses |
Parsed with cidr::IpInet. A bare address is a host address (/32 or /128). A parse failure reports bad tun address. |
routes: Vec<String> |
DeviceSpec::routes |
Parsed with cidr::IpCidr, so host bits must be zero (10.77.1.5/24 fails with bad tun route … host part of address was not zero). |
udp: Option<bool> |
TunConfig::udp |
true. |
udp_idle_timeout: Option<u64> |
TunConfig::udp_idle_timeout |
Seconds; DEFAULT_UDP_IDLE_TIMEOUT (60 s). |
max_flows: Option<usize> |
TunConfig::max_flows |
DEFAULT_MAX_FLOWS (65 536). |
Each build error, in this table and below, starts with inbound <tag>: . build_tun_inbound also refuses a listen or port on the inbound (tun owns a network interface and has no listener; remove listen/port) and, through reject_stream (checked first), any [inbound.stream] whose network is set to anything but empty or tcp, or whose security is set to anything but empty or none (protocol tun does not support stream network … / … stream security …). The inbound-level sniffing key (default true) decides between TunInbound::new and .without_sniffing().
The device
Section titled “The device”protocols/src/tun/device.rs:
pub struct DeviceSpec { pub name: Option<String>, pub mtu: u16, pub addresses: Vec<(IpAddr, u8)>, pub routes: Vec<(IpAddr, u8)>,}
pub fn check_platform(spec: &DeviceSpec) -> io::Result<()>;pub async fn open(spec: &DeviceSpec) -> io::Result<(OwnedFd, String)>;
#[cfg(target_os = "linux")]async fn install_routes(routes: &[(IpAddr, u8)], ifindex: u32) -> io::Result<()>;
pub(crate) struct TunDevice(Arc<AsyncFd<File>>);impl TunDevice { pub(crate) fn new(fd: OwnedFd) -> io::Result<Self>; pub(crate) fn handle(&self) -> Weak<AsyncFd<File>>;}-
check_platformrefuses a spec with routes on anything but Linux (io::ErrorKind::Unsupported,tun routes are installed only on Linux; add them with the OS route tool). Bothbuild_tun_inboundandopencall it, so--testand a real start reject the same configs. -
openbuilds the interface withtun_rs::DeviceBuilder:- sets the MTU and, when given, the name;
- gives the builder the first IPv4 address (
ipv4(addr, prefix, None)) and every IPv6 address (ipv6_tuple); - calls
build_sync, then adds every further IPv4 address live withadd_address_v4, since the builder takes only one; - switches the descriptor to non-blocking and reads back the name and interface index;
- on Linux, installs the routes;
- converts the device into an
OwnedFdand returns it with the name the kernel settled on.
Any failure drops the half-built device, which destroys it, so a failed
openleaves no interface behind. -
install_routesis the netlink equivalent ofip route add <net>/<prefix> dev <ifindex>. It opens anrtnetlinkconnection, spawns its driver task, adds each route withRouteMessageBuilder::<IpAddr>(destination_prefix,output_interface(ifindex),scope(RouteScope::Link)), then drops the handle and aborts the driver. An error names the route:route <net>/<prefix>: <netlink error>. Routes are never removed explicitly; they are scoped to the device and vanish with it. -
TunDeviceis the async adapter the stack reads through: anAsyncFd<File>where onereadis one packet and onewriteis one packet.std::io::ReadandWriteare implemented for&File, so it needs nounsafeorlibc. OnWouldBlockit clears readiness (try_io) and polls again.poll_flushandpoll_shutdownare no-ops.handle()hands out aWeak; once the stack’s task has dropped the device,upgradefails, which is howshutdownlearns the descriptor is closed.
The inbound
Section titled “The inbound”protocols/src/tun/inbound.rs:
pub struct TunInbound<T> { config: Arc<TunConfig<T>>, sniff: bool, device: Arc<Mutex<Option<Weak<AsyncFd<File>>>>>, live_tcp: Arc<AtomicUsize>,}
impl<T> TunInbound<T> { pub fn new(config: TunConfig<T>) -> Self; pub fn without_sniffing(mut self) -> Self; pub async fn shutdown(&self);}
impl<T: Send + Sync + 'static> TunInbound<T> { pub async fn run<C, F>( &self, fd: OwnedFd, make_connector: F, token: CancellationToken, ) -> io::Result<()> where F: Fn(IpAddr) -> C + Send + Sync + 'static, C: Connector<Flow<T>> + Send + 'static, C::Future: Send, C::Stream: Send, C::Datagram: DatagramLink<Addr = Destination> + Send;}TunInbound is Clone by hand (no T: Clone bound); clones share device and live_tcp, so a generation has exactly one of each. device is a parking_lot::Mutex filled by run when the descriptor arrives; live_tcp counts TCP streams that have not finished dropping.
make_connector is called once per TCP connection and once per UDP association, with the client’s IP. The app’s closure builds an AppConnector whose FlowContext carries the inbound tag and source: Some(ip), so routing rules on the source address work (see Outbounds).
run owns all per-flow work in a JoinSet, so dropping its future takes every live flow with it.
TCP: TrackedTcp, PassthroughCore and serve_stream
Section titled “TCP: TrackedTcp, PassthroughCore and serve_stream”protocols/src/tun/tracked.rs:
pub struct TrackedTcp { stream: Option<IpStackTcpStream>, live: Arc<AtomicUsize>,}
impl TrackedTcp { pub fn new(stream: IpStackTcpStream, live: Arc<AtomicUsize>) -> Self; pub fn stream(&self) -> Option<&IpStackTcpStream>;}TrackedTcp wraps an ipstack TCP stream and counts it in the inbound’s live_tcp. new increments the counter; Drop first drops the inner stream (self.stream.take()) and only then decrements it. The order matters: when its protocol task is still running, IpStackTcpStream’s own Drop blocks in tokio::task::block_in_place plus Handle::block_on until that task finishes, and the count must cover that blocking part. After the stream is gone, every AsyncRead/AsyncWrite method returns NotConnected (tun: tcp stream already dropped).
protocols/src/core/mod.rs → PassthroughCore is the core every TUN TCP connection runs:
pub struct PassthroughCore<T> { pending: Option<Flow<T>>, sniff: bool, prefix: SniffPrefix, relay: Passthrough<Single>, timing: Timing,}
impl<T> PassthroughCore<T> { pub const BUF_SIZE: usize = 8 * 1024; pub fn new(flow: Flow<T>) -> Self; pub fn sniffing(flow: Flow<T>) -> Self; pub fn is_established(&self) -> bool;}Its destination is known before the first byte, so it parses nothing: it opens the flow on the client’s first bytes (or on the client’s EOF, if it closes without sending) and relays verbatim in both directions (STAGING_RESERVE = 0). sniffing turns sniffing on only when worth_sniffing says the destination is a bare IP, which on a TUN device is always true. See Sniffing for the collector.
async fn serve_stream<T, C, P>( tcp: TrackedTcp, core: PassthroughCore<T>, connector: C,) -> io::Result<()>where T: Send + Sync + 'static, C: Connector<Flow<T>, Datagram = P>, P: DatagramLink<Addr = Destination>;serve_stream builds ProxyServerRuntime::<{ PassthroughCore::<()>::BUF_SIZE }, _, _, _>::new(tcp, core, connector).showing_progress() and polls it as a Stream of Result<Traffic, RuntimeError<io::Error>> steps. Until runtime.core().is_established(), each step is wrapped in tokio::time::timeout(HANDSHAKE_TIMEOUT, runtime.next()); after that, steps are polled without an outer timeout and the core’s own relay idle deadline applies. The outer timeout exists because a runtime delivers no event before the client’s first bytes, so a silent client would otherwise never arm the core’s deadline.
UDP: TunUdpLink and TunUdpCore
Section titled “UDP: TunUdpLink and TunUdpCore”protocols/src/tun/udp.rs:
pub const FLOW_QUEUE: usize = 16;
pub struct TunUdpLink { new_flows: mpsc::Receiver<IpStackUdpStream>, flows: HashMap<SocketAddr, IpStackUdpStream>, order: Vec<SocketAddr>, next: usize,}
impl TunUdpLink { pub fn new(first: IpStackUdpStream, new_flows: mpsc::Receiver<IpStackUdpStream>) -> Self;}
impl DatagramLink for TunUdpLink { type Addr = SocketAddr; // poll_send_to, poll_recv_from}ipstack opens one IpStackUdpStream per (source, destination) pair. TunUdpLink aggregates all streams of one client source into one DatagramLink addressed by the far end:
- Adopting.
newadopts the first flow. Every read first drainsnew_flows(adopt_pending) and adopts each queued stream under itspeer_addr()(inipstacknaming,peer_addris the packet’s destination andlocal_addrits source). If the channel is closed and no flow is left, the link reportsNotConnected(tun: every flow of this association is gone). - Reading.
poll_recv_frompolls the flows round-robin, starting atnext, and returns the first datagram with the far end it came from.nextthen moves past that flow, so one busy far end cannot starve the others. A flow whose read completes without data (itsipstackidle timeout fired, or its session is gone) is retired and the scan starts over. A zero-length datagram from the client looks the same, so it also retires its flow. Whenorderbecomes empty, the link reportsNotConnected, which ends the association. - Writing.
poll_send_tolooks up the flow named byto. No flow means the client never addressed that peer: the reply is dropped and reported as sent. A write error retires that flow and also reports success, so one dead far end never fails the association. One write is one packet:ipstacktruncates a reply to the MTU minus the IP and UDP headers instead of sending a second packet.
pub struct TunUdpCore<T> { user: Arc<T>, source: IpAddr, opened: bool, timing: Timing,}
impl<T> TunUdpCore<T> { pub const BUF_SIZE: usize = 8 * 1024; pub fn new(user: Arc<T>, source: IpAddr) -> Self; pub fn is_established(&self) -> bool;}
impl<T: Send + Sync + 'static> ProxyCoreDecode for TunUdpCore<T> { type Key = Single; type Target = Flow<T>; type Error = io::Error; type TransportAddr = SocketAddr; const STAGING_RESERVE: usize = 4096; const MAX_DATAGRAM: usize = 4096; // handle}TunUdpCore turns the link into one association with a single outbound key (Single):
| Event | Core action |
|---|---|
TransportDatagram (first one) |
Effect::Open with a Flow toward that far end (DialNetwork::Udp, Remote::IpAddr), the configured user and source; enter Phase::Relay. |
TransportDatagram |
Refresh the idle deadline; if the payload is not empty, Effect::SendTo toward the far end the packet was addressed to. |
Datagram (a reply) |
Refresh the idle deadline. An IP source is staged with fx.put_to(SocketAddr, data) so the link writes it to the flow of its origin; a domain source is dropped. |
ConnectFailed, OutboundError |
Effect::Finish, logged at debug as tun: association ended: …. |
TransportEof |
Effect::Finish. |
Deadline |
Idle: Timing pushes Effect::Finish. |
SendFailed, TransportSendFailed |
The packet is dropped; the association continues. |
With AppConnector, the datagram outbound is a FanOutLink, which routes each packet by its own destination, so one association can reach several outbounds.
App glue
Section titled “App glue”app/src/inbound/tun.rs:
pub fn build_tun_inbound(cfg: &InboundConfig) -> io::Result<(InboundKind, BindSpec)>;
pub async fn run_tun_inbound( inbound: TunInbound<()>, tag: CompactString, fd: OwnedFd, router: Arc<Router>, token: CancellationToken,);build_tun_inbound returns InboundKind::Tun(TunInbound<()>) and BindSpec::Tun(DeviceSpec). In app/src/instance.rs, bind_inbound calls etemenanki_protocols::tun::open for BindSpec::Tun, logs inbound <tag> owns tun device <name>, and returns Listener::Tun(OwnedFd). spawn_generation then spawns run_tun_inbound, which runs the inbound, logs tun inbound failed: <error> at error if run returns an error, and always awaits inbound.shutdown() before the task ends. See Serving and Generations and reload for the surrounding lifecycle.
Data flow
Section titled “Data flow”The accept loop
Section titled “The accept loop”TunInbound::run wraps the descriptor in a TunDevice, records its Weak handle, configures IpStackConfig (mtu, udp_timeout(udp_idle_timeout), and packet_information only on macOS and iOS, where the device prefixes each packet with a 4-byte header), and starts IpStack::new(config, device). It then selects over three branches: the cancellation token, stack.accept(), and flows.join_next().
flowchart TB
A["stack.accept()"] --> K{"IpStackStream"}
K -->|"Tcp"| P1{"permit?"}
P1 -->|"no"| D1["drop the flow"]
P1 -->|"yes"| T["TrackedTcp, PassthroughCore, spawn serve_stream"]
K -->|"Udp and udp = false"| D2["drop the stream"]
K -->|"Udp"| Q{"association for this source?"}
Q -->|"yes, queue has room"| S["try_send into its FLOW_QUEUE"]
Q -->|"yes, queue full"| D3["drop the flow"]
Q -->|"none, or its channel closed"| P2{"permit?"}
P2 -->|"no"| D4["drop the flow"]
P2 -->|"yes"| U["TunUdpLink, TunUdpCore, spawn runtime"]
K -->|"UnknownTransport or UnknownNetwork"| D5["drop the packet"]
associations: HashMap<SocketAddr, mpsc::Sender<IpStackUdpStream>> maps each client source (IP and port) to the queue of its association. A UDP association task returns Some(src); when the join_next branch sees it and the sender for that source is closed, the entry is removed. A try_send that finds the channel closed removes the entry too, and the flow opens a fresh association.
A TCP connection
Section titled “A TCP connection”sequenceDiagram participant C as Client host participant S as ipstack participant L as TunInbound run loop participant R as serve_stream participant O as Connector C->>S: SYN (routed into the device) S->>L: IpStackStream::Tcp (on the SYN) S-->>C: SYN-ACK, handshake completes in the stack L->>L: take a permit, wrap in TrackedTcp L->>R: spawn with PassthroughCore C->>S: first payload bytes S->>R: Event::Transport R->>R: sniff up to SNIFF_LIMIT or SNIFF_TIMEOUT R->>O: Effect::Open with the sniffed domain O-->>R: outbound stream R->>O: held prefix, then relay both ways
ipstack hands the stream to the accept loop on the SYN and completes the client’s TCP handshake itself, before anything is dialled. The outbound is dialled only once the client has sent bytes (or closed its side). Two consequences follow:
- A successful
connecton the client does not mean the destination is reachable. If the dial fails,PassthroughCorehandlesConnectFailedthroughPassthrough::on_outbound_gone, which shuts the transport down and finishes. - A protocol whose server speaks first (SMTP, MySQL) stalls through this inbound: the client waits for a greeting, the runtime waits for the client, and
serve_streamgives up afterHANDSHAKE_TIMEOUTwithtun: the client never spoke.
The phases of the core’s single deadline (Timing) on a TUN connection:
stateDiagram-v2 [*] --> Handshake: nothing armed Handshake --> Sniff: first bytes, sniffing on Handshake --> Relay: first bytes with sniffing off, or client EOF Sniff --> Relay: verdict, SNIFF_LIMIT, SNIFF_TIMEOUT or EOF, then Open Relay --> Closing: RELAY_IDLE_TIMEOUT passes Relay --> [*]: both sides closed, or the outbound failed Closing --> [*]
is_established() is true in Relay and Closing, which is when serve_stream drops its outer HANDSHAKE_TIMEOUT.
A UDP association
Section titled “A UDP association”sequenceDiagram participant C as Client source A participant L as TunInbound run loop participant K as TunUdpLink participant U as TunUdpCore participant O as Datagram outbound C->>L: flow A to X (new source) L->>K: new link, FLOW_QUEUE channel K->>U: TransportDatagram from X U->>O: Open, then SendTo X C->>L: flow A to Y (same source) L->>K: try_send into the queue K->>U: TransportDatagram from Y U->>O: SendTo Y O-->>U: Datagram from Y U->>K: put_to Y K-->>C: packet from Y to A O-->>U: Datagram from Z U->>K: put_to Z K->>K: no flow to Z, dropped
udp_idle_timeout is ipstack’s per-stream UDP timeout: the stack re-arms it whenever the stream is written or polled for reading, and a stream whose timer fires returns an error on its next read, which retires it from the link. Because TunUdpLink polls flows in turn on every read, activity elsewhere in the association can re-arm a quiet flow’s timer, so udp_idle_timeout is not a precise per-far-end limit. When the last flow retires and no new one is queued, poll_recv_from returns NotConnected, the runtime ends, and the association’s permit is released. The core’s RELAY_IDLE_TIMEOUT also finishes an association on which nothing moves at all.
Invariants
Section titled “Invariants”| Invariant | Enforced by | Pinned by |
|---|---|---|
One UDP association per client source; a second destination rides it without a second Open. |
associations map plus the FLOW_QUEUE channel in TunInbound::run; opened in TunUdpCore. |
udp_flows_share_one_association_per_source (protocols/tests/pipeline/tun.rs); the_first_packet_opens_the_association_and_every_packet_is_sent (protocols/tests/unit/tun/udp.rs) |
| A UDP reply returns as a packet from its origin to the client, addresses swapped; a domain-sourced reply is dropped. | fx.put_to in TunUdpCore::handle and the per-far-end lookup in TunUdpLink::poll_send_to. |
replies_go_back_as_packets_from_their_origin_and_domains_are_dropped (protocols/tests/unit/tun/udp.rs); udp_flows_share_one_association_per_source |
| A UDP association finishes when its outbound fails or idles. | Effect::Finish on ConnectFailed / OutboundError; Timing::expired in relay. |
the_association_finishes_when_the_outbound_fails_or_idles (protocols/tests/unit/tun/udp.rs) |
A TCP connection is dialled toward the packet’s destination, with the client’s source, only once the client has spoken; the first bytes are carried and the HTTP Host becomes the sniffed domain. |
PassthroughCore::on_transport and SniffPrefix; Effect::ForwardHeld. |
a_tcp_connection_becomes_a_stream (protocols/tests/pipeline/tun.rs); a_sniffing_passthrough_core_holds_the_prefix_and_opens_with_the_host, a_sniffing_passthrough_core_opens_on_the_sniff_deadline (protocols/tests/unit/core/mod.rs) |
| A silent TCP client is cut off. | tokio::time::timeout(HANDSHAKE_TIMEOUT, …) around each pre-open step in serve_stream. |
Not pinned by a dedicated test. |
At most max_flows TCP connections plus UDP associations run at once. |
Semaphore::new(max_flows) and try_acquire_owned; the OwnedSemaphorePermit moves into the spawned task and lives until it returns. |
Not pinned by a dedicated test. |
A TCP stream is counted until its blocking Drop has finished. |
TrackedTcp::drop takes and drops the stream before fetch_sub. |
Exercised by the teardown of both TUN pipeline tests (Fake::stop awaits shutdown). |
| Build rejects what a start would reject. | check_platform called from both build_tun_inbound and open; MTU, address and route parsing in build_tun_inbound. |
tun_owns_its_interface_and_takes_no_listener, tun_refuses_an_mtu_below_the_stack_floor, tun_parses_addresses_and_routes (app/tests/unit/inbound.rs) |
| The interface and its routes exist exactly as long as the process serves them. | Device-scoped routes (RouteScope::Link, output_interface); the kernel removes both when the last descriptor closes. |
a_routed_connect_is_answered_while_the_app_runs (app/tests/integration/e2e_tun.rs) |
Failure paths and cancellation
Section titled “Failure paths and cancellation”| Failure | Outcome |
|---|---|
open fails at startup (no CAP_NET_ADMIN, name in use, netlink error) |
bind_inbound returns the error and Instance::start fails with inbound <tag> bind tun <name> failed: … (tun auto when unnamed). The half-built device is already destroyed. |
open fails on reload |
spawn_generation runs non-strict: the error is logged as inbound <tag> bind tun <name> failed: … and the other inbounds start. The old generation is already gone, so the TUN inbound stays down until a later reload brings it up. |
TunDevice::new or IpStackConfig::mtu fails |
run returns the error; run_tun_inbound logs tun inbound failed: … and still awaits shutdown. |
stack.accept() returns an error |
The stack’s task has ended, which the code treats as the device being gone: the loop breaks and run returns Ok(()), so nothing is logged at error and the inbound serves nothing until the generation ends. |
| TCP client silent before open | serve_stream returns TimedOut (tun: the client never spoke), logged at debug as tun: tcp flow <src> -> <dst> ended: …. |
| Runtime error on a TCP flow | Converted with io::Error::other(e.to_string()) and logged at debug. |
| Dial fails on a TCP flow | The core shuts the transport down and finishes. |
| UDP association ends (all flows retired, outbound failed, idle) | The runtime’s result is logged at trace as tun: udp association of <src> ended: …; the task returns Some(src) and the map entry is removed. |
| Write to one UDP flow fails | That flow is retired; the association continues. |
Cancellation and shutdown
Section titled “Cancellation and shutdown”run is cancelled by its CancellationToken. Breaking out of the loop returns from run, which drops the JoinSet (aborting every flow task) and the IpStack. IpStack’s Drop aborts the stack task that owns the TunDevice, but an abort only lands when the tokio runtime next schedules that task, so the descriptor outlives the drop by a moment.
shutdown covers that gap. It takes the Weak device handle out of the mutex and polls every RELEASE_POLL (20 ms) until the handle no longer upgrades and live_tcp is zero, for at most RELEASE_TIMEOUT (3 s). On timeout it logs a warning (tun: device fd or <n> tcp flows still open after 3s) and returns.
sequenceDiagram participant I as Instance reload participant T as run_tun_inbound participant S as ipstack task participant N as next generation I->>T: token.cancel() T->>T: run returns, JoinSet and IpStack dropped T->>S: abort S-->>T: TunDevice dropped, fd closed T->>T: shutdown polls Weak and live_tcp T-->>I: accept handle completes I->>N: spawn_generation, open the same name
Instance::reload and Instance::shutdown await every accept handle after cancelling, and run_tun_inbound returns only after shutdown. That ordering lets the next generation recreate an interface with the same name, and it keeps a TCP stream from surviving into runtime shutdown. A reload therefore closes every TUN connection and recreates the interface and its routes.
Limits
Section titled “Limits”| Constant | Value | Where | Meaning |
|---|---|---|---|
DEFAULT_MTU |
1500 | protocols/src/tun/config.rs |
Interface and stack MTU when mtu is unset. |
| MTU floor | 1280 | build_tun_inbound; ipstack’s IpStackConfig::mtu |
The IPv6 minimum; lower values fail the build. |
DEFAULT_UDP_IDLE_TIMEOUT |
60 s | protocols/src/tun/config.rs |
ipstack’s per-stream UDP timeout, one stream per far end; passed to IpStackConfig::udp_timeout. |
DEFAULT_MAX_FLOWS |
65 536 | protocols/src/tun/config.rs |
Permits shared by TCP connections and UDP associations; sized like the TCP listener path, since each flow may cost the outbound a socket. |
FLOW_QUEUE |
16 | protocols/src/tun/udp.rs |
New UDP flows queued toward one association before further ones are dropped. |
RELEASE_TIMEOUT |
3 s | protocols/src/tun/inbound.rs |
Longest shutdown waits for the descriptor and live TCP streams. |
RELEASE_POLL |
20 ms | protocols/src/tun/inbound.rs |
Poll interval of that wait. |
HANDSHAKE_TIMEOUT |
10 s | protocols/src/core/mod.rs |
Per step before a TCP flow opens, applied by serve_stream. |
SNIFF_TIMEOUT |
300 ms | protocols/src/sniff/mod.rs |
Sniff window before a TCP flow opens without a domain. |
SNIFF_LIMIT |
4 KiB | protocols/src/sniff/mod.rs |
Bytes inspected for a domain. |
RELAY_IDLE_TIMEOUT |
300 s | protocols/src/core/mod.rs |
Idle limit of a relaying TCP connection or UDP association. |
PassthroughCore::BUF_SIZE |
8 KiB | protocols/src/core/mod.rs |
Each of the TCP runtime’s three buffers. |
TunUdpCore::BUF_SIZE |
8 KiB | protocols/src/tun/udp.rs |
Each of the UDP runtime’s three buffers. |
TunUdpCore::MAX_DATAGRAM |
4096 | protocols/src/tun/udp.rs |
Largest datagram delivered whole in either direction; a longer client datagram or reply is truncated, which only matters with an mtu above about 4 KiB. |
TunUdpCore::STAGING_RESERVE |
4096 | protocols/src/tun/udp.rs |
Staging room the runtime keeps free before it delivers an event, so each reply is staged whole as one packet. |
max_flows = 0 is accepted by the parser and makes the inbound drop every flow. udp_idle_timeout = 0 is accepted too; ipstack then times a flow out on the first read after its opening datagram, so the association ends almost at once.
| Test | File | What it proves |
|---|---|---|
udp_flows_share_one_association_per_source |
protocols/tests/pipeline/tun.rs |
One association per source; the reply comes back with swapped addresses and ports; a second destination rides the same association and gets its own replies. |
a_tcp_connection_becomes_a_stream |
protocols/tests/pipeline/tun.rs |
The stack answers the SYN; the dial happens after the request, carries it, and sniffs the Host; the reply returns as a segment. |
the_first_packet_opens_the_association_and_every_packet_is_sent |
protocols/tests/unit/tun/udp.rs |
The first datagram opens with the right destination and source, arms RELAY_IDLE_TIMEOUT, and every datagram becomes a SendTo. |
replies_go_back_as_packets_from_their_origin_and_domains_are_dropped |
protocols/tests/unit/tun/udp.rs |
Replies are staged as packets from their origin; domain-sourced replies are dropped. |
the_association_finishes_when_the_outbound_fails_or_idles |
protocols/tests/unit/tun/udp.rs |
ConnectFailed and the idle deadline both finish. |
tun_owns_its_interface_and_takes_no_listener, tun_refuses_an_mtu_below_the_stack_floor, tun_parses_addresses_and_routes |
app/tests/unit/inbound.rs |
Build-time validation and the DeviceSpec it produces. |
a_routed_connect_is_answered_while_the_app_runs |
app/tests/integration/e2e_tun.rs |
A real interface: a connect to a routed address is answered by the stack, and stops being answered once the app process and its interface are gone. |
The fake device
Section titled “The fake device”The pipeline tests need no privileges. protocols/tests/pipeline/tun.rs uses a UnixDatagram::pair() as the descriptor: one end becomes the OwnedFd passed to run, and each send or recv on the other end is exactly one IP packet, matching the one-packet-per-read contract of TunDevice. The tests build packets with etherparse::PacketBuilder (a dev-dependency of etemenanki-protocols) and parse the stack’s answers with SlicedPacket::from_ip.
A Capture connector stands in for the outbounds. For a TCP flow it returns one end of a tokio::io::duplex and hands the other to the test; for a UDP flow it returns a FakeLink whose sent packets and injected replies travel over unbounded test channels. Fake::stop mirrors the app’s teardown: cancel the token, await the task, drop the wire, then await TunInbound::shutdown.
The privileged end-to-end test
Section titled “The privileged end-to-end test”app/tests/integration/e2e_tun.rs runs the real binary with a tun inbound (address = ["10.77.0.1/24"], routes = ["10.77.1.0/24"]) and a blackhole outbound. Nothing can flow back in-process without a loop through the device, so the proof is the TCP handshake: the kernel routes the connect into the device, the stack answers the SYN, and connect succeeds until the process exits. The test first calls open itself and skips with SKIP: cannot create a tun device on PermissionDenied.
cargo test -p etemenanki-protocols --features tun --lib tun::cargo test -p etemenanki-protocols --features tun --test pipeline pipeline::tuncargo test -p etemenanki-app --bin etemenanki-app tun_sudo -E cargo test -p etemenanki-app --test integration e2e_tunSee Testing for the test layout across the workspace.