diff options
| author | srdusr <[email protected]> | 2024-05-17 19:54:00 +0200 |
|---|---|---|
| committer | srdusr <[email protected]> | 2024-05-17 19:54:00 +0200 |
| commit | e0f4c701028aa81026a17cf9ebfb36112184f4bc (patch) | |
| tree | 31c05e4ccbba0dd2ab4c0567630275ebfc6cd264 /PLAN.md | |
| parent | 08332a4195956611db80a2cfe3710d760cbd6acf (diff) | |
| download | packeteer-e0f4c701028aa81026a17cf9ebfb36112184f4bc.tar.gz packeteer-e0f4c701028aa81026a17cf9ebfb36112184f4bc.zip | |
Add privilege dropping, AF_PACKET demo, ICMP, checksum validation, --help, and TCP reassembly
Rounds out the build order in PLAN.md with six incremental additions:
drop root privileges immediately after opening the capture handle;
a standalone AF_PACKET/mmap ring-buffer demo (kept separate from
CaptureSession, see its header comment for why); ICMPv4/ICMPv6 type
and code decoding; opt-in IPv4/TCP/UDP checksum validation (-c);
CLI --help; and opt-in, in-order-only TCP stream reassembly (-a) so
HTTP requests/responses split across segments can be seen whole.
Each addition is unit-tested and, where it touches live traffic
behavior, verified against real captured packets - see PLAN.md's
Decisions section for the verification notes on each.
Diffstat (limited to 'PLAN.md')
| -rw-r--r-- | PLAN.md | 126 |
1 files changed, 126 insertions, 0 deletions
@@ -33,6 +33,8 @@ unowned buffers) via a real-world capture pipeline. interface + DNS + HTTP + TLS SNI done (wireframe/l7/); more protocols can still be added incrementally, by design 6. [done] Filtering (-f <expr>, libpcap's own BPF compiler - see Decisions) +7. [done] Drop privileges after opening the capture handle (see Decisions) +8. [done] TCP stream reassembly, opt-in via -a (see Decisions) ## Open questions None currently open. @@ -186,3 +188,127 @@ None currently open. now-static list, and 'q' closes it; Xvfb confirmed the same for the GUI, including a live process check across a multi-second wait to rule out a delayed auto-close. +- Privilege dropping (wireframe/privileges.hpp): after pcap_open_live() + succeeds - the only operation that actually needs CAP_NET_RAW - and + before the datalink check or a -w file is even created, drop from + root to the invoking user via sudo's SUDO_UID/SUDO_GID. setuid() to a + nonzero UID clears the process's Linux capability sets as a kernel + side effect too, so this covers both "ran via sudo" and "root's own + CAP_NET_RAW" without a separate libcap dependency, and as a side + benefit means -w's output file ends up owned by the real user, not + root (previously needed a manual chown after every capture - every + live test earlier this session did). Recommended usage skips this + path entirely: `sudo setcap cap_net_raw+ep <binary>` once, then run + unprivileged forever after, matching PLAN.md's original "use + CAP_NET_RAW via file capabilities instead of running as root." + Only the non-root no-op path is unit-testable without a test process + permanently dropping its own privileges mid-suite, which would be a + surprising thing for a unit test to do - so the real drop sequence + was verified live instead: running via sudo, /proc/<pid>/status + showed Uid go from 0 to the real UID and CapEff/CapPrm both go to + zero within about a second of startup, with capture continuing to + work correctly afterward (proving the already-open fd keeps working + regardless of the process's current privilege level, which is the + whole point of "drop after open"). Separately verified the + setcap-without-sudo path works with zero privilege escalation at any + point in the process's life. +- AF_PACKET/mmap ring buffer (src/afpacket_capture.cpp, wireframe_afpacket_demo, + Linux-only): PLAN.md's originally-listed alternative capture backend, + built as a standalone artifact rather than swapped into CaptureSession + - the existing pipeline has real, tested value riding on libpcap's + APIs (pcap_setfilter, pcap_stats, pcap_datalink) that a raw-socket + path would need to reimplement from scratch at every one of + CaptureSession's already-verified call sites, real risk to 90 passing + tests for a copy-avoidance benefit modern libpcap on Linux already + gets much of internally. TPACKET_V2 (simpler one-frame-per-slot + layout than V3's block-batching) mmap'd directly into the process, + packets read via std::span pointing straight into that kernel-shared + mapping - no read()/recv(), no buffer of our own, genuinely zero + copies between the NIC and summarize_packet() seeing the bytes. Reuses + drop_privileges_if_root() (same principle, same code, right after the + ring is mapped and bound). Verified against real traffic on both lo + and the physical wlp1s0 interface - full TCP handshakes, DNS, mDNS, + ICMPv6 all decoded correctly across a large volume of genuine + traffic, no crashes, no leaked sockets/mappings after exit, tests and + the rest of the build entirely unaffected by its addition. +- ICMP decoding (wireframe/net/icmp.hpp): previously every ICMPv4 + packet just showed "proto=1" with nothing further - no dissector + existed at all - despite ICMP being most of this session's own test + traffic (every ping). ICMPv6 was labeled but not decoded either. + ICMPv4 and ICMPv6 share the same first-4-byte shape (type/code/ + checksum) but a completely different type namespace - the same + number means something different in each (ICMPv4 type 8 is Echo + Request; ICMPv6's Echo Request is 128, and its own type 8 isn't + defined at all) - so they get separate parse functions and type-name + tables, not one shared by number. Neither protocol has ports, so this + doesn't fit L7Registry's port-keyed dispatch; both are handled + directly by IP protocol number in summarize_transport_and_above + instead. Verified live against real ping traffic on both lo (proto=1) + and ::1 (proto=58) - request/reply pairs decoded correctly on both, + including matching identifier/sequence numbers between each request + and its reply. +- -h/--help: both wireframe and wireframe_gui now print real usage + text (each binary's actual flag set - the GUI never had -x/-t/-g, + so its help doesn't claim it does) and exit 0 before touching a + device or any privilege at all. Previously -x -t -w -f -g -r all + existed with zero discoverability outside reading the source. +- Checksum validation (wireframe/net/checksum.hpp): RFC 1071 Internet + checksum, plus IPv4-header/TCP/UDP verification built on it (IPv6 + checksums use a different pseudo-header and different optionality + rules - not done here, a reasonable follow-on if wanted). UDP's + checksum is optional over IPv4: a transmitted value of exactly + 0x0000 means "not computed", reported as its own kNotPresent state, + not folded into invalid. Deliberately not part of summarize_packet's + shared output - opt-in via the CLI's -c flag only (same + plain-text-mode-only precedent -x/hex-dump already set), because + checksum offload means many outbound and loopback packets can + legitimately show an invalid checksum with nothing actually wrong: + the NIC computes the real one in hardware during DMA, which is often + after the capture point already saw the packet. Wireshark makes this + opt-in for the same reason. + internet_checksum() itself is verified against RFC 1071's own worked + example (an external reference value, not derived from this code), + not just internal self-consistency. Live-tested on lo and the + physical wlp1s0 - both showed IP=ok/UDP=ok/TCP=ok throughout; `ethtool + -k wlp1s0` shows tx-checksumming off on this machine's driver, which + is exactly why (no hardware offload means the kernel computes real + checksums in software) - so the "offload causes false BAD" case this + feature exists to route around couldn't be reproduced on this + specific sandboxed machine's NIC, but that's a property of this + hardware, not a gap in the reasoning: most real NICs ship tx-checksum + offload on by default, which is exactly the scenario -c's + opt-in-ness is meant to keep from reading as false positives. +- TCP stream reassembly (wireframe/net/tcp_reassembly.hpp): in-order-only + - out-of-order segments and retransmissions are dropped, not buffered + for later reordering. A real limitation, but an honest one for a + learning tool captured directly on an endpoint (lo/wlp1s0/tailscale0, + everything this project has actually run against), where segments + mostly do arrive in order; a capture point far from either endpoint + (a middlebox) would need real reorder buffering this doesn't attempt. + Deliberately kept out of summarize_packet()'s shared signature and the + TUI/GUI consumer loops - adding a TcpReassembler& parameter there + would ripple into every call site and both frontends' render paths, + risking the (at the time) 108 passing tests for a single opt-in + feature. Instead it's CLI-only, opt-in via -a, same + plain-text-mode-only precedent -x/-c already set: a separate + TcpReassembler instance lives in main(), and render_packet() does its + own minimal Ethernet/IPv4/TCP walk (mirroring checksum_status()) to + feed segments in and, when new contiguous bytes come back, re-runs + parse_http() (wireframe/l7/http.hpp) against the joined stream and + prints the result as a distinct "[reassembled ...]" line, not folded + into the per-packet summary. Deliberately calls parse_http() directly + rather than going through L7Registry, so it isn't gated to port 80 the + way the shared per-packet summary is - a deliberate difference, not + an oversight. Live-verified against real split traffic: a Python + client sent an HTTP request's request-line and its Host: header in + two separate sendall() calls 0.3s apart with TCP_NODELAY set (to stop + the kernel coalescing them back into one segment), captured on lo. + The first segment's reassembled view showed the request line with no + Host: (correct - it hadn't arrived yet); only once the second + segment landed did Host: appear, confirming the two segments were + actually joined rather than the dissector getting lucky on one + segment alone. Buffers are capped per direction (64 KiB default) and + the flow table is capped in total flow count, both to bound memory + without needing active FIN/RST-triggered flow teardown - simpler, + and stale entries past those caps don't affect correctness, just + bounded memory use. |