srdusr
aboutsummaryrefslogtreecommitdiffstats
path: root/PLAN.md
blob: 6defe1d60c5bf7cfc197c1c69fba3df225e4dc54 (plain) (blame)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
# wireframe - Packet Analyzer / Network TUI

## Overview
Terminal packet capture and analysis tool. Primary goal: learn the C++
memory model (byte layout, alignment, endianness, `std::span` over
unowned buffers) via a real-world capture pipeline.

## Stack
- Language: C++ (first of two C++ projects - build this one first)
- Capture: libpcap, or raw `AF_PACKET` with an mmap'd ring buffer to skip
  libpcap's copies
- Parsing: hand-rolled L2-L4 decoders over `std::span`, L7 dissectors as
  a small interface/vtable so protocols can be added incrementally
- Optional: `aya`-style in-kernel filtering isn't available in C++; if
  eBPF filtering is wanted later, that's a separate learning detour
- UI: TUI (library TBD - ftxui or notcurses are the usual C++ options)
- Output format: pcapng (not pcap) so interface metadata survives and
  files stay Wireshark-compatible

## Architecture sketch
- Capture thread (owns the pcap/AF_PACKET handle) -> bounded channel ->
  render/analysis thread. A traffic spike should drop packets, not
  block the UI.
- Drop privileges immediately after opening the capture handle; use
  `CAP_NET_RAW` via file capabilities instead of running as root.

## Build order
1. [done] Raw capture -> hex dump to stdout
2. [done] Ethernet/IP/TCP/UDP decoders + live packet list in TUI (-t)
3. [done] pcapng read/write
4. [done] Bounded channel + drop-on-backpressure between capture and render
5. [in progress] L7 dissector interface, add protocols incrementally --
   interface + DNS + HTTP + TLS SNI + mDNS + SSH banner done
   (wireframe/l7/); more protocols can still be added incrementally,
   by design
6. [done] Filtering (-f <expr>, libpcap's own BPF compiler - see Decisions)
7. [done] Drop privileges after opening the capture handle (see Decisions)
8. [done] TCP stream reassembly, opt-in via -a (see Decisions)

## Open questions
None currently open.

## Decisions
- TUI library: FTXUI (v7.0.3, fetched via CMake FetchContent). Chosen
  over notcurses for pure-C++ portability (no C build-system/dependency
  chain to fight on Gentoo/low-spec machines) and genuine native
  Windows console support, which notcurses lacks - both matter given
  this needs to work everywhere.
- GUI added as a secondary frontend - TUI stays primary (explicit
  user direction). Dear ImGui + SDL3 (v1.92.9b / release-3.4.14, both
  FetchContent, same approach as FTXUI/doctest - no system-package
  dependency, builds the same way everywhere). SDL3 over SDL2: this
  ecosystem already carries sdl2-compat as an SDL3-backed shim, so
  SDL3 is the live line, not legacy. SDL_Renderer backend, not raw
  OpenGL3 - avoids needing a separate GL function loader as another
  dependency, which matters more here than raw rendering performance
  does. src/gui_main.cpp; parity with the CLI/TUI is structural, not
  incidental - all three go through the same wireframe::CaptureSession
  (wireframe/capture_session.hpp) for device-open/datalink-validate/
  filter/pcapng/signal-handler setup, so the GUI can't silently skip a
  step (e.g. the DLT_RAW check) the way two hand-copied setups would
  eventually drift.
- Tests: doctest (v2.5.3, FetchContent), tests/ mirrors include/wireframe/.
  Every module gets unit tests as it's built, not backfilled later --
  `cmake --build build && ./build/wireframe_tests` (or `ctest`) should
  stay green at every commit.
- Filtering: libpcap's own pcap_compile()/pcap_setfilter() (tcpdump
  syntax, kernel-level via BPF), not a hand-rolled parser - the
  parser/compiler already exists, is correct, and reimplementing it has
  no bearing on this project's actual goal (the C++ memory model).
  wireframe/filter.hpp wraps compilation; testable without root via
  pcap_open_dead(). Verified live: -f "tcp port N" and -f icmp each
  correctly suppressed non-matching traffic that was actually present.
- pcap_stats(): CaptureSession::stats() surfaces kernel/interface-level
  drops (ps_recv/ps_drop/ps_ifdrop), shown in CLI/TUI/GUI whenever
  nonzero. Distinct from CaptureQueue::dropped() - verified live that
  the two really do measure different things: a short capture showed
  ps_recv=12 against only 4 packets actually rendered, i.e. packets the
  kernel had already received but that were never dispatched to our
  callback before shutdown, with queue-side drops at 0 throughout.
- Fuzzing: libFuzzer harnesses (fuzz/, clang + ASan/UBSan, opt-in via
  -DWIREFRAME_ENABLE_FUZZING=ON -DCMAKE_CXX_COMPILER=clang++, separate
  build-fuzz/ dir) for every hand-rolled decoder plus the pcapng reader
  and the full summarize_packet() pipeline - the highest-value tests
  in the repo given the project's actual goal (byte layout/alignment/
  UB on parsers over untrusted bytes), not an afterthought. Found and
  fixed a real bug on the first run: Reader::next_packet() allocated a
  block's claimed size (an untrusted 32-bit field straight from the
  file) before validating it, so a corrupted/hostile pcapng file could
  OOM the process. Fixed with a 1 MiB body-size cap (reader.hpp is
  explicitly scoped to pair with our own writer, whose packets are
  capped at a 65535 snaplen, so this is generous, not tight) and locked
  in with both a unit test and a passing re-fuzz of the exact crashing
  input. ~23M total fuzz executions across all 8 harnesses this
  session, one bug found and fixed, zero remaining crashes.
- HTTP L7 dissector (wireframe/l7/http.hpp): best-effort single-segment
  request/status-line parse (+ Host: header for requests), same scope
  DNS already has - no TCP stream reassembly, so a message split
  across packets is only partially visible. This is the first
  registered dissector to actually exercise L7Registry's TCP-payload
  path; DNS alone never did, since it only runs over UDP. Verified live
  against a real HTTP request/response (curl -> python http.server on
  port 80): both directions decoded correctly, including the
  dst-port-then-src-port fallback in l7_summarize (request matches on
  dst_port=80, response matches on src_port=80). Fuzzed separately
  (fuzz_http.cpp, 5.3M runs, no crashes) since the string_view request-
  line/header scanning is new hand-rolled logic distinct from anything
  fuzz_summarize's binary-format parsers already cover.
- TLS SNI L7 dissector (wireframe/l7/tls.hpp): parses a ClientHello's
  record/handshake/extensions structure (nested TLVs, every length
  bounds-checked against attacker-influenced fields at every level --
  the most structurally complex hand-rolled parser in the project) to
  extract the SNI extension. Answers what HTTP alone increasingly
  can't: most web traffic is TLS-encrypted, and the server name is the
  one thing still readable in cleartext, in every TLS version, before
  encryption starts. Same single-segment scope as DNS/HTTP. Verified
  against real, unsolicited internet traffic captured live on wlp1s0
  (not loopback/synthetic) - correctly extracted a genuine SNI from a
  real ClientHello. Fuzzed the hardest of any target so far given the
  nesting depth: fuzz_tls.cpp, 25.7M runs, no crashes.
- IPv6 extension headers: walk_ipv6_extension_headers() (ipv6.hpp)
  walks Hop-by-Hop, Routing, Destination Options, Fragment, and AH to
  find the real transport protocol underneath them, so e.g. TCP wrapped
  in a Hop-by-Hop options header is decoded instead of silently
  stopping. ESP is a deliberate hard stop, not an oversight: its own
  next-header field lives in a trailer after the encrypted payload, at
  an offset unknowable without decrypting first - reported as
  "ESP (encrypted)" rather than guessed at. parse_ipv6() itself stays
  an unconditional decode of just the fixed 40-byte header; the walk is
  a separate, composable function summarize.hpp calls, so parse_ipv6's
  existing tests didn't need to change. Verified end-to-end (a
  Hop-by-Hop-wrapped TCP frame decodes through to the TCP layer via
  summarize_packet, not just the walker in isolation) and fuzzed
  (extended fuzz_ipv6.cpp, 6.3M runs; fuzz_summarize.cpp indirectly
  covers it too, 4.3M more) - no crashes. This was the last item on
  the known-gaps list; none remain.
- Post-capture search: wireframe/search.hpp's matches_search() is a
  display filter, deliberately distinct from -f's capture filter --
  -f decides what's captured (and written to -w); search decides what's
  shown, without touching either, same distinction Wireshark draws
  between a capture filter and a display filter. CLI: -g <term> (only
  suppresses what's printed; -w output is unaffected). TUI: '/' opens
  live-filtered search (Enter keeps the filter and returns to
  browsing, Esc clears it), verified interactively via a real terminal
  (tmux capture-pane) - typing, backspace, both Enter and Esc paths,
  and confirmed 'q' still quits correctly afterward. GUI: a search box
  next to the capture-info line, using io.WantCaptureKeyboard to route
  Esc to "clear the search" while the box has focus vs. "quit the app"
  otherwise - verified visually via Xvfb, including the focused/
  unfocused Esc distinction actually working both ways.
  One real methodology lesson from building this: an initial pty-based
  interactive test of the TUI (raw-byte capture, regex-matched against
  unparsed ANSI escape sequences) appeared to show a redraw bug --
  typing "abc" only ever displayed "a". Chasing it added an unnecessary
  PostEvent "fix" before re-verifying under tmux (which properly
  resolves escape sequences via a real terminal emulator) showed the
  original code was correct all along; the first test method just
  wasn't reliable enough to trust. The PostEvent change was reverted --
  correct code, not narrowly-passing code, was the actual goal.
- Replay mode (-r <file>): reads a previously-saved pcapng file back
  through the exact same CaptureQueue/render/search pipeline as a live
  capture - the render/consumer side only ever talks to a
  CaptureQueue, so it can't tell whether packets are arriving from
  pcap_loop or being read back from disk. All in CaptureSession, so
  every frontend gets it for free rather than needing a second code
  path. Reader gained link_type() (the datalink from the file's IDB,
  previously discarded) so replayed packets decode with the correct
  DLT_EN10MB/DLT_RAW branch instead of an assumption; CaptureQueue
  gained a blocking push() alongside the existing drop-on-full
  try_push(), because a live capture thread can't be allowed to stall
  but a file has no real-time pressure forcing a drop - dropping from
  what's supposed to be a faithful replay of a fixed record would
  defeat the point of replaying it. -f (capture filter) is rejected
  outright when combined with -r, with an actionable error pointing at
  -g, rather than silently ignored.
  Verified live end-to-end, not just via unit tests: captured real
  traffic with -w on both DLT_EN10MB (lo) and DLT_RAW (tailscale0),
  replayed each file with -r with no root/live device needed, and the
  output matched the original capture exactly, including L7 dissection
  (DNS) surviving the round-trip. The replay thread finishing (not
  killed via request_stop()) means the file is exhausted, not that the
  user wants to quit - CaptureSession::stop_requested() distinguishes
  the two, and TUI/GUI both leave the window open on natural
  end-of-file (the point of replaying into an interactive frontend is
  browsing/searching afterward, not watching it flash by), closing only
  on an explicit 'q'/Esc/window-close or an external signal. Verified
  interactively in both: tmux capture-pane confirmed the TUI stays open
  with "[replay finished]" shown, search still works against the
  now-static list, and 'q' closes it; Xvfb confirmed the same for the
  GUI, including a live process check across a multi-second wait to
  rule out a delayed auto-close.
- Privilege dropping (wireframe/privileges.hpp): after pcap_open_live()
  succeeds - the only operation that actually needs CAP_NET_RAW - and
  before the datalink check or a -w file is even created, drop from
  root to the invoking user via sudo's SUDO_UID/SUDO_GID. setuid() to a
  nonzero UID clears the process's Linux capability sets as a kernel
  side effect too, so this covers both "ran via sudo" and "root's own
  CAP_NET_RAW" without a separate libcap dependency, and as a side
  benefit means -w's output file ends up owned by the real user, not
  root (previously needed a manual chown after every capture - every
  live test earlier this session did). Recommended usage skips this
  path entirely: `sudo setcap cap_net_raw+ep <binary>` once, then run
  unprivileged forever after, matching PLAN.md's original "use
  CAP_NET_RAW via file capabilities instead of running as root."
  Only the non-root no-op path is unit-testable without a test process
  permanently dropping its own privileges mid-suite, which would be a
  surprising thing for a unit test to do - so the real drop sequence
  was verified live instead: running via sudo, /proc/<pid>/status
  showed Uid go from 0 to the real UID and CapEff/CapPrm both go to
  zero within about a second of startup, with capture continuing to
  work correctly afterward (proving the already-open fd keeps working
  regardless of the process's current privilege level, which is the
  whole point of "drop after open"). Separately verified the
  setcap-without-sudo path works with zero privilege escalation at any
  point in the process's life.
- AF_PACKET/mmap ring buffer (src/afpacket_capture.cpp, wireframe_afpacket_demo,
  Linux-only): PLAN.md's originally-listed alternative capture backend,
  built as a standalone artifact rather than swapped into CaptureSession
  - the existing pipeline has real, tested value riding on libpcap's
  APIs (pcap_setfilter, pcap_stats, pcap_datalink) that a raw-socket
  path would need to reimplement from scratch at every one of
  CaptureSession's already-verified call sites, real risk to 90 passing
  tests for a copy-avoidance benefit modern libpcap on Linux already
  gets much of internally. TPACKET_V2 (simpler one-frame-per-slot
  layout than V3's block-batching) mmap'd directly into the process,
  packets read via std::span pointing straight into that kernel-shared
  mapping - no read()/recv(), no buffer of our own, genuinely zero
  copies between the NIC and summarize_packet() seeing the bytes. Reuses
  drop_privileges_if_root() (same principle, same code, right after the
  ring is mapped and bound). Verified against real traffic on both lo
  and the physical wlp1s0 interface - full TCP handshakes, DNS, mDNS,
  ICMPv6 all decoded correctly across a large volume of genuine
  traffic, no crashes, no leaked sockets/mappings after exit, tests and
  the rest of the build entirely unaffected by its addition.
- ICMP decoding (wireframe/net/icmp.hpp): previously every ICMPv4
  packet just showed "proto=1" with nothing further - no dissector
  existed at all - despite ICMP being most of this session's own test
  traffic (every ping). ICMPv6 was labeled but not decoded either.
  ICMPv4 and ICMPv6 share the same first-4-byte shape (type/code/
  checksum) but a completely different type namespace - the same
  number means something different in each (ICMPv4 type 8 is Echo
  Request; ICMPv6's Echo Request is 128, and its own type 8 isn't
  defined at all) - so they get separate parse functions and type-name
  tables, not one shared by number. Neither protocol has ports, so this
  doesn't fit L7Registry's port-keyed dispatch; both are handled
  directly by IP protocol number in summarize_transport_and_above
  instead. Verified live against real ping traffic on both lo (proto=1)
  and ::1 (proto=58) - request/reply pairs decoded correctly on both,
  including matching identifier/sequence numbers between each request
  and its reply.
- -h/--help: both wireframe and wireframe_gui now print real usage
  text (each binary's actual flag set - the GUI never had -x/-t/-g,
  so its help doesn't claim it does) and exit 0 before touching a
  device or any privilege at all. Previously -x -t -w -f -g -r all
  existed with zero discoverability outside reading the source.
- Checksum validation (wireframe/net/checksum.hpp): RFC 1071 Internet
  checksum, plus IPv4-header/TCP/UDP verification built on it (IPv6
  checksums use a different pseudo-header and different optionality
  rules - not done here, a reasonable follow-on if wanted). UDP's
  checksum is optional over IPv4: a transmitted value of exactly
  0x0000 means "not computed", reported as its own kNotPresent state,
  not folded into invalid. Deliberately not part of summarize_packet's
  shared output - opt-in via the CLI's -c flag only (same
  plain-text-mode-only precedent -x/hex-dump already set), because
  checksum offload means many outbound and loopback packets can
  legitimately show an invalid checksum with nothing actually wrong:
  the NIC computes the real one in hardware during DMA, which is often
  after the capture point already saw the packet. Wireshark makes this
  opt-in for the same reason.
  internet_checksum() itself is verified against RFC 1071's own worked
  example (an external reference value, not derived from this code),
  not just internal self-consistency. Live-tested on lo and the
  physical wlp1s0 - both showed IP=ok/UDP=ok/TCP=ok throughout; `ethtool
  -k wlp1s0` shows tx-checksumming off on this machine's driver, which
  is exactly why (no hardware offload means the kernel computes real
  checksums in software) - so the "offload causes false BAD" case this
  feature exists to route around couldn't be reproduced on this
  specific sandboxed machine's NIC, but that's a property of this
  hardware, not a gap in the reasoning: most real NICs ship tx-checksum
  offload on by default, which is exactly the scenario -c's
  opt-in-ness is meant to keep from reading as false positives.
- TCP stream reassembly (wireframe/net/tcp_reassembly.hpp): in-order-only
  - out-of-order segments and retransmissions are dropped, not buffered
  for later reordering. A real limitation, but an honest one for a
  learning tool captured directly on an endpoint (lo/wlp1s0/tailscale0,
  everything this project has actually run against), where segments
  mostly do arrive in order; a capture point far from either endpoint
  (a middlebox) would need real reorder buffering this doesn't attempt.
  Deliberately kept out of summarize_packet()'s shared signature and the
  TUI/GUI consumer loops - adding a TcpReassembler& parameter there
  would ripple into every call site and both frontends' render paths,
  risking the (at the time) 108 passing tests for a single opt-in
  feature. Instead it's CLI-only, opt-in via -a, same
  plain-text-mode-only precedent -x/-c already set: a separate
  TcpReassembler instance lives in main(), and render_packet() does its
  own minimal Ethernet/IPv4/TCP walk (mirroring checksum_status()) to
  feed segments in and, when new contiguous bytes come back, re-runs
  parse_http() (wireframe/l7/http.hpp) against the joined stream and
  prints the result as a distinct "[reassembled ...]" line, not folded
  into the per-packet summary. Deliberately calls parse_http() directly
  rather than going through L7Registry, so it isn't gated to port 80 the
  way the shared per-packet summary is - a deliberate difference, not
  an oversight. Live-verified against real split traffic: a Python
  client sent an HTTP request's request-line and its Host: header in
  two separate sendall() calls 0.3s apart with TCP_NODELAY set (to stop
  the kernel coalescing them back into one segment), captured on lo.
  The first segment's reassembled view showed the request line with no
  Host: (correct - it hadn't arrived yet); only once the second
  segment landed did Host: appear, confirming the two segments were
  actually joined rather than the dissector getting lucky on one
  segment alone. Buffers are capped per direction (64 KiB default) and
  the flow table is capped in total flow count, both to bound memory
  without needing active FIN/RST-triggered flow teardown - simpler,
  and stale entries past those caps don't affect correctness, just
  bounded memory use.
- GUI parity for -c/-a: checksum_status() and reassembled_http_status()
  moved out of main.cpp into a new shared header
  (wireframe/packet_diagnostics.hpp) rather than duplicated into
  gui_main.cpp - the same reasoning wireframe::CaptureSession exists
  for at the setup layer, applied here to the diagnostics layer. GUI's
  hex dump was already always-on for the selected row (no -x-equivalent
  flag needed); checksum status is computed lazily when a row is
  selected (stateless, cheap); reassembled HTTP status has to be
  computed at consume time instead (reassembly needs in-order state
  across packets), so PacketRow gained an
  optional<string> reassembled_http field set once in consumer_loop.
  Both are still opt-in via the same -c/-a flag names as the CLI, off
  by default. Visually verified under Xvfb (python-xlib synthetic
  click) with the same split-segment scenario used to verify -a on the
  CLI: selecting the packet whose segment completed the request showed
  both "checksums: IP=ok TCP=BAD" and "reassembled request: GET
  /split-test Host: split.example.com (72 bytes so far)" together in
  the details pane, and a second run with neither flag confirmed both
  lines are absent by default. The TCP=BAD reading on lo in that
  screenshot is the checksum-offload caveat working as documented, not
  a bug - Linux's loopback receive path typically never computes a
  real TCP checksum at all (CHECKSUM_UNNECESSARY), which is exactly the
  false-positive scenario -c's opt-in-ness exists to guard against.
- mDNS (wireframe/l7/mdns.hpp) and SSH banner (wireframe/l7/ssh.hpp)
  dissectors, registered alongside DNS/HTTP/TLS in l7_registry().
  mDNS reuses parse_dns() outright - RFC 6762 keeps DNS's exact wire
  format, just over UDP 5353 instead of 53 - and deliberately omits
  the id= field DNS's own summary shows, since RFC 6762 18.1 has
  multicast queries send it as zero, which would just be "id=0" noise
  on every real packet. SSH's identification banner (RFC 4253 4.2) is
  the one part of an SSH connection ever sent in the clear - a single
  line before key exchange encrypts everything else - so unlike every
  other dissector here, there's structurally nothing further to add to
  it later. Live-verified against real traffic: SSH against this
  machine's actual running sshd via /dev/tcp on lo, correctly decoding
  "SSH 2.0 OpenSSH_10.4" (matching the real installed OpenSSH version);
  mDNS via a real DNS-wire-format query sent to 127.0.0.1:5353 (no
  avahi/mDNS responder running on this sandboxed machine, so a
  synthetic-but-wire-format-real packet substituted for organic
  traffic), correctly decoding "mDNS query myhost.local type=1" with no
  id= field present.
- Fuzzing: two new libFuzzer harnesses added alongside the original
  nine - fuzz_checksum (internet_checksum/verify_ipv4_checksum
  directly, plus verify_tcp/udp_checksum_ipv4 with the first 8 input
  bytes providing src/dst addresses) and fuzz_tcp_reassembly (unlike
  every other harness, drives *one* TcpReassembler with a whole
  sequence of segments parsed out of a single input, since the
  interesting bugs in cross-call state - the flow map, per-direction
  sequence tracking, the buffer cap - don't show up from one segment
  alone). All 12 harnesses (the original nine plus these two) run
  clean - no crashes, no ASan/UBSan errors, no leak/timeout artifacts
  - across roughly 90 million total executions in a 20-second-each
  pass; summarize's own harness also now exercises the ICMP decode path
  added earlier this session and the new mDNS/SSH dissectors, since all
  of that lives inside summarize_packet()'s call graph already.