srdusr
aboutsummaryrefslogtreecommitdiffstats
path: root/PLAN.md
blob: a070b20949ba4f37dc75c811befa79bb79f8edee (plain) (blame)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
# mitmux - Intercepting Proxy TUI (Burp/Caido replacement)

## Overview
Daily-driver intercepting proxy for manual pentest work, terminal-based.
Prior art to read before writing code: Cruster (Rust, built on
hudsucker) - same problem, worth studying even though this build is Go.

## Stack
- Language: Go - memory safety on hostile input matters here more than
  in the other projects, since this parses attacker-adjacent traffic
- TLS interception: Go's own `crypto/tls` + a CA cert generator
  (analogous to `rcgen`) for per-domain leaf certs
- Proxy core: `net/http` + manual `CONNECT` handling, or a MITM proxy
  library if one fits without fighting Go's aggressive header
  normalization. Upstream requests are round-tripped manually (write
  the request, read the response off the same connection) rather than
  through `http.Transport` - Transport's automatic HTTP/2 dispatch keys
  off a literal `*tls.Conn` type assertion on the dialed connection,
  which a raw-byte-capturing wrapper around that connection defeats
  (found by testing: it silently parsed HTTP/2 framing as HTTP/1.1).
- Storage: SQLite in WAL mode - blob columns for raw request/response
  bytes, FTS5 index for search across bodies
- UI: Bubble Tea + Lipgloss (TUI), same family as the packet analyzer's
  Go sibling if that ever gets built

## Architecture sketch (important - don't skip this)
- Split proxy engine from TUI. Headless daemon owns the listening
  socket and the DB; TUI is a client over a Unix socket. The proxy
  keeps running when the UI restarts, and a web UI or CLI scanner can
  be bolted on later without touching the engine.
- Store raw bytes as the source of truth. Parse into a display view,
  never re-serialize for storage - request smuggling, header injection,
  and parser-differential bugs depend on the original malformed framing
  surviving. For Repeater specifically, write requests as raw bytes
  over the socket rather than through a normalizing HTTP client.

## Build order
1. Proxy + CA cert generation + plaintext HTTP passthrough
2. TLS interception (per-host cert generation, install CA)
3. History view (SQLite storage, raw bytes preserved) in the TUI
4. Repeater (raw-byte send/resend, the feature used daily)
5. Search/filter (FTS5)
6. Match-and-replace rules
7. Intruder-equivalent (last, optional)

## Open questions
- HTTP/2: handle natively (decided) - full fidelity over MITM'd
  connections rather than downgrading to HTTP/1.1. Adds complexity to
  CONNECT handling, stream framing, and step 3 storage (multiplexed
  streams over one connection need per-stream request/response
  boundaries, not just per-connection ones).
- CA install UX per OS (Linux/macOS/Windows trust stores)
- Whether WebSocket interception is v1 or a later addition
- Step 6 match-and-replace shipped headers-only. Body rules are a
  separate, harder problem: request-body capture currently relies on
  streaming the body straight from the client connection to the
  upstream write (that's what makes it exact, byte for byte); a body
  rule needs to materialize, transform, and re-send it instead, which
  means deciding what "exact" even means for a rule-modified request
  before touching that path again. Also still single-line-text-field
  limited in the TUI (bubbles/textinput can't hold a literal CRLF), so
  even with header rules, injecting a brand-new header line via the
  form isn't possible yet - only rewriting/removing existing ones. The
  underlying engine (rules.ApplyHeaders) already supports arbitrary
  text-block edits; it's specifically the form UI that's constrained.
- Step 7 (Intruder-equivalent) shipped Sniper only: one payload set,
  one §marked§ position fuzzed at a time, every other marked position
  held at its base value - the mode that covers most real Intruder
  usage. Battering ram / pitchfork / cluster bomb aren't implemented.
  Sequential sending only (no concurrency), capped at 1000 generated
  requests as a fixed safety limit against an accidental huge wordlist
  combined with several positions. Reuses the Repeater send primitive
  (proxy.Server.sendRaw) directly - an attack is just that primitive
  run in a loop with generated bytes - and results land in the same
  history table tagged source="intruder", same as Repeater's
  source="repeater", rather than a separate results store.

## Post-build-order: Burp/ZAP/Caido parity pass

Build order 1-7 is done. Researched what those three actually offer
(features and basic UI/UX) and triaged the gap into "should build soon"
/ "worth considering" / "skip" - see commit history for the full list;
tracking what's shipped vs. deferred here.

Shipped: vi-modal editing for the raw request textareas (table and
viewport already had vi nav by default - this was specifically about
textarea/textinput, which don't); a persistent status bar and a '?'
keybinding reference; display-only response JSON pretty-printing;
structured search filters (status:, source:, flagged:) alongside the
existing FTS5 text search; a flagged marker (★) for "revisit this" -
deliberately simpler than full free-text notes/comments, which would
need their own text-input overlay for comparatively modest extra value
over a boolean; noted as a real follow-up, not dropped silently; a
Comparer tool - mark an entry with 'c' (from history list or detail
view), 'c' again on a different entry opens a unified diff (git-diff
style, colored) of either side's request or response. Unified rather
than Burp's side-by-side: a two-column layout fights terminal width for
anything but a narrow window, and unified reuses the same scrollable-
viewport pattern already used everywhere else in the TUI. CRLF is
normalized to LF before diffing (display-only, same reasoning as the
JSON pretty-printer) so an HTTP/1.1 exact capture doesn't show every
line as changed from an invisible trailing \r; a standalone Decoder
tool ('d') - URL/Base64/Hex/HTML encode and decode, live output as you
type, tab to cycle transforms. Deliberately single-transform, not
chained/pipelined like Burp's Decoder - v1 scope, and pipeline-building
UI is real added complexity for a feature that's already useful without
it. Base64 decode tries standard/URL-safe/padded/unpadded variants in
turn rather than making the user pick, since real pasted data is as
likely to be one as the other. URL encode/decode uses strict RFC 3986
percent-encoding (space <-> %20), not Go's url.QueryEscape's form-
encoding behavior (space <-> '+'), since "URL encode" for a pentester
almost always means the former.

Shipped since: a standalone Decoder tool ('d') - see above; multiple
concurrent Repeater tabs - 'r' from the history list or detail view now
opens a NEW tab rather than overwriting whatever was already open,
`]`/`[` switch tabs, `ctrl+w` closes the active one (all three gated to
normal mode, so they're inert while typing - `[`/`]` are common JSON
body characters and `ctrl+w` is a textarea binding for delete-word-
backward that must still work while composing a request). Each tab owns
its own request buffer, response view, and send-in-flight state; a slow
send whose result lands after the user has switched away still updates
the correct tab (send results carry the tab index they belong to), and
the status line/response pane it's shown in only updates live if that
tab is still the one on screen.

Shipped since: Intruder payload processing and grep-match/grep-extract.
Payload processing - an optional case rule (upper/lower) and an optional
encode rule (URL/Base64/Hex/HTML), cycled with `c`/`e` - is applied
client-side to each payload line before it ever crosses the IPC socket,
case first then encode (encoding an already-case-folded value is safe;
the reverse would corrupt e.g. Base64 padding), since it's a pure string
transform with no proxy-side state involved and reuses the Decoder's own
`urlEncodeAll`. Grep-match/grep-extract are optional Go regexps
(`m`/`v` to edit, both gated to normal mode and both revert-on-esc /
validate-on-enter the same way the history list's `/` search box
works), evaluated server-side in `internal/ipc/server.go`'s "intrude"
handler against each result's actual response bytes - chosen over a
client-side implementation because the daemon already has `entry.
ResponseRaw` in hand right where the result is built, and Burp's own
grep options work the same way (matched against the real response, not
a client-refetched copy). Grep-match flags a result (shown as a Match
column); grep-extract captures the first submatch (or the whole match
if the pattern has no capturing group) into an Extract column. Both are
configured once before `ctrl+r` starts an attack and apply for that run
only - matching Burp, which doesn't retroactively re-grep already-fired
requests if you change the options mid-attack.

Shipped since: CA install UX per OS (`mitmuxd -install-ca`) - generates
the CA if needed, prints copy-pasteable install steps for the detected
platform, and exits without starting the proxy. Deliberately
instructions-only, never auto-executing: trust-store tooling varies
enough across Linux distros that guessing wrong and running the wrong
command unattended is worse than asking, and installing a root CA is a
system-wide trust change that affects every TLS connection on the
machine, not just mitmux's own traffic - the user running the printed
command themselves keeps them in control of that. On Linux, detects
`trust` (p11-kit - Arch, and Fedora also ships it)/`update-ca-trust`
(RHEL/Fedora/CentOS)/`update-ca-certificates` (Debian/Ubuntu/Gentoo)
via PATH lookup and picks whichever is actually present, plus separate
`certutil` (NSS) instructions for Firefox/Chrome's own certificate
store, which doesn't always follow the system trust store on Linux.
macOS (`security add-trusted-cert`) and Windows (`certutil -addstore` /
`Import-Certificate`) instructions are implemented but, unlike the
Linux path, not verified live - no macOS/Windows machine was available
to test against; only the command text itself (sourced from each
platform's standard, documented tooling) is confirmed correct by
inspection.

Shipped since: multiple proxy listeners and upstream proxy chaining -
the last two items from the original "worth considering" list.

Multiple listeners: `-listen` takes a comma-separated address list
(`-listen "127.0.0.1:8080,127.0.0.1:8081"`); all bound addresses share
the same handler, history store, CA and rules - one logical proxy
reachable on more than one address/port, not several independent
proxies in one process. Every address is bound up front before any of
them start serving, so a bad address fails startup immediately instead
of leaving the daemon partially listening.

Upstream proxy chaining: `-upstream-proxy host:port` (optional
`http://` prefix, stripped) routes every outbound connection through
another HTTP CONNECT proxy instead of dialing origins directly -
chaining mitmux into Burp, a corporate proxy, or any other
CONNECT-speaking proxy. SOCKS5 upstreams aren't implemented. For the
CONNECT/HTTPS path, chaining is transparent below the tunnel: once the
CONNECT handshake to the upstream proxy succeeds, TLS and the
request/response on top of it are identical to a direct connection, so
no other code needed to change. The plain-HTTP path is different: an
absolute-form request line ("GET http://host/path HTTP/1.1") has to be
sent to the upstream proxy instead of origin-form, so `roundTripH1`
gained a `proxyForm` parameter and `forward()` selects it based on
whether the request is plain HTTP and an upstream proxy is configured.

Chaining into another intercepting/MITM proxy (including another
mitmuxd instance) will fail TLS verification unless that proxy's own CA
is separately trusted - expected, not a mitmux-specific gap: the
upstream MITM terminates and re-signs the connection with its own CA,
which mitmux's outbound TLS client (verifying against the system root
store) has no reason to trust. Confirmed live: chaining through a
genuine passthrough CONNECT proxy (tunnels raw bytes, no MITM) works
correctly for both plain HTTP and HTTPS; chaining through a second
mitmuxd instance correctly fails with a clear
"certificate signed by unknown authority" error recorded in history,
rather than hanging or crashing.

This closes every item from the original "worth considering" list.

## Post-audit hardening and history management

A hands-on robustness audit (two parallel passes, backend and frontend,
actually driving the daemon/TUI against adversarial input rather than
reading code - "run it, don't read it") found and fixed six real bugs:
raw ANSI/control-character injection from captured traffic reaching the
operator's actual terminal (severe - confirmed a malicious Host value
changed the real tmux pane title); a table-cursor desync that left
enter/r/i/f/c inert on a live-captured entry until an unrelated
navigation keypress; Intruder silently hanging 60s per payload on
body-parameter fuzzing because Content-Length was never recalculated
after marker substitution; match-and-replace rules accepting an invalid
regex with zero validation or feedback, silently never firing; captures
that hit the 10 MiB cap being marked "exact" anyway, hiding data loss
from exactly the kind of investigation that needs the tail of a large
body; and no timeouts anywhere, so a slow-loris connection or a client
that completed CONNECT and never sent a TLS ClientHello held a
connection and goroutine open forever. See commit history for full
detail on each - every fix was verified against the actual failure
mode, not just code-reviewed.

Also added: history deletion. `Store` had full CRUD for rules but no
way to delete or prune history - it only ever grew, with no way to
remove an accidental capture or start fresh for a new engagement short
of manually deleting the DB file outside the tool. `x` deletes the
selected entry, `X` clears the entire database (explicitly not scoped
to an active search filter - the confirmation always states the true
total count, since understating it would make the prompt itself
misleading about what's about to happen). Both gated behind a `y`/`n`
confirmation: a small reusable confirmPrompt/confirmYes pattern in the
TUI model, checked first in the history list's key handling, so any key
other than y/Y safely cancels rather than falling through to whatever
that key normally does.

Shipped since: single-entry export. `e` from Detail view writes the
selected entry's raw request and response bytes to a plain-text file -
a modal path-prompt (same pattern as the Intruder grep-match/extract
edit buffers: enter writes and confirms, esc cancels), prefilled with a
sensible default filename. Deliberately plain text, not a structured
format: for "attach this to a report" or "grep it later," the raw bytes
as text are the point, matching this tool's own raw-bytes-first
philosophy rather than reformatting them into something else. Each
side is annotated when it isn't a wire-exact capture (reconstructed vs.
truncated, matching the Detail view's own labels) so the exported file
carries the same trust information the UI already shows, not a blanker
claim.

Shipped since: bulk export. `E` from the history list exports the
current view (respecting an active search filter - explicitly the
filtered set, not always everything, unlike `X` clear-all which is
deliberately the opposite) as a HAR 1.2 file, chosen specifically for
interop: DevTools, Burp, Postman, and others can all import it, which a
mitmux-specific format couldn't do. Building it means parsing each
entry's raw request/response bytes back into structured HAR fields
(method, url, headers, status, body) via the same net/http parsing the
capture path and prettyResponse already use - reused, not
reimplemented. A binary body is base64-encoded (HAR's "encoding" field)
rather than passed through as a JSON string, which would silently
corrupt it: encoding/json replaces invalid UTF-8 with U+FFFD by
default, exactly the failure mode that would quietly corrupt an
exported image or protobuf body with no error anywhere. An entry that
fails to fetch (daemon round trip) or parse (a deliberately malformed
Repeater request, say) is skipped rather than aborting the whole
export - the status line reports how many, so a partial export is
visible, not silent.

Shipped since: target scope. `s` from the history list opens scope
management - add/toggle/delete rules matching a host by substring
(case-insensitive, so "example.com" matches "www.example.com" and
"api.example.com" too, covering "this domain and its subdomains"
without a separate wildcard syntax) or by regex, mirroring the same
Match-text-or-regex toggle match-and-replace rules already use for one
consistent mental model. No rules configured (or none enabled) means
everything is recorded - today's behavior before scope existed at all,
unchanged, so a fresh install or a user who never opens the scope view
keeps recording everything rather than silently nothing.

Scope only filters what gets recorded, not what gets proxied: an
out-of-scope request still reaches its destination and its response
still reaches the client completely normally (see `internal/proxy`'s
`forward()` - the response is already written to the client by the
time the scope check runs; skipping the record step only skips
storage). This was a deliberate choice over blocking out-of-scope
traffic outright, which would be a materially different, much riskier
feature - an access-control mechanism, not a noise filter, and a wrong
scope pattern could silently break the very traffic the user is trying
to test. Repeater and Intruder deliberately bypass the scope check
entirely (recordRaw, a separate code path from the passive-capture
record()): a user explicitly resending or fuzzing a specific request
wants to see the result regardless of scope, which exists to cut
passive-capture noise (CDNs, analytics, trackers, unrelated third-party
hosts), not to second-guess a deliberate action. Verified live: added a
substring scope rule for one host, confirmed a request to a
non-matching host still proxied successfully (200 response reached the
client) but was never recorded, confirmed the matching host's requests
were recorded, confirmed a Repeater resend of the excluded host WAS
recorded despite being out of scope, confirmed toggling the rule off
resumed recording everything, and confirmed both the substring and
regex pattern forms save and match correctly.

Shipped since: CSV export and copy-as-curl, both extending the existing
export system by dispatching on the file extension the user types
rather than adding a separate format-selection control - ".har" (the
existing default) or ".csv" for bulk export from the history list,
".txt" (the existing default) or ".sh"/".curl" for single-entry export
from Detail view. CSV is deliberately a lighter, faster path than HAR:
a summary table (id/method/host/path/status/sizes/timing/flag/source)
built straight from the already-loaded Summary rows, no per-entry fetch
from the daemon needed, matching Burp's own "export as CSV" being a
listing rather than a full-fidelity capture - HAR already covers that.
Copy-as-curl parses the raw request (same net/http parsing used
everywhere else in this codebase) and re-serializes it as a runnable
curl command line rather than another copy of the raw bytes, which the
plain-text format already gives you; verified by actually executing a
generated command against the real target and confirming the response
matched the original.

CSV export required one more thing HAR didn't: guarding against CSV/
formula injection. Method, host, path, and error all ultimately trace
back to a request line or Host header - content this tool exists
specifically to inspect from potentially hostile traffic - and a field
starting with =, +, -, @, tab, or CR is a formula to Excel/LibreOffice/
Sheets when the exported file is later opened. csvSafe prefixes any
such field with a single quote, the standard mitigation (OWASP's own
guidance), so every affected spreadsheet application treats it as
literal text instead. This is the same class of bug as the terminal-
injection fix from the robustness audit, just for a different output
format - captured content controlling the tool that later processes it,
rather than the terminal that renders it.

Still open from the expanded "worth considering" list: import (no path
back in yet). Requested explicitly; not started yet.

Skipped deliberately (from the research, matches this tool's stated
scope): active/passive vulnerability scanning, plugin marketplace,
Collaborator/OAST, team collaboration, CI integration, client TLS
(mutual-TLS) certs, invisible/non-proxy-aware proxying.