Skip to content

pcap 文件格式深度解析 ​

这篇用实际抓到的 pcap,逐字节拆解给你看 r0capture 写出的到底是什么。

实验样本 ​

用 --selftest 抓 example.com:

bash
python3 r0capture.py -H 127.0.0.1:27042 7033 --selftest -p selftest.pcap -v

得到 selftest.pcap(3552 字节)。接下来逐层拆解。

全局头(24 字节) ​

python
import struct
with open('selftest.pcap','rb') as f:
    data = f.read()

magic, ver_maj, ver_min, thiszone, sigfigs, snaplen, linktype = \
    struct.unpack('<IHHiIII', data[:24])
字段值含义
magic0xA1B2C3D4pcap 魔法数,标识这是 pcap 文件(小端)
version2.4pcap 格式版本
thiszone-28800时区修正(秒),-28800 = UTC+8
sigfigs0时间戳精度,通常为 0
snaplen65535单个包最大捕获长度
linktype228LINKTYPE_IPV4:每条记录直接是 IPv4 包

为什么是 228 不是 1

linktype=1 是 LINKTYPE_ETHERNET,每条记录带 14 字节以太网帧头。linktype=228 是 LINKTYPE_IPV4,直接从 IP 头开始,没有以太网头。r0capture 不关心链路层,只关心 IP/TCP 之上的明文,所以用 228 更省事。

记录头(16 字节/条) ​

每条记录紧跟一个 16 字节头:

| 时间戳秒(4) | 时间戳微秒(4) | 包含长度(4) | 原始长度(4) |
  • 时间戳:抓这条记录的时刻
  • 包含长度:实际存了几个字节(可能被 snaplen 截断)
  • 原始长度:没截断前多大

一条记录的完整解剖 ​

selftest.pcap 的第 4 条记录(0.0.0.0:0,明文 GET 请求),逐字节:

Python 完整解析脚本 ​

python
import struct, gzip

with open('selftest.pcap', 'rb') as f:
    data = f.read()

off = 24  # 跳过全局头
while off + 16 <= len(data):
    # 记录头
    ts_sec, ts_usec, inc_len, orig_len = struct.unpack('<IIII', data[off:off+16])
    off += 16
    pkt = data[off:off+inc_len]
    off += inc_len

    # IPv4 头
    if (pkt[0] >> 4) != 4: continue
    ihl = (pkt[0] & 0xf) * 4
    src = '.'.join(str(b) for b in pkt[12:16])
    dst = '.'.join(str(b) for b in pkt[16:20])

    # TCP 头
    tcp_off = ihl
    sport, dport = struct.unpack('>HH', pkt[tcp_off:tcp_off+4])
    doff = ((pkt[tcp_off+12] >> 4) & 0xf) * 4
    payload = pkt[tcp_off + doff:]  # ← 这就是明文载荷

    print(f'{src}:{sport} -> {dst}:{dport} [{len(payload)}字节]')
    if payload.startswith(b'HTTP/'):
        # 响应:找body,解压gzip
        hdr_end = payload.find(b'\r\n\r\n')
        body = payload[hdr_end+4:]
        # chunked: 第一行是hex长度
        nl = body.find(b'\r\n')
        chunk_size = int(body[:nl], 16)
        chunk = body[nl+2:nl+2+chunk_size]
        html = gzip.decompress(chunk)  # 解压还原 HTML
        print(html.decode())

运行结果 ​

172.17.0.5:55252 -> 104.20.23.154:443 [520字节]   ← TLS握手(密文层)
104.20.23.154:443 -> 172.17.0.5:55252 [218字节]   ← TLS握手(密文层)
0.0.0.0:0 -> 0.0.0.0 [176字节]                    ← 明文请求 GET / HTTP/1.1
0.0.0.0:0 -> 0.0.0.0 [700字节]                    ← 明文响应 HTTP/1.1 200 OK
0.0.0.0:0 -> 0.0.0.0 [5字节]                      ← chunked结束 0\r\n\r\n

最后一段 gzip 解压得到 example.com 的 HTML:

html
<!doctype html>
<html lang="en"><head><title>Example Domain</title>...

从加密流量到明文 HTML,端到端还原成功。

用标准工具验证 ​

bash
$ tcpdump -r selftest.pcap -n
reading from file selftest.pcap, link-type IPV4 (Raw IPv4)
15:41:33.050626 IP 172.17.0.5.55252 > 104.20.23.154.443: Flags [P.], length 520
15:41:33.431246 IP 0.0.0.0.0 > 0.0.0.0: Flags [P.], length 176
...

link-type IPV4、seq/ack/flags 全对——pcap 格式完全标准,Wireshark 可直接打开。

下一篇:Frida 17 兼容性 →

基于 VitePress 构建 · 教学用途