↓快轉到主要內容
  1. 教學文章/

Python struct 實戰:二進位資料打包、解析與檔案格式入門

·7 分鐘· loading · loading · ·
Python Struct Binary Standard-Library Developer-Tools
每日拍拍
作者
每日拍拍
科學家 X 科技宅宅
目錄
Python 學習 - 本文屬於一個選集。
§ 126: 本文

featured

一. 前言:bytes 不是亂碼,是資料有自己的排隊方式
#

Python 開發時,我們很常把資料想成文字:JSON、TOML、CSV、Markdown、log,一打開就能讀。 但很多資料不是文字,而是一段 bytes。 圖片、音訊、網路封包、感測器輸出、遊戲存檔、C 程式寫出的 record,都可能長得像這樣:

b'PYPY\x01\x00\x03\x00*\x00\x00\x00'

看起來像亂碼,其實只是它的規則不靠逗號和引號,而是靠「第幾個 byte 代表什麼」。 今天拍拍君要介紹標準庫 struct。 它可以把 Python 數字打包成 bytes,也可以把 bytes 解析回 Python 數字。 你會用它處理固定長度 header、二進位 record、endianness,最後做一個迷你 .pypybin 檔案格式。 這篇不是叫你所有資料都改用 binary。 JSON 很好,人類看得懂。 但遇到 binary protocol、舊系統資料或低階檔案格式時,struct 是那把小螺絲起子。

二. 安裝:標準庫內建,先建立練習專案
#

struct 是 Python 標準庫,不需要安裝。 我們建立一個乾淨練習目錄:

uv init struct-lab
cd struct-lab

如果你不用 uv,直接跑也可以:

python basic_pack.py

重點不是環境工具,而是理解 bytes 的布局。 一旦布局清楚,binary data 就不神秘。

三. 第一個 pack:把數字塞進 bytes
#

先建立 basic_pack.py:

from __future__ import annotations
import struct

payload = struct.pack("I", 42)
print(payload)
print(payload.hex())
print(len(payload))

執行後可能看到:

b'*\x00\x00\x00'
2a000000
4

struct.pack("I", 42) 的意思是:用 unsigned int 的格式,把 42 打包成 bytes。 "I" 通常是 4-byte unsigned integer。 所以 42 不是存成文字 "42",而是存成一個固定 4 bytes 的整數。 文字格式比較直覺,binary layout 則提供固定寬度。 固定的代價是:你必須知道規格。 沒有規格時,bytes 像謎語;有規格時,它只是資料在排隊。

四. 第一個 unpack:把 bytes 讀回數字
#

打包之後,當然要能拆回來:

from __future__ import annotations
import struct

payload = struct.pack("I", 42)
value = struct.unpack("I", payload)
print(value)

結果是:

(42,)

struct.unpack() 永遠回傳 tuple。 就算只有一個欄位,也是 (42,)。 通常會這樣拆:

(count,) = struct.unpack("I", payload)
print(count)

多個欄位也可以一次處理:

payload = struct.pack("I f", 42, 3.5)
count, score = struct.unpack("I f", payload)
print(count)
print(score)

格式字串 "I f" 表示第一個欄位是 unsigned int,第二個欄位是 float。 空白只是讓人好讀,"If" 也可以。 拍拍君偏好在教學和複雜格式裡加空白,未來的你會比較不想翻桌。

五. Format string:struct 的小型資料布局語言
#

struct 的核心是 format string。 常用格式大概是這些:

格式 意義 常見大小
b signed char 1 byte
B unsigned char 1 byte
h signed short 2 bytes
H unsigned short 2 bytes
i signed int 4 bytes
I unsigned int 4 bytes
q signed long long 8 bytes
Q unsigned long long 8 bytes
f float 4 bytes
d double 8 bytes
s fixed bytes 固定長度
x padding byte 1 byte
你可以用 calcsize() 看格式需要多少 bytes:
import struct
print(struct.calcsize("I"))
print(struct.calcsize("I f"))
print(struct.calcsize("4s I f"))

"4s I f" 表示 4-byte bytes 欄位、unsigned int、float。 注意 4s 和 4B 不一樣:

print(struct.unpack("4s", b"PYPY"))
print(struct.unpack("4B", b"PYPY"))

結果:

(b'PYPY',)
(80, 89, 80, 89)

4s 是一個 4-byte 欄位。 4B 是四個 1-byte 整數。 這個差別很小,但 debug binary format 時很重要。

六. Endianness:同一個數字,byte 順序可以不同
#

來看一個經典坑:

import struct
number = 0x12345678
little = struct.pack("<I", number)
big = struct.pack(">I", number)
print(little.hex())
print(big.hex())

結果:

78563412
12345678

同一個整數,byte 順序可以不同。 little endian 把低位 byte 放前面;big endian 把高位 byte 放前面。 format string 前面可以加 prefix:

Prefix 意義
@ native layout
= native endian, standard size
< little endian
> big endian
! network byte order,也就是 big endian
如果你在設計自己的檔案格式或 protocol,請明確寫 < 或 >。
不要讓「本機預設」混進規格。
例如:
HEADER_FORMAT = "<4s H I"

這表示 little endian、4-byte magic、2-byte unsigned version、4-byte unsigned count。 只要規格清楚,不同機器解析結果就一致。

七. 固定長度 Header:設計一個可讀的 binary 開頭
#

很多 binary 檔案一開始都有 header。 header 通常回答這幾件事:這是什麼格式、版本是多少、後面有多少資料。 建立 header.py:

from __future__ import annotations
from dataclasses import dataclass
import struct

MAGIC = b"PYPY"
HEADER = struct.Struct("<4s H I")

@dataclass(frozen=True)
class Header:
    version: int
    record_count: int

def pack_header(header: Header) -> bytes:
    return HEADER.pack(MAGIC, header.version, header.record_count)

def unpack_header(data: bytes) -> Header:
    if len(data) != HEADER.size:
        raise ValueError(f"header must be {HEADER.size} bytes")
    magic, version, record_count = HEADER.unpack(data)
    if magic != MAGIC:
        raise ValueError(f"invalid magic: {magic!r}")
    return Header(version=version, record_count=record_count)

試跑:

header = Header(version=1, record_count=3)
payload = pack_header(header)
print(payload.hex())
print(unpack_header(payload))

讀檔時也很直接:

with open("data.pypybin", "rb") as f:
    header_bytes = f.read(HEADER.size)
    header = unpack_header(header_bytes)

Struct 會預先保存格式,也直接提供 .size。 同一個 layout 會重複使用時,比到處傳 format string 更不容易寫錯。 規格說 header 幾個 bytes,就讀幾個 bytes。 這種清楚感,是 binary format 最可愛的地方之一。

八. Record Layout:把一筆資料固定成幾個欄位
#

接著設計 record。 假設我們要存一批任務評分資料:

  • task_id:unsigned int
  • duration_ms:unsigned int
  • score:float 程式可以這樣寫:
from __future__ import annotations
from dataclasses import dataclass
import struct

RECORD = struct.Struct("<I I f")

@dataclass(frozen=True)
class TaskRecord:
    task_id: int
    duration_ms: int
    score: float

def pack_record(record: TaskRecord) -> bytes:
    return RECORD.pack(
        record.task_id,
        record.duration_ms,
        record.score,
    )

def unpack_record(data: bytes) -> TaskRecord:
    if len(data) != RECORD.size:
        raise ValueError(f"record must be {RECORD.size} bytes")
    task_id, duration_ms, score = RECORD.unpack(data)
    return TaskRecord(task_id=task_id, duration_ms=duration_ms, score=score)

試試看:

record = TaskRecord(task_id=7, duration_ms=1250, score=0.95)
payload = pack_record(record)
print(payload.hex())
print(unpack_record(payload))

你可能會看到 float 有一點點誤差,例如 0.949999988079071。 這正常。 f 是 32-bit float,不是十進位精準小數。 存金額請不要用 float,拍拍君會皺眉。

九. 做一個迷你 .pypybin 檔案格式
#

把 header 和 records 合起來,建立 tiny_format.py:

from __future__ import annotations
from dataclasses import dataclass
from pathlib import Path
import struct

MAGIC = b"PYPY"
VERSION = 1
MAX_RECORDS = 100_000
HEADER = struct.Struct("<4s H I")
RECORD = struct.Struct("<I I f")

@dataclass(frozen=True)
class TaskRecord:
    task_id: int
    duration_ms: int
    score: float

def write_records(path: Path, records: list[TaskRecord]) -> None:
    if len(records) > MAX_RECORDS:
        raise ValueError("too many records")
    with path.open("wb") as f:
        f.write(HEADER.pack(MAGIC, VERSION, len(records)))
        for record in records:
            f.write(RECORD.pack(record.task_id, record.duration_ms, record.score))

def read_records(path: Path) -> list[TaskRecord]:
    with path.open("rb") as f:
        header_bytes = f.read(HEADER.size)
        if len(header_bytes) != HEADER.size:
            raise ValueError("file is too small to contain header")
        magic, version, record_count = HEADER.unpack(header_bytes)
        if magic != MAGIC:
            raise ValueError("not a PYPY binary file")
        if version != VERSION:
            raise ValueError(f"unsupported version: {version}")
        if record_count > MAX_RECORDS:
            raise ValueError("record count exceeds safety limit")
        records: list[TaskRecord] = []
        for _ in range(record_count):
            record_bytes = f.read(RECORD.size)
            if len(record_bytes) != RECORD.size:
                raise ValueError("file ended before all records were read")
            task_id, duration_ms, score = RECORD.unpack(record_bytes)
            records.append(TaskRecord(task_id=task_id, duration_ms=duration_ms, score=score))
        if f.read(1):
            raise ValueError("file has trailing bytes")
    return records

使用方式:

from pathlib import Path
from tiny_format import TaskRecord, read_records, write_records

path = Path("tasks.pypybin")
write_records(
    path,
    [
        TaskRecord(task_id=1, duration_ms=120, score=0.82),
        TaskRecord(task_id=2, duration_ms=980, score=0.91),
    ],
)
print(read_records(path))

這個格式很小,但概念完整:magic number、version、record count、fixed-size records、數量上限與 trailing bytes check。 實務上你可能還會加 checksum、compression flag、created timestamp 或 variable-length section。 但核心就是這樣。

十. unpack_from 與 iter_unpack:處理較大的 buffer
#

unpack() 要求輸入長度剛好符合格式。 如果 header 藏在較大的 buffer 裡,可以指定 offset:

packet = b"META" + HEADER.pack(MAGIC, 1, 2) + b"payload"
magic, version, count = HEADER.unpack_from(packet, offset=4)

如果整個 buffer 都是相同 record,iter_unpack() 會逐筆解析:

payload = b"".join(
    RECORD.pack(i, i * 100, 0.5)
    for i in range(3)
)
for task_id, duration_ms, score in RECORD.iter_unpack(payload):
    print(task_id, duration_ms, score)

buffer 長度必須是 record size 的整數倍;不完整的尾巴會直接觸發 struct.error。 這比自己手算每個 slice 清楚,也適合搭配 mmap 讀固定長度資料。

十一. 用 pytest 保護格式合約
#

binary parser 最怕「看起來能讀」,遇到截斷或錯版本才爆炸。 至少測 round trip 與壞資料:

from pathlib import Path
import pytest
from tiny_format import TaskRecord, read_records, write_records

def test_round_trip(tmp_path: Path) -> None:
    path = tmp_path / "tasks.pypybin"
    expected = [TaskRecord(1, 120, 0.5), TaskRecord(2, 980, 0.75)]
    write_records(path, expected)
    assert read_records(path) == expected

def test_rejects_truncated_file(tmp_path: Path) -> None:
    path = tmp_path / "broken.pypybin"
    path.write_bytes(b"PYPY\x01")
    with pytest.raises(ValueError, match="too small"):
        read_records(path)

測試不是裝飾。magic、version、長度與數量上限,都是檔案格式的公開合約。

十二. 常見踩雷整理
#

第一個坑:沒指定 endian。

struct.pack("I", value)

如果資料要跨平台或長期保存,請改成:

struct.pack("<I", value)

第二個坑:把 s 和 B 搞混。

struct.unpack("4s", b"PYPY")  # 一個 bytes 欄位
struct.unpack("4B", b"PYPY")  # 四個整數欄位

第三個坑:float 精準度。 f 是 32-bit float,d 是 64-bit double。 第四個坑:相信檔案裡宣告的長度。 外部資料可能損壞或惡意,解析前要限制 count、payload size 與支援的 version。 第五個坑:忘記版本演進。 binary format 一旦發出去,就像 API。 不是不能改,是要有遷移策略。 最後,如果你只是在處理單一整數,也可以考慮 int.to_bytes() / int.from_bytes()。 但只要同時有多個欄位、固定 layout、float 或 C-like record,struct 會比較清楚。

結語
#

今天我們用 struct 做了幾件事:

  • 用 pack() 把 Python 數值打包成 bytes
  • 用 unpack() 從 bytes 讀回欄位
  • 用 format string 描述 binary layout
  • 明確指定 little endian / big endian
  • 設計固定長度 header
  • 寫入和讀取 fixed-size records
  • 用 unpack_from() 與 iter_unpack() 解析 buffer
  • 做了一個迷你 .pypybin 檔案格式
  • 用測試保護 parser 行為 拍拍君覺得 struct 的價值不只在「可以處理 binary」。 更重要的是,它會逼你思考資料的形狀:每個欄位幾 bytes、signed 還是 unsigned、byte order 是什麼、版本如何演進、壞資料要怎麼拒絕。 這些問題在高階工具裡常常被包起來。 但只要你做過一次,就會更懂網路、檔案格式、資料庫頁面和序列化工具在處理什麼。 下次你看到一串 b'\x00\x01\x02',不要急著說它是亂碼。 它可能只是還沒被正確介紹給你。

延伸閱讀
#

Python 學習 - 本文屬於一個選集。
§ 126: 本文

相關文章

Python socket 實戰:TCP client/server、timeout 與簡易通訊協定
·8 分鐘· loading · loading
Python Socket TCP Networking Standard-Library Developer-Tools
Python shutil 實戰:檔案複製、搬移、壓縮與安全清理
·7 分鐘· loading · loading
Python Shutil Filesystem Automation Standard-Library Developer-Tools
Python inspect 實戰:看懂函式簽名、物件結構與開發工具自動化
·6 分鐘· loading · loading
Python Inspect Introspection Standard-Library Developer-Tools
Python JSON 實戰:解析、序列化、自訂型別與大型資料
·7 分鐘· loading · loading
Python Json Serialization Standard-Library JSON Lines Data-Engineering Developer-Tools
Python csv 實戰:DictReader、Dialect 與串流清理資料
·7 分鐘· loading · loading
Python CSV Data-Cleaning Standard-Library ETL Developer-Tools
Python sysconfig 實戰:安裝路徑、編譯資訊與環境診斷
·6 分鐘· loading · loading
Python Sysconfig Standard-Library Packaging Virtualenv Developer-Tools