一. 前言:bytes 不是亂碼,是資料有自己的排隊方式 #
Python 開發時,我們很常把資料想成文字:JSON、TOML、CSV、Markdown、log,一打開就能讀。 但很多資料不是文字,而是一段 bytes。 圖片、音訊、網路封包、感測器輸出、遊戲存檔、C 程式寫出的 record,都可能長得像這樣:
b'PYPY\x01\x00\x03\x00*\x00\x00\x00'
看起來像亂碼,其實只是它的規則不靠逗號和引號,而是靠「第幾個 byte 代表什麼」。
今天拍拍君要介紹標準庫 struct。
它可以把 Python 數字打包成 bytes,也可以把 bytes 解析回 Python 數字。
你會用它處理固定長度 header、二進位 record、endianness,最後做一個迷你 .pypybin 檔案格式。
這篇不是叫你所有資料都改用 binary。
JSON 很好,人類看得懂。
但遇到 binary protocol、舊系統資料或低階檔案格式時,struct 是那把小螺絲起子。
二. 安裝:標準庫內建,先建立練習專案 #
struct 是 Python 標準庫,不需要安裝。
我們建立一個乾淨練習目錄:
uv init struct-lab
cd struct-lab
如果你不用 uv,直接跑也可以:
python basic_pack.py
重點不是環境工具,而是理解 bytes 的布局。 一旦布局清楚,binary data 就不神秘。
三. 第一個 pack:把數字塞進 bytes
#
先建立 basic_pack.py:
from __future__ import annotations
import struct
payload = struct.pack("I", 42)
print(payload)
print(payload.hex())
print(len(payload))
執行後可能看到:
b'*\x00\x00\x00'
2a000000
4
struct.pack("I", 42) 的意思是:用 unsigned int 的格式,把 42 打包成 bytes。
"I" 通常是 4-byte unsigned integer。
所以 42 不是存成文字 "42",而是存成一個固定 4 bytes 的整數。
文字格式比較直覺,binary layout 則提供固定寬度。
固定的代價是:你必須知道規格。
沒有規格時,bytes 像謎語;有規格時,它只是資料在排隊。
四. 第一個 unpack:把 bytes 讀回數字
#
打包之後,當然要能拆回來:
from __future__ import annotations
import struct
payload = struct.pack("I", 42)
value = struct.unpack("I", payload)
print(value)
結果是:
(42,)
struct.unpack() 永遠回傳 tuple。
就算只有一個欄位,也是 (42,)。
通常會這樣拆:
(count,) = struct.unpack("I", payload)
print(count)
多個欄位也可以一次處理:
payload = struct.pack("I f", 42, 3.5)
count, score = struct.unpack("I f", payload)
print(count)
print(score)
格式字串 "I f" 表示第一個欄位是 unsigned int,第二個欄位是 float。
空白只是讓人好讀,"If" 也可以。
拍拍君偏好在教學和複雜格式裡加空白,未來的你會比較不想翻桌。
五. Format string:struct 的小型資料布局語言
#
struct 的核心是 format string。
常用格式大概是這些:
| 格式 | 意義 | 常見大小 |
|---|---|---|
b |
signed char | 1 byte |
B |
unsigned char | 1 byte |
h |
signed short | 2 bytes |
H |
unsigned short | 2 bytes |
i |
signed int | 4 bytes |
I |
unsigned int | 4 bytes |
q |
signed long long | 8 bytes |
Q |
unsigned long long | 8 bytes |
f |
float | 4 bytes |
d |
double | 8 bytes |
s |
fixed bytes | 固定長度 |
x |
padding byte | 1 byte |
你可以用 calcsize() 看格式需要多少 bytes: |
import struct
print(struct.calcsize("I"))
print(struct.calcsize("I f"))
print(struct.calcsize("4s I f"))
"4s I f" 表示 4-byte bytes 欄位、unsigned int、float。
注意 4s 和 4B 不一樣:
print(struct.unpack("4s", b"PYPY"))
print(struct.unpack("4B", b"PYPY"))
結果:
(b'PYPY',)
(80, 89, 80, 89)
4s 是一個 4-byte 欄位。
4B 是四個 1-byte 整數。
這個差別很小,但 debug binary format 時很重要。
六. Endianness:同一個數字,byte 順序可以不同 #
來看一個經典坑:
import struct
number = 0x12345678
little = struct.pack("<I", number)
big = struct.pack(">I", number)
print(little.hex())
print(big.hex())
結果:
78563412
12345678
同一個整數,byte 順序可以不同。 little endian 把低位 byte 放前面;big endian 把高位 byte 放前面。 format string 前面可以加 prefix:
| Prefix | 意義 |
|---|---|
@ |
native layout |
= |
native endian, standard size |
< |
little endian |
> |
big endian |
! |
network byte order,也就是 big endian |
如果你在設計自己的檔案格式或 protocol,請明確寫 < 或 >。 |
|
| 不要讓「本機預設」混進規格。 | |
| 例如: |
HEADER_FORMAT = "<4s H I"
這表示 little endian、4-byte magic、2-byte unsigned version、4-byte unsigned count。 只要規格清楚,不同機器解析結果就一致。
七. 固定長度 Header:設計一個可讀的 binary 開頭 #
很多 binary 檔案一開始都有 header。
header 通常回答這幾件事:這是什麼格式、版本是多少、後面有多少資料。
建立 header.py:
from __future__ import annotations
from dataclasses import dataclass
import struct
MAGIC = b"PYPY"
HEADER = struct.Struct("<4s H I")
@dataclass(frozen=True)
class Header:
version: int
record_count: int
def pack_header(header: Header) -> bytes:
return HEADER.pack(MAGIC, header.version, header.record_count)
def unpack_header(data: bytes) -> Header:
if len(data) != HEADER.size:
raise ValueError(f"header must be {HEADER.size} bytes")
magic, version, record_count = HEADER.unpack(data)
if magic != MAGIC:
raise ValueError(f"invalid magic: {magic!r}")
return Header(version=version, record_count=record_count)
試跑:
header = Header(version=1, record_count=3)
payload = pack_header(header)
print(payload.hex())
print(unpack_header(payload))
讀檔時也很直接:
with open("data.pypybin", "rb") as f:
header_bytes = f.read(HEADER.size)
header = unpack_header(header_bytes)
Struct 會預先保存格式,也直接提供 .size。
同一個 layout 會重複使用時,比到處傳 format string 更不容易寫錯。
規格說 header 幾個 bytes,就讀幾個 bytes。
這種清楚感,是 binary format 最可愛的地方之一。
八. Record Layout:把一筆資料固定成幾個欄位 #
接著設計 record。 假設我們要存一批任務評分資料:
task_id:unsigned intduration_ms:unsigned intscore:float 程式可以這樣寫:
from __future__ import annotations
from dataclasses import dataclass
import struct
RECORD = struct.Struct("<I I f")
@dataclass(frozen=True)
class TaskRecord:
task_id: int
duration_ms: int
score: float
def pack_record(record: TaskRecord) -> bytes:
return RECORD.pack(
record.task_id,
record.duration_ms,
record.score,
)
def unpack_record(data: bytes) -> TaskRecord:
if len(data) != RECORD.size:
raise ValueError(f"record must be {RECORD.size} bytes")
task_id, duration_ms, score = RECORD.unpack(data)
return TaskRecord(task_id=task_id, duration_ms=duration_ms, score=score)
試試看:
record = TaskRecord(task_id=7, duration_ms=1250, score=0.95)
payload = pack_record(record)
print(payload.hex())
print(unpack_record(payload))
你可能會看到 float 有一點點誤差,例如 0.949999988079071。
這正常。
f 是 32-bit float,不是十進位精準小數。
存金額請不要用 float,拍拍君會皺眉。
九. 做一個迷你 .pypybin 檔案格式
#
把 header 和 records 合起來,建立 tiny_format.py:
from __future__ import annotations
from dataclasses import dataclass
from pathlib import Path
import struct
MAGIC = b"PYPY"
VERSION = 1
MAX_RECORDS = 100_000
HEADER = struct.Struct("<4s H I")
RECORD = struct.Struct("<I I f")
@dataclass(frozen=True)
class TaskRecord:
task_id: int
duration_ms: int
score: float
def write_records(path: Path, records: list[TaskRecord]) -> None:
if len(records) > MAX_RECORDS:
raise ValueError("too many records")
with path.open("wb") as f:
f.write(HEADER.pack(MAGIC, VERSION, len(records)))
for record in records:
f.write(RECORD.pack(record.task_id, record.duration_ms, record.score))
def read_records(path: Path) -> list[TaskRecord]:
with path.open("rb") as f:
header_bytes = f.read(HEADER.size)
if len(header_bytes) != HEADER.size:
raise ValueError("file is too small to contain header")
magic, version, record_count = HEADER.unpack(header_bytes)
if magic != MAGIC:
raise ValueError("not a PYPY binary file")
if version != VERSION:
raise ValueError(f"unsupported version: {version}")
if record_count > MAX_RECORDS:
raise ValueError("record count exceeds safety limit")
records: list[TaskRecord] = []
for _ in range(record_count):
record_bytes = f.read(RECORD.size)
if len(record_bytes) != RECORD.size:
raise ValueError("file ended before all records were read")
task_id, duration_ms, score = RECORD.unpack(record_bytes)
records.append(TaskRecord(task_id=task_id, duration_ms=duration_ms, score=score))
if f.read(1):
raise ValueError("file has trailing bytes")
return records
使用方式:
from pathlib import Path
from tiny_format import TaskRecord, read_records, write_records
path = Path("tasks.pypybin")
write_records(
path,
[
TaskRecord(task_id=1, duration_ms=120, score=0.82),
TaskRecord(task_id=2, duration_ms=980, score=0.91),
],
)
print(read_records(path))
這個格式很小,但概念完整:magic number、version、record count、fixed-size records、數量上限與 trailing bytes check。 實務上你可能還會加 checksum、compression flag、created timestamp 或 variable-length section。 但核心就是這樣。
十. unpack_from 與 iter_unpack:處理較大的 buffer
#
unpack() 要求輸入長度剛好符合格式。
如果 header 藏在較大的 buffer 裡,可以指定 offset:
packet = b"META" + HEADER.pack(MAGIC, 1, 2) + b"payload"
magic, version, count = HEADER.unpack_from(packet, offset=4)
如果整個 buffer 都是相同 record,iter_unpack() 會逐筆解析:
payload = b"".join(
RECORD.pack(i, i * 100, 0.5)
for i in range(3)
)
for task_id, duration_ms, score in RECORD.iter_unpack(payload):
print(task_id, duration_ms, score)
buffer 長度必須是 record size 的整數倍;不完整的尾巴會直接觸發 struct.error。
這比自己手算每個 slice 清楚,也適合搭配 mmap 讀固定長度資料。
十一. 用 pytest 保護格式合約 #
binary parser 最怕「看起來能讀」,遇到截斷或錯版本才爆炸。 至少測 round trip 與壞資料:
from pathlib import Path
import pytest
from tiny_format import TaskRecord, read_records, write_records
def test_round_trip(tmp_path: Path) -> None:
path = tmp_path / "tasks.pypybin"
expected = [TaskRecord(1, 120, 0.5), TaskRecord(2, 980, 0.75)]
write_records(path, expected)
assert read_records(path) == expected
def test_rejects_truncated_file(tmp_path: Path) -> None:
path = tmp_path / "broken.pypybin"
path.write_bytes(b"PYPY\x01")
with pytest.raises(ValueError, match="too small"):
read_records(path)
測試不是裝飾。magic、version、長度與數量上限,都是檔案格式的公開合約。
十二. 常見踩雷整理 #
第一個坑:沒指定 endian。
struct.pack("I", value)
如果資料要跨平台或長期保存,請改成:
struct.pack("<I", value)
第二個坑:把 s 和 B 搞混。
struct.unpack("4s", b"PYPY") # 一個 bytes 欄位
struct.unpack("4B", b"PYPY") # 四個整數欄位
第三個坑:float 精準度。
f 是 32-bit float,d 是 64-bit double。
第四個坑:相信檔案裡宣告的長度。
外部資料可能損壞或惡意,解析前要限制 count、payload size 與支援的 version。
第五個坑:忘記版本演進。
binary format 一旦發出去,就像 API。
不是不能改,是要有遷移策略。
最後,如果你只是在處理單一整數,也可以考慮 int.to_bytes() / int.from_bytes()。
但只要同時有多個欄位、固定 layout、float 或 C-like record,struct 會比較清楚。
結語 #
今天我們用 struct 做了幾件事:
- 用
pack()把 Python 數值打包成 bytes - 用
unpack()從 bytes 讀回欄位 - 用 format string 描述 binary layout
- 明確指定 little endian / big endian
- 設計固定長度 header
- 寫入和讀取 fixed-size records
- 用
unpack_from()與iter_unpack()解析 buffer - 做了一個迷你
.pypybin檔案格式 - 用測試保護 parser 行為
拍拍君覺得
struct的價值不只在「可以處理 binary」。 更重要的是,它會逼你思考資料的形狀:每個欄位幾 bytes、signed 還是 unsigned、byte order 是什麼、版本如何演進、壞資料要怎麼拒絕。 這些問題在高階工具裡常常被包起來。 但只要你做過一次,就會更懂網路、檔案格式、資料庫頁面和序列化工具在處理什麼。 下次你看到一串b'\x00\x01\x02',不要急著說它是亂碼。 它可能只是還沒被正確介紹給你。