Learn
HTTP/18-build-your-own

实战:手写极简 HTTP 服务器与客户端

毕业考试:不用任何 HTTP 库,只用 socket 和字符串处理,写出能被 curl 正常访问的服务器和能与标准服务器对话的客户端。写完这两个程序,前 17 章的知识就真正长在你手上了。

1. 服务器要做的五件事

回顾第 4 章的报文结构,一个最小可用的 HTTP/1.1 服务器只需:

  1. bind + listen + accept(第 3 章的三次握手由内核代劳);
  2. 读取字节直到 \r\n\r\n(头部结束标志);
  3. 解析请求行(方法、路径、版本)与头部;
  4. 路由到处理函数,生成状态行 + 头 + 体;
  5. 按 Connection 头决定复用还是关闭连接(第 11 章)。

2. 极简服务器:50 行核心代码

手写 HTTP 服务器,curl 直接访问
python3 -c '
import socket, threading, json
 
def handle(conn):
    try:
        # ── 1. 读到头部结束 ──────────────────────────
        buf = b""
        while b"\r\n\r\n" not in buf:
            chunk = conn.recv(4096)
            if not chunk:
                return
            buf += chunk
        head, _, rest = buf.partition(b"\r\n\r\n")
        lines = head.decode("iso-8859-1").split("\r\n")
 
        # ── 2. 解析请求行与头部 ──────────────────────
        method, path, version = lines[0].split(" ", 2)
        headers = {}
        for line in lines[1:]:
            k, _, v = line.partition(":")
            headers[k.strip().lower()] = v.strip()
 
        # ── 3. 按 Content-Length 读完请求体 ──────────
        need = int(headers.get("content-length", 0))
        body = rest
        while len(body) < need:
            body += conn.recv(4096)
 
        # ── 4. 路由 ─────────────────────────────────
        if path == "/hello":
            status, ctype, out = "200 OK", "text/plain; charset=utf-8", b"Hello, HTTP!\n"
        elif path == "/headers":
            out = json.dumps(headers, ensure_ascii=False, indent=1).encode()
            status, ctype = "200 OK", "application/json"
        elif path == "/echo" and method == "POST":
            status, ctype, out = "200 OK", "application/octet-stream", body
        else:
            status, ctype, out = "404 Not Found", "text/plain", b"no such route\n"
 
        # ── 5. 拼响应报文 ────────────────────────────
        resp = ("HTTP/1.1 " + status + "\r\n"
                "Content-Type: " + ctype + "\r\n"
                "Content-Length: " + str(len(out)) + "\r\n"
                "Server: handmade/0.1\r\n"
                "Connection: close\r\n\r\n").encode() + out
        conn.sendall(resp)
    finally:
        conn.close()
 
def serve():
    srv = socket.socket()
    srv.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
    srv.bind(("127.0.0.1", 8000))
    srv.listen(16)
    while True:
        c, _ = srv.accept()
        threading.Thread(target=handle, args=(c,), daemon=True).start()
 
threading.Thread(target=serve, daemon=True).start()
import time; time.sleep(0.3)
 
import subprocess
subprocess.run(["curl", "-s", "http://127.0.0.1:8000/hello"], check=False)
subprocess.run(["curl", "-s", "http://127.0.0.1:8000/headers", "-H", "X-Course: http"], check=False)
print()
subprocess.run(["curl", "-s", "-X", "POST", "http://127.0.0.1:8000/echo", "-d", "round-trip!"], check=False)
print()
subprocess.run(["curl", "-s", "-o", "/dev/null", "-w", "未知路径 -> %{http_code}\n",
                "http://127.0.0.1:8000/nope"], check=False)
'

几处对应前面章节的设计决策:

  • 用 \r\n\r\n 判断头部结束、用 Content-Length 读体(第 4 章的定界铁律);
  • 头部字段名统一转小写再查(大小写不敏感,第 4 章);
  • 每连接一个线程(第 11 章:真实服务器会用线程池/事件循环,思想相同);
  • 简化起见回 Connection: close(练习里会让你升级成 keep-alive)。

3. 手写客户端

客户端的难点在解析响应:读状态行、逐行读头、再按 Content-Length 精确读体。这次对面换成标准的 python 文件服务器,验证我们的客户端"说的是标准 HTTP":

手写 HTTP 客户端,访问标准服务器
echo "served by stdlib server" > data.txt
python3 -m http.server 8000 --bind 127.0.0.1 >/dev/null 2>&1 &
for _i in $(seq 1 50); do (exec 3<>/dev/tcp/127.0.0.1/8000) 2>/dev/null && break; sleep 0.1; done
 
python3 -c '
import socket
 
def http_get(host, port, path):
    s = socket.create_connection((host, port), timeout=5)
    req = ("GET " + path + " HTTP/1.1\r\n"
           "Host: " + host + "\r\n"
           "User-Agent: handmade-client/0.1\r\n"
           "Connection: close\r\n\r\n")
    s.sendall(req.encode())
 
    # 收齐头部
    buf = b""
    while b"\r\n\r\n" not in buf:
        chunk = s.recv(4096)
        if not chunk:
            break
        buf += chunk
    head, _, body = buf.partition(b"\r\n\r\n")
    lines = head.decode("iso-8859-1").split("\r\n")
 
    version, code, reason = lines[0].split(" ", 2)
    headers = {}
    for line in lines[1:]:
        k, _, v = line.partition(":")
        headers[k.strip().lower()] = v.strip()
 
    # 按 Content-Length 读完剩余的体
    need = int(headers.get("content-length", 0))
    while len(body) < need:
        chunk = s.recv(4096)
        if not chunk:
            break
        body += chunk
    s.close()
    return int(code), reason, headers, body
 
code, reason, headers, body = http_get("127.0.0.1", 8000, "/data.txt")
print("状态:", code, reason)
print("Content-Type:", headers.get("content-type"))
print("Content-Length:", headers.get("content-length"), "实收:", len(body))
print("体:", body.decode().strip())
 
code, reason, _, _ = http_get("127.0.0.1", 8000, "/missing.txt")
print("再试一个不存在的文件:", code, reason)
'

一个健壮的真实客户端还要处理:chunked 解码(第 10 章)、gzip 解压、3xx 跟随(第 6 章)、keep-alive 连接池(第 11 章)、TLS(第 12 章)——每一项你都已经知道原理了。

4. 终极联调:自己的客户端访问自己的服务器

毕业合影:手写客户端 x 手写服务器
python3 -c '
import socket, threading, time
 
# ── 迷你服务器:只有一个路由,但支持 keep-alive ──
def handle(conn):
    while True:                          # 循环处理同一连接上的多个请求
        buf = b""
        while b"\r\n\r\n" not in buf:
            chunk = conn.recv(4096)
            if not chunk:
                conn.close(); return
            buf += chunk
        req_line = buf.split(b"\r\n", 1)[0].decode()
        out = ("你请求了: " + req_line + "\n").encode()
        conn.sendall(("HTTP/1.1 200 OK\r\n"
                      "Content-Length: " + str(len(out)) + "\r\n"
                      "\r\n").encode() + out)   # 不发 Connection: close -> 保持连接
 
def serve():
    srv = socket.socket()
    srv.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
    srv.bind(("127.0.0.1", 8000)); srv.listen(4)
    while True:
        c, _ = srv.accept()
        threading.Thread(target=handle, args=(c,), daemon=True).start()
 
threading.Thread(target=serve, daemon=True).start()
time.sleep(0.2)
 
# ── 迷你客户端:一条连接串行发 3 个请求(验证 keep-alive)──
s = socket.create_connection(("127.0.0.1", 8000))
for path in ("/first", "/second", "/third"):
    s.sendall(("GET " + path + " HTTP/1.1\r\nHost: local\r\n\r\n").encode())
    buf = b""
    while b"\r\n\r\n" not in buf:
        buf += s.recv(4096)
    head, _, body = buf.partition(b"\r\n\r\n")
    need = 0
    for line in head.split(b"\r\n"):
        if line.lower().startswith(b"content-length"):
            need = int(line.split(b":")[1])
    while len(body) < need:
        body += s.recv(4096)
    print(body.decode().strip(), "| 复用同一条连接:", s.getsockname())
s.close()
'

三个请求、同一个本地端口号——一条 TCP 连接被复用了三次,我们徒手实现了 keep-alive。至此:分层(1)、报文(4)、方法(5)、状态码(6)、定界(4/10)、连接管理(11)全部亲手落地。

5. 从玩具到生产还差什么

维度玩具版生产级
并发模型每连接一线程epoll/kqueue 事件循环(Nginx)、协程(Go/asyncio)
健壮性信任输入头部大小限制、慢速攻击防护(slowloris)、超时全覆盖
协议完整度GET/POST + Content-Lengthchunked、100-continue、Range、TLS、h2/h3
安全无路径穿越校验(../../etc/passwd)、请求走私防御
⚠️玩具代码不要见公网

上面的服务器没有任何输入校验与资源限制,一个畸形报文或慢速连接就能让它难受。手写 HTTP 的价值是理解协议,生产环境请永远站在 Nginx/Caddy 与成熟框架的肩膀上。

💡下一步去哪

读一遍 RFC 9110(HTTP 语义)的目录会惊讶于它的可读性;抓一次真实站点的包(Wireshark/Chrome DevTools)对照本课每一章;或者给你的玩具服务器加上 chunked 与 Range 支持——协议能力是排障能力的上限。

小结

  • HTTP/1.1 服务器五步:accept → 读到空行 → 解析 → 路由 → 拼响应;
  • 客户端解析对称:状态行 → 头 → 按 Content-Length 收体;
  • keep-alive 的本质就是"响应完不关连接,循环再读下一个请求";
  • 定界(Content-Length / chunked)是手写实现中最容易出错也最重要的环节;
  • 玩具实现用于理解,生产依赖久经考验的服务器与框架。
🎯毕业练习
  1. 给 Playground 1 的服务器加上 keep-alive:解析 Connection 头,非 close 时循环处理同一连接的后续请求(参考 Playground 3)。
  2. 给服务器加一个静态文件路由:读取 /work 下的文件返回,并用第 7 章的知识按扩展名设置正确的 Content-Type;务必阻止路径中出现 .. 的穿越攻击。
  3. 给客户端加上 302 跟随:检测 3xx 状态码与 Location 头,最多跟随 5 次防止循环重定向。
  4. 综合题:给服务器加上 ETag 支持(对响应体取 md5),客户端第二次请求带 If-None-Match,验证收到 304 且无体——把第 9 章的协商缓存完整复刻一遍。