โหมดมืด
Distributed Systems
ส่วนหนึ่งของ Beginner Book | เสริมจาก System Design
หนังสือเล่มนี้พาคุณไปถึงไหน
อ่านจบ + ทำ checkpoint หมด คุณจะ:
- เข้าใจปัญหาพื้นฐานของ distributed system (8 fallacies, failure models, two generals)
- เข้าใจ CAP (แค็ป), PACELC (เพ-เซล-ซี), consistency models — รูปแบบความสอดคล้องของข้อมูล (linearizable=ลิ-เนีย-ไรซ์-อะ-เบิล → eventual=อีเวนช่วล หลวมที่สุด) — ทุกศัพท์มีอธิบายตั้งแต่ศูนย์ในบท 1 และมี Glossary บท 11 เปิดควบคู่ได้
- รู้จัก consensus algorithm (Raft, Paxos, BFT) — ชื่อพวกนี้อธิบายตั้งแต่ศูนย์ในบท 2 ไม่ต้องรู้มาก่อน
- ออกแบบ event-driven + saga + outbox + CQRS + event sourcing
- จัดการ time, ordering (Lamport (แลม-พอร์ต), Vector clock = นาฬิกาเวกเตอร์, HLC = Hybrid Logical Clock นาฬิกาผสม, TrueTime ของ Google) — อธิบายตั้งแต่หลักการพื้นฐาน ไม่ต้อง pre-req
- เข้าใจ Service Discovery, Load Balancing, Rate Limiting, Circuit Breaker
- รู้จัก Service Mesh, CRDT (Conflict-free Replicated Data Type — โครงสร้างข้อมูลที่ merge แล้วไม่ขัดกัน), Edge Computing, Modern Patterns
หนังสือนี้ออกแบบมาให้ใคร
✅ Backend developer ที่ทำงาน microservices
✅ คนที่อยาก senior+ role
✅ คนที่อ่าน DDIA (Designing Data-Intensive Applications — หนังสือ reference ของวงการ distributed systems โดย Martin Kleppmann, O'Reilly) แล้วงง — เล่มนี้ช่วยปูพื้นบางส่วนเป็นภาษาไทยก่อน; DDIA ยังจำเป็นถ้าจะลึก senior+
✅ คนที่ออกแบบ multi-region system
❌ Researcher / distributed systems PhD (เล่มนี้ practical)
หมายเหตุเรื่องภาษา: เล่มนี้ "อธิบายเป็นไทย" และอ่านเข้าใจง่ายกว่า DDIA มาก แต่ศัพท์เทคนิคหลัก ๆ ยังคงเป็นภาษาอังกฤษ (เพราะเป็นคำที่ใช้ในวงการจริง) และมีกล่องโค้ด/ตัวอย่างบางส่วนเป็นอังกฤษ — คุณ ควรอ่านศัพท์เทคนิคภาษาอังกฤษได้ในระดับหนึ่ง ถ้าเจอคำย่อที่ไม่คุ้น เปิด Glossary บท 11 ควบคู่ไปได้เลย
Prerequisites — ควรรู้ก่อนเปิดเล่มนี้
| หัวข้อ | ระดับ | ที่หาความรู้เพิ่ม |
|---|---|---|
| HTTP, TCP/IP, DNS | พื้นฐาน | Networking |
| SQL + Relational DB | พื้นฐาน | Database book |
| อย่างน้อย 1 ภาษา backend | กลาง | Java/Go/Node/Python — ภาษาไหนก็ได้ |
| REST API, JSON | พื้นฐาน | — |
| Docker (option) | กลาง | สำหรับ checkpoint ที่ใช้ container |
ถ้ายังไม่ครบ — แนะนำลุย Beginner Book ก่อน
บทเรียน
| # | บท | สอนอะไร |
|---|---|---|
| 0 | Mindset + Fallacies | 8 fallacies, failure models, two generals, failure detection, networking deep |
| 1 | Consistency + CAP | Linearizable→eventual, CAP, PACELC, Quorum, CRDT, distributed locks |
| 2 | Consensus + Replication | Raft, Paxos, BFT, Spanner, Patroni walkthrough |
| 3 | Distributed Transactions | 2PC, 3PC, Saga (orchestration/choreography), Idempotency patterns, Outbox/Inbox, Event Sourcing, CQRS |
| 4 | Time + Ordering | Physical/Logical clock, Lamport, Vector clock, HLC, TrueTime, CRDT types |
| 5 | Event-Driven (Kafka) | Kafka deep, partitions, consumer groups, DLQ (Dead Letter Queue — คิวเก็บ message ที่ process ไม่ได้), Backpressure (กลไกถอยกลับเมื่อ consumer ตามไม่ทัน), Schema Registry (ที่เก็บ schema กลาง) |
| 6 | Modern Patterns | Service Mesh (Istio configs), Sidecar/BFF, CRDT apps, Gossip, DHT, Edge, Wasm |
| 7 | Service Discovery + LB + Rate Limit | DNS-based discovery, LB algorithms, Rate limiting (Token bucket, Sliding window), Circuit breaker |
| 8 | Distributed Caching | Cache patterns (cache-aside, read/write-through, write-back), invalidation, stampede, Redis Cluster, CDN |
| 9 | Failure Detection + DR | SWIM, Phi accrual, Raft heartbeats (de-facto detector ใน consensus systems เช่น etcd/CockroachDB/TiKV), RTO/RPO, DR strategies, Capacity planning, Observability (logs/metrics/traces), Chaos engineering |
| 10 | Distributed Data Structures | Bloom Filter, HyperLogLog, Count-Min Sketch, Merkle Tree, Consistent Hashing, LSM-Tree, SWIM |
| 11 | Glossary + Cheat Sheet | A-Z glossary, decision trees, architecture review checklist, anti-patterns, recommended reading |
Learning Path (3 ระดับ)
Junior — เริ่มเข้าใจปัญหา
ลุยบท 0-2 ก่อน — ได้ mental model + รู้ว่าทำไม distributed มันยาก
text
สัปดาห์ 1: บท 0 (Mindset + Fallacies)
สัปดาห์ 2: บท 1 (Consistency + CAP) — focus quorum, linearizable, eventual
สัปดาห์ 3: บท 2 (Raft) — ทำ raft.github.io interactiveMid-level — ออกแบบ microservices ได้
ต่อด้วยบท 3-5 — ทำ event-driven + saga + outbox
text
สัปดาห์ 4: บท 3 (Transactions) — implement saga + outbox ใน project ตัวเอง
สัปดาห์ 5: บท 4 (Time + Ordering) — Lamport + Vector clock
สัปดาห์ 6: บท 5 (Kafka) — setup local cluster + producer/consumerSenior — เข้าใจ trade-off ระดับลึก
จบบท 6-11 + อ่าน DDIA ที่บ้านเสริม
text
สัปดาห์ 7: บท 6 (Modern Patterns) — service mesh, CRDT
สัปดาห์ 8: บท 7 (Service Discovery + LB) — production patterns
สัปดาห์ 9: บท 8 (Caching) — multi-tier + CDN
สัปดาห์ 10: บท 9 (Failure Detection + DR) — chaos + observability
สัปดาห์ 11: บท 10 (Data Structures) — Bloom, HLL, Merkle, Consistent Hash
สัปดาห์ 12: บท 11 (Glossary + Cheat Sheet) — quick reference
หลังจากนั้น: DDIA + Spanner paper + production hands-onวิธีอ่านที่แนะนำ
- อ่าน 8 Fallacies ก่อน — กระตุ้นให้คิดเหมือน distributed engineer
- ทำ checkpoint ทุกบท — ลองทำเอง อย่าแค่อ่าน
- เรียน Kafka + Raft hands-on — concept + practice
- อ่าน DDIA หลังจบ — เล่มนี้ + DDIA = ทรงพลัง
📌 คำแนะนำสำหรับการอ่านครั้งแรก (basic / mid-level): เล่มนี้มีเนื้อหา ~30,000+ บรรทัดรวม 2 โฟลเดอร์ ครอบคลุมตั้งแต่ basic ถึง senior+ — ไม่ต้องอ่านครบทั้งหมดในรอบแรก!
- รอบแรก: อ่านเฉพาะ §1-5 ของแต่ละบท + ทุกที่ที่ไม่มี marker ⚠️
- ข้าม: ทุก section ที่ marked "⚠️ (ขั้นต่อยอด...)" / "(Senior+ level)" / "(production team)" — กลับมาเมื่อทำงานในบริบทนั้นจริง
- รอบ 2-3: หลังลงมือทำ project จริง กลับมาอ่าน sections ที่ข้ามไปก่อน — จะเข้าใจมากกว่า 5-10 เท่า
Senior+ ที่กลับมาทบทวน → flip ดู advanced sections ได้ตามต้องการ
แหล่งอ้างอิงเพิ่มเติม
หนังสือ
- DDIA — "Designing Data-Intensive Applications" (Kleppmann) ⭐
- Database Internals (Petrov)
- Building Microservices (Sam Newman)
- Site Reliability Engineering (Google, free online)
Papers (เลือกอ่านที่สนใจ)
💡 papers พวกนี้เป็น optional/advanced — มือใหม่ข้ามได้ เล่มนี้สรุปให้แล้ว กลับมาอ่านตอนอยากลึกถึงรากต้นฉบับ
- "Time, Clocks, and the Ordering of Events" (Lamport, 1978)
- "Paxos Made Simple" (Lamport, 2001)
- "In Search of an Understandable Consensus Algorithm" (Raft, 2014)
- "Spanner: Google's Globally-Distributed Database" (2012)
- "Dynamo: Amazon's Highly Available Key-value Store" (2007)
- "Conflict-free Replicated Data Types" (Shapiro, Preguiça, Baquero, Zawirski, 2011) — canonical CRDT paper
- "Logical Physical Clocks and Consistent Snapshots in Globally Distributed Databases" (Kulkarni et al., 2014) — HLC paper