Skip to content

SRE

SLI, SLO, ইনসিডেন্ট রেসপন্স, chaos engineering

0/21 অধ্যায় · 0 XP অর্জিত

  1. 00 8 সপ্তাহের রোডম্যাপ: Fullstack → SRE একজন কর্মরত fullstack engineer-কে junior-SRE-ready operator বানানোর জন্য দুই মাসের একটি সলিড প্ল্যান। প্রতিদিনের ব্রেকডাউন, রিয়েল ল্যাব, আর একটি ফাইনাল ক্যাপস্টোন।
  2. 01 SRE আসলে কী Class SRE implements DevOps। error-budget contract, toil cap, আর embedded engineer model যা Google-এর reliability-কে কাজ করায়।
  3. 02 SLIs, SLOs & Error Budgets সঠিক SLI বেছে নাও, এমন একটা SLO সেট করো যা lawyer review-তে টেকে, আর Google-এর CRE team যেভাবে করে সেভাবে budget burn করো।
  4. 03 Golden Signals, RED, and USE যে তিনটা monitoring framework আসলে গুরুত্বপূর্ণ, কখন কোনটা ব্যবহার করবে, আর যে Prometheus + Grafana stack সেগুলো সব expose করে।
  5. 04 Incident Response & On-Call ICS roles, severity classification, comms cadence, আর যে on-call rotation engineer-দের burn out করে না।
  6. 05 Blameless Postmortems Google, Etsy, আর Stripe যে full template ব্যবহার করে, সাথে action-item discipline যা একই incident দ্বিতীয়বার আটকায়।
  7. 06 Capacity Planning & Load Testing Little's Law, Universal Scalability Law, headroom, আর একটা রিয়েল k6 + Locust load test যা তুমি আজই চালাতে পারো।
  8. 07 Production Readiness Reviews যে PRR চেকলিস্ট প্রতিরোধযোগ্য launch-into-fire-এর 80% ঠেকিয়ে দেয়, একটি বাস্তব launch gate Terraform module সহ।
  9. 08 Chaos Engineering Netflix Simian Army থেকে Chaos Mesh পর্যন্ত hypothesis-driven failure injection, বাস্তব experiment আর একটা safety harness সহ।
  10. 09 Disaster Recovery & Backups RTO, RPO, multi-region failover, আর সেই restore drill যেটা প্রমাণ করে আপনার backup আসলেই আছে।
  11. 10 Toil & Automation toil মাপা, 50% cap, আর one-off script থেকে self-healing operator পর্যন্ত automation taxonomy।
  12. 11 Scaling & Distributed Systems — 8-Week Companion Roadmap শূন্য থেকে scalable system ডিজাইন, বিল্ড আর অপারেট করা পর্যন্ত। দিনে 1–2 ঘণ্টা, 8 সপ্তাহ, 8টি বাস্তব project, 5টি case study, mini-YouTube capstone।
  13. 12 Linux Performance Mastery `top` থেকে `perf`, `bpftrace`, আর flame graph পর্যন্ত। production-এ কিছু restart না করে kernel level-এ latency, CPU, memory, আর I/O diagnose করার senior-SRE toolkit।
  14. 13 Network Engineering for SREs BGP, anycast, ECMP, CDN internals, packet capture, আর scale-এ TCP। যে networking layer-এ 'random' production অদ্ভুততা আসলে বাস করে।
  15. 14 SRE-দের জন্য Database Internals MVCC, replication lag, hot rows, query plans, B-tree vs LSM, স্কেলে connection pools। যে DB জ্ঞান 'আমি Postgres চালাই'-কে 'আমি চাপের মুখেও Postgres টিকিয়ে রাখি' থেকে আলাদা করে।
  16. 15 SRE-দের জন্য Distributed Systems Theory CAP, PACELC, FLP, Raft, Paxos, gossip, vector clocks, CRDTs, fencing tokens। যে theory ব্যাখ্যা করে আপনার distributed system কেন যেভাবে ভাঙে সেভাবে ভাঙে।
  17. 16 স্কেলে Kubernetes 1,000+ node cluster, multi-tenancy, RBAC, NetworkPolicy, OPA/Kyverno, GitOps, etcd tuning। যখন 'just run kubectl apply' আর কোনো strategy নয়, তখনকার operating model।
  18. 17 Service Mesh Internals Envoy, Istio, Linkerd, sidecar vs ambient, mTLS, xDS, retries, circuit breakers, traffic shifting। একটা mesh আসলে কী করে আর কখন এর জটিলতা পুষিয়ে দেয়।
  19. 18 FinOps ও Cost Engineering Unit economics, rightsizing, spot, savings plans, cost-aware SLOs। যে senior SRE skill 'cloud bill অনেক বেশি'-কে একটা tracked, owned, কমতে থাকা সংখ্যায় পরিণত করে।
  20. 19 Reliability Culture ও SRE Org Design Staff+ SRE work, embedding, charters, blame-aware org, mentoring, sustainable on-call। যে non-technical lever প্রতিটা reliability program বানায় বা ভাঙে।
  21. 20 12-মাসের Mastery Roadmap — Junior SRE → Senior/Staff যে year-long plan 8-week roadmap যেখানে থামে সেখান থেকে তুলে নেয়। Monthly milestone, বাস্তব production project, deep reading, আর যে artifact staff-level capability প্রমাণ করে।