You’re reading the free preview.Unlock every conversation for €2.99
FreeFree previewChapter 11

Chapter 11 · Free preview
Observability: The Watchtower
When Production Becomes a Black Box
What you will understand
- Describe the three pillars of observability (logs, metrics, traces)
- Explain how the ELK stack collects logs and makes them searchable
- Describe how Prometheus scrapes metrics and Grafana visualises them
- Explain distributed tracing and know when to use it
- Explain the difference between SLI, SLO and SLA with an example
Written by EduVerse for this preview; the slides below are the book’s own words.
Slides from the chapter
Each slide is a passage from the book with the figure or listing it talks about. Swipe, use the arrow keys or the buttons.
Slide 1 of 4
Log Levels
11.2.3 · p. 238
Conversation from The Software Realm, Decoded
Chapter 11 · 11.2.3 Log Levels · p. 238

Senior developer
Not everything is equally important. Use log levels.
Explore this figure
Pick one to highlight it and read what the book says about it.
Other labels in the figure (11)
Description
Stacked levels: five boxes from top to bottom, 'DEBUG' (detailed info, verbose), 'INFO' (normal operations), 'WARN' (something unusual), 'ERROR' (something failed) and 'FATAL' (service crashing), with a downward arrow labelled 'Increasing Severity'.
Description written by EduVerse; the figure itself is from the book.

Peter
So in production, I only show INFO and above, not DEBUG?

Senior developer
Yes, correct. DEBUG is too noisy. Only enable it when actively debugging a specific issue.
Peter's Slow Request Mystery
11.5.1 · p. 247
Conversation from The Software Realm, Decoded
Chapter 11 · 11.5.1 Peter's Slow Request Mystery · p. 247

Peter
All services are fast! But the user experienced 10 seconds. Where did the time go?

Senior developer
This is the distributed systems problem. A request touches multiple services. Metrics show each service individually, but you need to see the full journey. That’s what distributed tracing solves.
What the book shows with it
Figure from The Software Realm, Decoded
Chapter 11 · 11.5.2 The Tracing Concept · p. 247
A trace = the complete journey of one request through all services
Explore this figure
Pick one to highlight it and read what the book says about it.
Other labels in the figure (15)
Description
Trace diagram: 'API Gateway' leads to 'Order Service', then 'Kafka', which branches to 'Email Service' and 'Inventory Service'. Timed spans run beside them: 'Span 1: 55ms', 'Span 2: 5ms' and, from Email Service, 'Span 3: 2,045ms' marked 'SLOW! Bottleneck!'. A note says Email Service took 2 seconds, so check its logs.
Description written by EduVerse; the figure itself is from the book.
Error Budgets: How Much Failure is OK?
11.4.6 · p. 246
Figure from The Software Realm, Decoded
Chapter 11 · 11.4.6 Error Budgets: How Much Failure is OK? · p. 246
Explore this table
Pick one to highlight it and read what the book says about it.
Other cells in the table (19)
Description
Table: SLOs with allowed downtime per month and per year. 90% allows 3 days or 36 days; 99% allows 7.2 hours or 3.65 days; 99.9% (three nines) 43 minutes or 8.76 hours; 99.99% (four nines) 4.3 minutes or 52 minutes; 99.999% (five nines) 26 seconds or 5.26 minutes.
Description written by EduVerse; the figure itself is from the book.
Slide 1 of 4
Try it yourself
A simulation built by EduVerse around this chapter. It runs in your browser; nothing is sent anywhere.
The full chapter
This preview shows 5 of the chapter’s 46 passages. The full chapter has:
- 7 sections
- 26 conversations
- 18 figures and tables
- 2 What They Say boxes
- 7 knowledge-check questions
Sections in this chapter
- 11.1The Three Pillars of Observability
- 11.2Logs: The Event Journal
- 11.3Metrics: The Health Dashboard
- 11.4Service Level Objectives: Defining Success
- 11.5Distributed Tracing: Following the Request
- 11.6APM Tools: All-in-One Observability
- 11.7Peter's Takeaways