LabHub

Blog

Observability 2025 Complete Guide: OpenTelemetry, Grafana / Datadog / Honeycomb / SigNoz, SLO and Error Budget, LLM Observability (2025)

한국어English日本語中文

Season 5 Ep 8 — If governance is "management", observability is "grasping reality". In 2025 observability does not stop at log, metric, and trace — it takes in LLMs, data pipelines, and user experience as well.

Prologue — "Observability is not insurance, it is the product"

From 2015 to 2020 many teams treated observability as "insurance for incident response". 2025 is different:

The era of "observability is a cost" is over. "Observability is part of product quality" is the 2025 standard.


Chapter 1 · Three Signals + α

1.1 The basic three

1.2 Extended signals

1.3 Correlation


Chapter 2 · OpenTelemetry — a Standard Comes of Age

2.1 What it is

2.2 Structure

2.3 Why it matters

2.4 How hard it is to adopt


Chapter 3 · The Grafana Stack (Open)

3.1 Components

3.2 Grafana Cloud

3.3 Strengths

3.4 Limits


Chapter 4 · Datadog, New Relic, Splunk, Dynatrace — the SaaS Giants

4.1 Datadog

4.2 New Relic

4.3 Splunk

4.4 Dynatrace

4.5 Comparison

ToolStrengthWeakness
DatadogIntegrationCost
New RelicTransparent pricingFeatures spread thin
SplunkLogs and securityCost and complexity
DynatraceAutomationLearning curve
Grafana CloudValue for moneyIntegrated experience

Chapter 5 · The New Generation — Honeycomb, SigNoz, Axiom, Tinybird

5.1 Honeycomb

5.2 SigNoz

5.3 Axiom

5.4 Tinybird

5.5 OpenObserve

5.6 What they share


Chapter 6 · Implementing Logs, Metrics, and Traces

6.1 Metric design

6.2 Log design

6.3 Trace design

6.4 RUM

6.5 Synthetic


Chapter 7 · SLO, SLI, Error Budget

7.1 Concepts

7.2 Why it matters

7.3 Metrics in practice

7.4 Error budget policy

7.5 Tools


Chapter 8 · LLM Observability (an extension of Ep 6)

8.1 Additional signals

8.2 Major tools

8.3 OpenLLMetry

8.4 Operating patterns


Chapter 9 · Chaos Engineering and Resilience

9.1 Philosophy

9.2 Tools

9.3 In practice

9.4 Applying it to data pipelines


Chapter 10 · Alerting, On-call, Incident Response

10.1 Alert design

10.2 On-call culture

10.3 Incident response

10.4 Postmortem

10.5 Tools


Chapter 11 · Cost Optimization

11.1 Log cost

11.2 Metric cost

11.3 Trace cost

11.4 SaaS cost

11.5 A realistic target


Chapter 12 · Observability at Korean Companies

12.1 Where things stand

12.2 Adopting LLM observability

12.3 Korean particulars

12.4 Reference cases


Chapter 13 · 10 Antipatterns

13.1 "Add observability after the deploy"

It is only worth something if you instrument from the start.

13.2 Every log at DEBUG

Cost explodes and important signals get buried.

13.3 Alert spam

An alert on every error → alert fatigue syndrome.

13.4 Operating without SLOs

Political fights over "development vs stability".

13.5 Analyzing an incident without traces

Hours spent, cause unknown.

13.6 High-cardinality metric labels

Prometheus blows up.

13.7 Sensitive information in logs

Audit findings and leak incidents.

13.8 The same incident repeating with no postmortem

There is no learning loop.

13.9 Leaving SaaS cost unattended

End-of-month invoice shock.

13.10 Zero LLM observability

Hallucination, refusals, and runaway cost all go undetected.


Chapter 14 · Checklist — 12 Things Before an Observability Launch


Chapter 15 · Next Post Preview — Season 5 Ep 9: "Data Team Organization and Careers"

With technology, governance, and observability stacked up, next come the people who build them. Ep 9 is about the organization and careers of data and AI teams.

"What matters more than tools is the team" — where data organizations stand in 2025.

See you in the next post.


Summary: In 2025 observability expanded into "three signals + α" (metric, log, trace plus profile, RUM, synthetic, LLM events), and as OpenTelemetry became the collection standard, vendor lock-in is dissolving fast. The Grafana stack is open and good value, Datadog, New Relic, Splunk, and Dynatrace are integrated and enterprise, and Honeycomb, SigNoz, and Axiom are the new generation of low cost and high explorability. SLO and Error Budget became the language for balancing development against stability, and LLM observability (LangFuse, Phoenix, Helicone, Datadog LLM) went mainstream as an extension of Ep 6. The declaration that "observability is not insurance, it is product quality" is the 2025 default. Korean companies are fusing network separation, the Korean language, and SLO culture, and the next post is about the data teams and people who build that observability.

Comments

No comments yet.

Sign in to leave a comment