Monitoring and Observability
Monitoring provides visibility into a system, a core requirement for assessing service health and diagnosing problems when they arise. The purpose of implementing monitoring is based on:
-
Generating alerts for conditions that require attention.
-
Investigating and diagnosing problems.
-
Visualizing system information graphically.
-
Gaining insights into resource usage trends or service health for long-term planning.
-
Comparing system behavior before and after a change, or between two groups in an experiment.
Observability goes far beyond being a synonym for monitoring. The term "observability" comes from Rudolf Kalman's control theory and refers to the ability to infer a system's internal state based on its external outputs. Applied to software systems, this concept refers to the ability to understand an application's internal state based on its telemetry. Not all systems allow or provide enough information to be "observed"; we will therefore classify as observable those that do. Being observable is one of the fundamental attributes of cloud-native systems.
Additional Information
Classroom: https://classroom.google.com/w/NDg2MTQ2MTQwMTYy/t/all
Monitoring and Observability: Monitoring & Observability
The observability tool at LATAM is Grafana Cloud. If you need access, follow these steps:
The access request process is as follows:
- Log in with your LATAM account on the Grafana Cloud stack:
- After logging in, create a ticket at:
https://projectmanagement.appslatam.com/servicedesk/customer/portal/19
Indicating which domain/squad you belong to.
- The Monitoring and Observability squad will activate your Grafana account and you will receive ticket closure notification.