The blueprint
A mock exam here has 60 questions in 90 minutes, split across the domains in the same proportions as the official exam guide. The official pass mark is 75%. FixOps scores practice against a target of 75%.
-
01
Observability Concepts
-
02
Prometheus Fundamentals
-
03
PromQL
-
04
Instrumentation and Exporters
-
05
Alerting & Dashboarding
Sample questions
Three of the 10 questions in the free diagnostic. Open one to see the answer and why.
In monitoring, what is a metric?
- A numeric measurement of a system, recorded over time
- A record of the path one request took through several services
- A snapshot of a process's memory at the moment it crashed
- A text record of a single event, with a timestamp and a free-form message
Answer: A numeric measurement of a system, recorded over time. A metric is a number that describes some aspect of a system, such as requests served or memory in use, sampled repeatedly so that it forms a time series. A log line describes one event, and a trace follows one request.
What does a distributed trace show?
- The CPU usage of each node in a cluster
- The list of alerts that are currently firing
- The path of one request through the services it touched
- The total number of requests a service received in a day
Answer: The path of one request through the services it touched. A trace follows a single request across process and network boundaries. It is made of spans, and it shows where time was spent and where an error occurred along the way.
What does the Prometheus server do?
- It collects log lines from applications and indexes their text
- It renders dashboards for end users and manages their accounts and permissions
- It scrapes targets, stores the samples and evaluates queries and rules
- It sends notifications to email, chat and paging systems
Answer: It scrapes targets, stores the samples and evaluates queries and rules. The server is the core: it discovers and scrapes targets, writes samples to its local time series database, answers PromQL queries and evaluates recording and alerting rules. Notifications are the Alertmanager's job and dashboards are usually Grafana's.
Revision notes: Observability Concepts
The notes for one domain, free to read here and in the app. FixOps Pro has them for all 5 domains.
What metrics, logs and traces each tell you, how pull and push collection differ, how targets are discovered, and how SLIs, SLOs, SLAs and error budgets fit together.
Metrics, logs and traces
- A metric is a numeric measurement recorded over time; a time series is the samples of one metric name with one set of labels.
- A log records one event with its details; a trace follows one request across services as a tree of spans.
- Metrics are cheap for trends and alerting because their volume does not grow with every request.
- Per-request detail, such as the parameters of a failed call, is found in logs or traces, not in metrics.
- Every distinct combination of label values is a separate series, so unbounded labels such as user ids exhaust memory.
- Prometheus stores numeric samples; log lines and traces need other systems.
- RED (rate, errors, duration) describes request-driven services; USE (utilisation, saturation, errors) describes resources.
- The four golden signals are latency, traffic, errors and saturation.
Tracing
- A span is one timed operation in a trace and refers to its parent span.
- Services join spans into one trace by propagating the trace context with each request.
- Tracing systems sample because recording every request in full costs too much.
- An exemplar attaches a trace id to a metric sample, linking a graph to an example request.
Pull and push
- Prometheus pulls: it scrapes each target's HTTP metrics endpoint at the scrape interval.
- Pulling makes a dead target visible, because the scrape fails and up becomes 0.
- A job that ends between scrapes is never seen by pulling; such jobs push to the Pushgateway.
- The Pushgateway is for short-lived, service-level batch jobs, not a general push path.
- The Pushgateway keeps pushed series until they are deleted, so old values stay visible.
Service discovery
- Service discovery keeps the target list current as instances appear and disappear.
- static_configs suits a small set of fixed targets; dynamic environments need discovery.
- Kubernetes discovery has the roles node, service, pod, endpoints, endpointslice and ingress.
- File-based and HTTP discovery let any external system supply targets.
- Discovery adds __meta_ labels; they are dropped after relabeling unless copied to ordinary labels.
SLIs, SLOs and SLAs
- An SLI is a measurement, such as the share of successful requests or of requests under 300 ms.
- An SLO is a target for an SLI over a period, such as 99.9% over 30 days.
- An SLA is a commitment to customers with consequences when it is missed.
- The error budget is what the SLO leaves: 0.1% for a 99.9% objective.
- A 100% objective leaves no budget for change and is not noticeably better for users.
- A burn rate of 10 consumes the budget ten times faster than the period allows.
- An error ratio is the rate of failed requests divided by the rate of all requests, each summed across instances.
Easy to mix up
- A metric counts what happened; a log or a trace shows one occurrence.
- An SLI is measured, an SLO is a target, and an SLA is a promise with consequences.
- The Pushgateway is for batch jobs; it does not turn Prometheus into a push system.
- RED is for services; USE is for resources.