Marcio Cunha

Implementing Canary Deployments with Flagger and Prometheus

Learn how to mitigate production risks using progressive delivery and telemetry metrics in modern Kubernetes environments.

Marcio Cunha•4 min
Also available in:EspañolPortuguês
Summary
  • Progressive rollouts drastically reduce the impact of unforeseen production failures
  • Flagger automates the validation cycle using Prometheus as the telemetry source
  • Latency and error rate metrics drive the success or automatic rollback of traffic
  • Declarative configuration via CRDs simplifies integration with service meshes
  • Integrated load tests ensure resilience before the final promotion of the version

The challenge of stability in distributed systems

Releasing new software versions in modern production environments is often a tense moment for engineering teams. In practice, this means that even with rigorous automated tests, subtle bugs or memory leaks slip into the real environment, affecting end users. The microservices architecture and the use of Kubernetes, which is the system for managing and scaling software containers, multiply the complexity of releases.

When an update fails, the traditional reaction is usually manual rescue or a hasty rollback of the entire system. This process generates downtime and erodes trust in the release cycle. To solve this problem, modern engineering adopts progressive release strategies, known as canary deployments, which test new code on a reduced fraction of the user base.

The concept of Canary Deployments in practice

The term canary refers to ancient miners who brought birds to detect toxic gases in mines before they affected humans. In computing, the idea is identical: send a small percentage of real traffic to the new version of the service while the old version handles the rest. In practice, if the new code shows instability, only a tiny group of users will be affected, containing the damage instantly.

However, monitoring this transition manually is unfeasible in systems with thousands of requests per second. This is where observability-based automation tools come in. Instead of relying on human eyes fixed on dashboards, the infrastructure needs to collect real-time telemetry data and make autonomous decisions about the health of the running service.

Flagger and Prometheus: the guardians of telemetry

Flagger is a Kubernetes operator, a software that runs inside the cluster of servers to automate the lifecycle of canary releases. It works by integrating with service meshes or ingress controllers to manipulate traffic surgically. Essentially, Flagger acts as a conductor adjusting traffic pointers based on strict performance rules.

To know if the system is healthy, Flagger queries Prometheus, which is a monitoring system and database focused on collecting numerical metrics from applications and servers. Prometheus stores vital indicators, such as HTTP error rates and request latency. Flagger cross-references this data with predefined thresholds to decide whether the new version deserves more traffic or should be discarded.

Architecture of the progressive delivery flow

When a developer updates a container image in the repository, the continuous integration system applies the change to the cluster. Flagger detects the change and automatically creates secondary Kubernetes objects, including a stable primary version and the canary version itself. At this point, traffic still points entirely to the old code.

Next, the operator starts the test cycle, incrementing the traffic directed to the canary version in controlled steps, usually ten percent at a time. At each time interval, Flagger asks Prometheus if the error rate remains below an acceptable limit, such as one percent. If the metric breaches the limit, the process is aborted and traffic returns to the safe base.

Step-by-step configuration with Custom Resources

To put this logic into operation in Kubernetes, we use custom resources that define Flagger's behavior. Below is an example configuration monitoring the error rate and latency of a web microservice:

apiVersion: flagger.app/v1beta1
kind: Canary
metadata:
  name: my-service
  namespace: production
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: my-service
  service:
    port: 80
    gateways:
    - public-gateway.istio-system.svc.cluster.local
    hosts:
    - api.example.com
  analysis:
    interval: 30s
    threshold: 5
    maxWeight: 50
    stepWeight: 10
    metrics:
    - name: request-success-rate
      thresholdRange:
        min: 99
      interval: 30s
    - name: request-duration
      thresholdRange:
        max: 500
      interval: 30s

In this manifest, we configure a thirty-second interval for each check. Flagger increases the canary traffic weight by ten percent each cycle until it reaches the maximum limit of fifty percent. If the success rate drops below ninety-nine percent or latency exceeds five hundred milliseconds, the mechanism automatically discards the version.

Operational considerations and conclusion

Implementing telemetry-based canary deployments removes the human factor and intuition from release processes, replacing them with objective data. In practice, this means teams gain speed without sacrificing operational safety. Flagger and Prometheus form a robust duo that automates surveillance and protects the end user experience against unexpected regressions in production.

In short, adopting this architecture transforms deployment from a stressful and manual event into a transparent and resilient routine. By relying on real metrics to guide code promotion, organizations build systems capable of self-protection and high availability under any circumstance.