Marcio Cunha

Building Distributed Tracing Pipelines with OpenTelemetry in Hybrid Microservices

Learn how to architect end-to-end observability in hybrid microservices architectures using OpenTelemetry, ensuring precise request tracking across legacy and cloud environments.

Marcio Cunha•3 min
Also available in:EspañolPortuguês
Summary
  • The OpenTelemetry ecosystem unifies metrics, logs, and traces collection without locking the application into specific vendors.
  • Hybrid environments require rigorous propagation of HTTP tracing contexts to prevent orphaned traces in the cloud.
  • Choosing the right collector reduces application performance impact by centralizing telemetry processing.
  • Efficient exporters optimize network traffic by grouping data batches before sending them to storage backends.
  • Standardizing transaction identifiers accelerates root cause analysis in highly complex distributed systems.

The Visibility Challenge in Hybrid Architectures

When a monolithic application is split into dozens of microservices, the simplicity of tracking a request disappears. In hybrid systems, where part of the workloads runs on physical on-premise servers and another part in public clouds, this complexity multiplies. In practice, this means a single user click can trigger services scattered across continents and different technologies, making bottleneck identification a monumental task without proper instrumentation.

To solve this problem, modern software engineering adopted distributed tracing, a technique that tags a request with a unique identifier right at the system edge. As this request travels through APIs, message queues, and databases, each component adds a timestamp and metadata. This complete history forms a detailed timeline, allowing operations teams to discover exactly where an error occurred or which step took the longest to execute.

OpenTelemetry as the Industry Standard for Telemetry

Historically, each monitoring tool required installing a proprietary agent directly into the application code, creating tight technological coupling. If the company decided to switch observability vendors, all code needed rewriting. OpenTelemetry, an open-source project maintained by the Cloud Native Computing Foundation, solves this pain by establishing a single, universal standard for telemetry data collection, unifying metrics, logs, and traces in one place.

In practice, OpenTelemetry acts as a universal power plug. You instrument your code using standardized libraries that generate performance data, and this data can be sent to any compatible analysis platform, such as Jaeger, Prometheus, or commercial services. This eliminates vendor lock-in and ensures the engineering team has the freedom to choose the best market tools without rewriting monitoring logic.

Data Collection Pipeline Architecture

Building a reliable data pipeline requires separating trace generation from processing and storage. The first layer of this pipeline is automatic or manual instrumentation within microservices, which collects events asynchronously to avoid hurting user response times. This raw data is sent to an intermediate component called the OpenTelemetry Collector, which acts as an intelligent postal worker tasked with organizing the flow.

The OpenTelemetry Collector typically runs as a separate process or container on each cluster node. It receives data from applications, performs filtering operations to discard irrelevant noise, masks sensitive customer information for security reasons, and finally packages and ships this information to the time-series database or trace visualization tool. This separation protects the application against sudden outages in the monitoring system.

Context Propagation Across Technological Boundaries

The heart of distributed tracing lies in context propagation, the mechanism by which tracing metadata is passed from one service to another. When microservice A makes an HTTP request to microservice B, it injects specific headers into network protocols containing the current trace ID and the specific operation ID. Microservice B extracts these headers upon receiving the call and continues the execution tree, ensuring the link is not lost.

In hybrid systems, this propagation faces additional challenges due to the mix of legacy and modern protocols. Queue-based messaging systems like RabbitMQ or Apache Kafka require trace metadata to be embedded directly into the metadata of messages sent to topics. If a single component along the way fails to pass these headers, the timeline breaks, generating orphaned nodes that hinder troubleshooting in production.

Practical Implementation with Collector Configuration

To bring the architecture into operation, the first practical step involves configuring the OpenTelemetry Collector rules file, defining how data will be received, processed, and exported. The snippet below illustrates a typical YAML configuration where we receive data via the standard OTLP protocol, apply a batch filter to optimize network usage, and send everything to an open-source collector backend.

receivers:  otlp:    protocols:      grpc:      http:processors:  batch:    send_batch_size: 1024    timeout: 1s  memory_limiter:    check_interval: 1s    limit_percentage: 80    spike_limit_percentage: 20exporters:  otlp/backend:    endpoint: