Atlassian is overhauling the internal architecture of its large-scale metrics platform by moving key backend components to OpenTelemetry while keeping its existing StatsD interface available to application teams. The platform processes telemetry from around 100,000 hosts across 14 regions and operates against a 99.95% service-level objective. Rather than forcing every application to change instrumentation immediately, Atlassian chose to modernise the platform first and leave application-side migration for a later phase.
This approach allows existing services to continue sending metrics in the same way while the collection, routing, aggregation and forwarding layers are rebuilt underneath them. By preserving the familiar StatsD interface, Atlassian avoided making a fleet-wide application change a prerequisite for replacing the telemetry infrastructure. It also reduced the operational risk of trying to change both the backend platform and application instrumentation at the same time.
Moving Beyond gostatsd
For much of the past decade, Atlassian relied on gostatsd, an open-source StatsD implementation maintained by the company. The software handled metrics collection as a sidecar running alongside individual hosts and also supported aggregation further downstream in the pipeline. This architecture served Atlassian well, but its UDP-based design was focused on metrics and did not provide native support for traces or logs.
At the same time, more of Atlassian's internal telemetry platform was moving towards OpenTelemetry. With OpenTelemetry now providing standardised support for metrics, traces and logs, the company saw an opportunity to consolidate its observability infrastructure around a common platform. Instead of beginning with application reinstrumentation, Atlassian decided to keep the existing interface and replace the underlying components first.
Four-Stage Collector-Based Pipeline
The redesigned metrics platform is divided into four main stages: collection, ingest, aggregation and forwarding. Each stage runs a purpose-built distribution of the OpenTelemetry Collector, giving Atlassian the flexibility to replace and optimise individual components independently. This modular approach also makes it easier to evolve the platform without requiring a complete redesign every time one layer changes.
The collection layer accepts both traditional StatsD traffic and data sent through the OpenTelemetry Protocol, or OTLP. Existing applications can therefore continue sending StatsD metrics over UDP, while newer services can adopt OpenTelemetry directly. This gives teams the freedom to migrate application instrumentation gradually instead of forcing every service to switch at once.
Metrics And Tracing Sidecars Are Combined
Atlassian had already been using the OpenTelemetry Collector as a central part of its tracing pipeline and for host-level metrics before the wider metrics migration started. This made it possible to consolidate StatsD metrics and tracing into a single Collector-based sidecar. The combined sidecar continues accepting legacy StatsD traffic while also supporting OTLP data.
Removing the separate gostatsd sidecar produced measurable infrastructure savings. Across Atlassian's most expensive Micros services, average CPU usage was reduced by about 3.9% per service. Atlassian estimates that this translated into roughly a 30% reduction in overall sidecar costs across the fleet.
OpenTelemetry also provides common conventions across metrics, traces and logs, including shared resource information and mechanisms for correlating telemetry using trace context. This gives Atlassian a more consistent foundation for observability as more services gradually move away from older telemetry formats.
Rethinking Routing For Stateful Aggregation
One of the more complex parts of the migration involved the ingest and aggregation layers. Metric aggregation is inherently stateful because datapoints belonging to the same time series must reach the same aggregator. If data from a single series is spread across multiple aggregators, the platform cannot reliably calculate the correct combined result.
Atlassian previously used an internal proxy called nomad, which distributed traffic by hashing the service and environment. The problem was that metric volumes varied dramatically between services. A large service could therefore place significantly more load on one shard than another, resulting in uneven CPU usage across the aggregation cluster.
The company replaced this approach with the OpenTelemetry Collector's load-balancing exporter. Instead of routing at the service level, the exporter can hash traffic using a stream ID that identifies a specific metric time series. This means individual metric streams stay on a consistent aggregator, while different streams from the same high-volume service can be distributed across several shards.
More Even Load And Much Lower Data Volume
The new routing approach resulted in a more even distribution of CPU usage across aggregation shards. Because traffic is spread at a finer level than before, the aggregation pool can also scale down more effectively when activity drops. This improves both operational efficiency and infrastructure utilisation.
Aggregation itself dramatically reduces the amount of telemetry Atlassian needs to store. The pipeline receives around 4.8 billion datapoints per minute, but after aggregation only approximately 220 million datapoints are written downstream. That represents a reduction of roughly 96%, significantly lowering the volume of data that ultimately reaches storage systems.
Much of Atlassian's metrics traffic uses delta temporality, where each datapoint represents the change since the previous measurement rather than an accumulated total. Existing OpenTelemetry components did not provide exactly the aggregation behaviour required by Atlassian's services, so its engineers developed a custom delta aggregation processor. The company subsequently released that processor through Atlassian Labs.
Aggregation CPU Usage Cut Roughly In Half
According to Atlassian, the redesigned aggregation tier now consumes around half the CPU required by the previous system when handling the same level of traffic. Several improvements contributed to this reduction, including the removal of gostatsd parsing from the aggregators and more balanced traffic distribution between shards. The use of OpenTelemetry Collector components also simplified parts of the processing pipeline.
The forwarding layer was similarly replaced with a stateless Collector distribution called metrics-gateway. This component sends telemetry to destinations including SignalFx and Amazon S3 while taking advantage of built-in OpenTelemetry Collector capabilities such as retries, queuing and backpressure handling. By relying on standard Collector features, Atlassian can reduce the amount of custom infrastructure it needs to maintain.
Supporting AWS Lambda Without A Sidecar
Serverless environments introduced another challenge because a conventional sidecar cannot run alongside AWS Lambda functions in the same way it can with long-running services. To address this, Atlassian developed a separate OpenTelemetry extension specifically for Lambda workloads. The extension retains the same StatsD address and environment variables that were used previously.
Maintaining the existing configuration was another deliberate effort to avoid forcing application teams to make code changes during the backend migration. Lambda workloads can therefore continue sending telemetry to the same endpoint using the same settings while Atlassian changes the implementation underneath. This keeps the migration focused on infrastructure rather than application rewrites.
Production Traffic Revealed What Benchmarks Missed
Atlassian found that laboratory benchmarks did not expose the full runtime cost of the new pipeline. Certain behaviours and resource usage patterns only became visible under real production traffic, so the engineering team used continuous profiling during the migration. This helped identify expensive components and areas where additional optimisation was required.
The company also avoided replacing the entire platform in a single deployment. New components were first tested in development and staging environments before being introduced to less critical workloads. Production rollout then increased gradually from around 1% to 10%, then 50%, before eventually reaching full deployment.
For much of that period, the old and new systems operated side by side. This meant Atlassian had to ensure that the replacement platform behaved consistently with the system it was replacing while the migration expanded. Given the platform's role in monitoring thousands of services and its 99.95% service-level objective, maintaining operational parity was critical.
Significant Infrastructure Savings
Atlassian said gostatsd aggregators and its nomad routing service together accounted for around 38% of CPU requests across its metrics clusters. Nomad alone represented approximately 13% of total resource usage. Replacing these components therefore offered a significant opportunity to reduce infrastructure overhead.
The migration has so far focused primarily on the backend while keeping application-facing interfaces stable. This strategy means application teams are not forced into immediate changes simply because the central platform has been rebuilt. It also allows Atlassian to realise infrastructure savings before the wider application migration is complete.
Application Reinstrumentation Comes Next
The next phase will involve moving application-side telemetry away from StatsD, DogStatsD and other vendor-specific or internally developed clients. Atlassian intends to shift these workloads towards native OpenTelemetry SDKs over time. Importantly, that transition can now happen independently of the platform migration.
This staged strategy reduces the risk and complexity of modernising observability across such a large environment. Instead of coordinating a massive application rewrite with an infrastructure replacement, Atlassian can migrate services progressively according to their own schedules. The backend is already prepared to support both legacy StatsD traffic and newer OpenTelemetry instrumentation during that transition.
Final Thoughts
Atlassian's OpenTelemetry migration demonstrates how a large organisation can modernise observability infrastructure without forcing immediate changes across thousands of applications. By preserving the StatsD interface while replacing the collection, routing, aggregation and forwarding layers underneath it, the company reduced migration risk while also achieving meaningful efficiency gains.
The results are significant, including roughly 30% lower sidecar costs, around 50% less CPU usage in the aggregation tier and a more balanced distribution of workloads across shards. With a platform handling approximately 4.8 billion datapoints per minute from around 100,000 hosts, those improvements can translate into substantial infrastructure savings. Atlassian's next challenge will be moving application instrumentation itself to OpenTelemetry, but the hardest part of the backend transition is already well underway.


Comments 0