# Prometheus replication - Is federation the way to go?

**URL:** <https://discuss.prometheus.io/t/prometheus-replication-is-federation-the-way-to-go/367>\
**Category:** Prometheus server\
**Created:** [July 9, 2021, 1:29pm UTC](https://discuss.prometheus.io/t/prometheus-replication-is-federation-the-way-to-go/367 "2021-07-09T13:29:31Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Aldar](https://avatars.discourse-cdn.com/v4/letter/a/b487fb/32.png) [@Aldar](https://discuss.prometheus.io/u/Aldar)\
**Post date:** [July 9, 2021, 1:29pm UTC](https://discuss.prometheus.io/t/prometheus-replication-is-federation-the-way-to-go/367/1 "2021-07-09T13:29:31Z")

</div>

Hello,

I am trying to build a complete replica of our primary prometheus instance so that other applications, not under our complete control, could access the data without posing any risk to our primary node that we use internally ourselves, and that is production-critical.

To facilite this data replication, I used the Prometheus’ built-in federation solution. The replica’s complete configuration is as follows:

```auto
global:
  scrape_interval: 4m
  scrape_timeout: 3m

scrape_configs:
    - job_name: 'federate'
      tls_config:
        insecure_skip_verify: true

      honor_labels: true
      metrics_path: '/federate'
      scheme: 'https'

      params:
        'match[]':
          - '{ __name__ =~".+"}'

      static_configs:
        - targets:
          - '*master-hostname*'

```

As you can see, I had to set a quite large scrape interval values, as I’m asking the primary node for _all_ the labels and values it has - Such a HTTP request takes anywhere between 2 and 3 minutes to complete. That is problematic, as the replica then provides fewer datapoints than the primary, resulting in less detail in the final graphs we plot from the datapoints.

Is there a better way to replicate data from one Prometheus node to another than federation? Even if through a 3rd party solution.

---

<div class="post-metadata">

**Author:** ![stuart](https://dub1.discourse-cdn.com/flex017/user_avatar/discuss.prometheus.io/stuart/32/18_2.png) [@stuart](https://discuss.prometheus.io/u/stuart)\
**Post date:** [July 9, 2021, 2:32pm UTC](https://discuss.prometheus.io/t/prometheus-replication-is-federation-the-way-to-go/367/2 "2021-07-09T14:32:27Z")

</div>

You could look at using remote write, or just scrape the underlying  
targets directly from the second instance.

---

<div class="post-metadata">

**Author:** ![SuperQ](https://dub1.discourse-cdn.com/flex017/user_avatar/discuss.prometheus.io/superq/32/4_2.png) [@SuperQ](https://discuss.prometheus.io/u/SuperQ)\
**Post date:** [July 9, 2021, 3:13pm UTC](https://discuss.prometheus.io/t/prometheus-replication-is-federation-the-way-to-go/367/3 "2021-07-09T15:13:14Z")

</div>

Federation is not meant for replication. Like @stuart said, enable remote write with `--enable-feature=remote-write-receiver`

> **[Disabled Features | Prometheus](https://prometheus.io/docs/prometheus/latest/disabled_features/#remote-write-receiver)**
>
> An open-source monitoring system with a dimensional data model, flexible query language, efficient time series database and modern alerting approach.

---

<div class="post-metadata">

**Author:** ![Aldar](https://avatars.discourse-cdn.com/v4/letter/a/b487fb/32.png) [@Aldar](https://discuss.prometheus.io/u/Aldar)\
**Post date:** [July 15, 2021, 4:07pm UTC](https://discuss.prometheus.io/t/prometheus-replication-is-federation-the-way-to-go/367/4 "2021-07-15T16:07:24Z")

</div>

Okay, thank you both for the suggestion. I am sorry it took me a while to reply, however, even after enabling and setting up the remote writes from one of our testing instances onto my prometheus instance, it still does not produce identical datasets.

I tried increasing the max\_samples\_per\_send to something ridiculous (Like 25k) and Capacity to 10x that.  
Attached is a comparison of the two instances - left is configured to remote-write samples to the right instance.

 ![prometheus-datasets](https://europe1.discourse-cdn.com/flex017/uploads/prometheus/original/1X/f750919dc40a3dac86f5b4e2f629712d5d9076ba.png)

The current setting in use is:

```auto
remote_write:
    - url: http://127.0.0.1:19090/api/v1/write
      max_samples_per_send: 25000
      capacity: 250000

```

The URL points to an stunnel4 tunnel leading to the replica (For security sake)

I am at a bit of a loss at how to do what I want - a 1:1 dataset replica. The two nodes are connected through a gigabit connection, meaning bandwidth shouldn’t be too much of a problem. And even if network issues were to arise, I don’t mind dropped samples (E.g.: Don’t need a HA solution), but I need to have ± the same datasets on the two nodes.
