Skip to content

Troubleshooting etcd

Info

The whole procedure must be followed as root. To elevate your privileges to root, you can use the su - command.

etcd is a distributed key-value store used as the coordination point between the nodes of the PostgreSQL cluster. It reliably centralizes and shares the state of the cluster, including the identity of the leader, the state of the members and the locking mechanisms.

patroni relies on etcd to drive the election of the leader node and to trigger failovers automatically when an incident occurs. Installed on all nodes, etcd guarantees the overall consistency of the cluster, without taking part in storing or processing PostgreSQL data.

patroni therefore depends on etcd, and cannot work properly if etcd is not working.

Info

The etcd configuration is stored in the /etc/default/etcd file on each node.

Logs

The whole procedure relies on the logs of the etcd service, which you get with the command below:

1
journalctl -u etcd

Tip

To follow the logs live, add the -f option to the command:

1
journalctl -fu etcd

Info

The etcd logs are also available in /var/log/syslog.

Error log explanations

health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: i/o timeout

If the TCP/2380 flow is not open between two nodes, the connection times out and etcd writes the following logs:

1
2
3
4
5
6
Apr 15 11:48:09 PSQL_1 etcd[1342933]: health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: i/o timeout
Apr 15 11:48:09 PSQL_1 etcd[1342933]: health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: i/o timeout
Apr 15 11:48:14 PSQL_1 etcd[1342933]: health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: i/o timeout
Apr 15 11:48:14 PSQL_1 etcd[1342933]: health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: i/o timeout
Apr 15 11:48:19 PSQL_1 etcd[1342933]: health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: i/o timeout
Apr 15 11:48:19 PSQL_1 etcd[1342933]: health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: i/o timeout

Where <PEER_PSQL_NODE_ID> is the ID assigned by etcd to the target node, and <PEER_IP_PSQL_NODE> the IP address of the target node.

Example

In our example, the PostgreSQL cluster is configured as follows:

Node Node IP etcd ID
PSQL_1 192.168.1.1 53d2c129945ccb8b
PSQL_2 192.168.1.2 7176dd381f583d83
PSQL_3 192.168.1.3 e3d5ef565a5bb46c

Running the journalctl -fu etcd command from the PSQL_1 server gives the following logs:

1
2
3
4
5
Apr 15 11:50:39 PSQL_1 etcd[1342933]: health check for peer 7176dd381f583d83 could not connect: dial tcp 192.168.1.2:2380: i/o timeout
Apr 15 11:50:44 PSQL_1 etcd[1342933]: health check for peer 7176dd381f583d83 could not connect: dial tcp 192.168.1.2:2380: i/o timeout
Apr 15 11:50:44 PSQL_1 etcd[1342933]: health check for peer 7176dd381f583d83 could not connect: dial tcp 192.168.1.2:2380: i/o timeout
Apr 15 11:50:49 PSQL_1 etcd[1342933]: health check for peer 7176dd381f583d83 could not connect: dial tcp 192.168.1.2:2380: i/o timeout
Apr 15 11:50:49 PSQL_1 etcd[1342933]: health check for peer 7176dd381f583d83 could not connect: dial tcp 192.168.1.2:2380: i/o timeout

These logs indicate that the connection from the PSQL_1 server to the server with IP address 192.168.1.2 (PSQL_2) and etcd ID 7176dd381f583d83, on port TCP/2380, did not succeed: it timed out. We can deduce that the flow from the PSQL_1 server to the PSQL_2 server is dropped (DROP) by a firewall.

This flow can be dropped (DROP) by the local firewall of the PostgreSQL appliance: check the appliance firewall configuration on the nodes involved. If this configuration is correct, a firewall of the infrastructure is dropping (DROP) the TCP/2380 flow between the two nodes.

health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: connect: connection refused

When a connection between two nodes is refused, etcd reports it with the following logs:

1
2
3
4
5
6
Apr 15 14:04:30 PSQL_1 etcd[1342933]: health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: connect: connection refused
Apr 15 14:04:30 PSQL_1 etcd[1342933]: health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: connect: connection refused
Apr 15 14:04:35 PSQL_1 etcd[1342933]: health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: connect: connection refused
Apr 15 14:04:35 PSQL_1 etcd[1342933]: health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: connect: connection refused
Apr 15 14:04:40 PSQL_1 etcd[1342933]: health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: connect: connection refused
Apr 15 14:04:40 PSQL_1 etcd[1342933]: health check for peer <PEER_PSQL_NODE_ID> could not connect: dial tcp <PEER_IP_PSQL_NODE>:2380: connect: connection refused

Where <PEER_PSQL_NODE_ID> is the ID assigned by etcd to the target node, and <PEER_IP_PSQL_NODE> the IP address of the target node.

The refusal can have several causes:

  • The etcd service is malfunctioning or stopped on the target node. In this case, check the status (systemctl status etcd) and the logs of the etcd service on that node.
  • A firewall rejects (REJECT) the TCP/2380 flow between the two nodes. It can be the local firewall of the appliance or a firewall of the infrastructure.
Example

In our example, the PostgreSQL cluster is configured as follows:

Node Node IP etcd ID
PSQL_1 192.168.1.1 53d2c129945ccb8b
PSQL_2 192.168.1.2 7176dd381f583d83
PSQL_3 192.168.1.3 e3d5ef565a5bb46c

Running the journalctl -fu etcd command from the PSQL_1 server gives the following logs:

1
2
3
4
5
Apr 15 14:04:15 PSQL_1 etcd[1342933]: health check for peer e3d5ef565a5bb46c could not connect: dial tcp 192.168.1.3:2380: connect: connection refused
Apr 15 14:04:15 PSQL_1 etcd[1342933]: health check for peer e3d5ef565a5bb46c could not connect: dial tcp 192.168.1.3:2380: connect: connection refused
Apr 15 14:04:20 PSQL_1 etcd[1342933]: health check for peer e3d5ef565a5bb46c could not connect: dial tcp 192.168.1.3:2380: connect: connection refused
Apr 15 14:04:20 PSQL_1 etcd[1342933]: health check for peer e3d5ef565a5bb46c could not connect: dial tcp 192.168.1.3:2380: connect: connection refused
Apr 15 14:04:25 PSQL_1 etcd[1342933]: health check for peer e3d5ef565a5bb46c could not connect: dial tcp 192.168.1.3:2380: connect: connection refused

These logs indicate that the connection from the PSQL_1 server to the server with IP address 192.168.1.3 (PSQL_3) and etcd ID e3d5ef565a5bb46c, on port TCP/2380, was refused.

On the PSQL_3 server, the systemctl status etcd command shows that etcd is stopped:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
○ etcd.service - etcd - highly-available key value store
    Loaded: loaded (/lib/systemd/system/etcd.service; enabled; preset: enabled)
    Active: inactive (dead) since Wed 2026-04-15 16:09:34 CEST; 38s ago
Duration: 1min 36.087s
    Docs: https://etcd.io/docs
            man:etcd
    Process: 1400558 ExecStart=/usr/bin/etcd $DAEMON_ARGS (code=killed, signal=TERM)
Main PID: 1400558 (code=killed, signal=TERM)
        CPU: 49.071s

Apr 15 12:26:32 PSQL_3 etcd[1197556]: finished scheduled compaction at 1665243 (took 583.297µs)
Apr 15 13:26:32 PSQL_3 etcd[1197556]: store.index: compact 1665949
Apr 15 13:26:32 PSQL_3 etcd[1197556]: finished scheduled compaction at 1665949 (took 574.496µs)
Apr 15 14:26:32 PSQL_3 etcd[1197556]: store.index: compact 1666658
Apr 15 14:26:32 PSQL_3 etcd[1197556]: finished scheduled compaction at 1666658 (took 567.196µs)
Apr 15 15:26:32 PSQL_3 etcd[1197556]: store.index: compact 1667368
Apr 15 15:26:32 PSQL_3 etcd[1197556]: finished scheduled compaction at 1667368 (took 568.096µs)
Apr 15 16:09:34 PSQL_3 systemd[1]: etcd.service: Deactivated successfully.
Apr 15 16:09:34 PSQL_3 systemd[1]: Stopped etcd.service - etcd - highly-available key value store.
Apr 15 16:09:34 PSQL_3 systemd[1]: etcd.service: Consumed 49.071s CPU time.

Once etcd is started, the problem is solved.

rejected connection from "<PEER_IP_PSQL_NODE>:35956" (error "remote error: tls: bad certificate", ServerName "PSQL_1")

If the PostgreSQL cluster certificates have expired, or if they were not properly installed or renewed on the node from which you are reading the logs, the following error appears in the etcd logs:

1
2
3
4
5
6
Apr 15 16:38:54 PSQL_1 etcd[1407644]: rejected connection from "<PEER_IP_PSQL_NODE>:49274" (error "remote error: tls: bad certificate", ServerName "PSQL_1")
Apr 15 16:38:54 PSQL_1 etcd[1407644]: rejected connection from "<PEER_IP_PSQL_NODE>:49294" (error "remote error: tls: bad certificate", ServerName "PSQL_1")
Apr 15 16:38:54 PSQL_1 etcd[1407644]: rejected connection from "<PEER_IP_PSQL_NODE>:49288" (error "remote error: tls: bad certificate", ServerName "PSQL_1")
Apr 15 16:38:55 PSQL_1 etcd[1407644]: rejected connection from "<PEER_IP_PSQL_NODE>:49316" (error "remote error: tls: bad certificate", ServerName "PSQL_1")
Apr 15 16:38:55 PSQL_1 etcd[1407644]: rejected connection from "<PEER_IP_PSQL_NODE>:49304" (error "remote error: tls: bad certificate", ServerName "PSQL_1")
Apr 15 16:38:55 PSQL_1 etcd[1407644]: rejected connection from "<PEER_IP_PSQL_NODE>:60958" (error "remote error: tls: bad certificate", ServerName "PSQL_1")

Where <PEER_IP_PSQL_NODE> is the IP address of one of the cluster nodes.

This log indicates that, when the node you are connected to tried to establish a connection with the other cluster nodes (<PEER_IP_PSQL_NODE>), they refused it, reporting an error on the certificate of the current node.

You should therefore check the validity of the cluster certificates on the node you are connected to.

Warning

If the certification authority configured on a remote node (in /etc/etcd/ca.crt) is incorrect, the error also appears on the node from which you are reading the logs: it then indicates that the remote node cannot validate the certificate presented, because it does not have the right CA.

Example

In our example, the PostgreSQL cluster is configured as follows:

Node Node IP etcd ID
PSQL_1 192.168.1.1 53d2c129945ccb8b
PSQL_2 192.168.1.2 7176dd381f583d83
PSQL_3 192.168.1.3 e3d5ef565a5bb46c

Running the journalctl -fu etcd command from the PSQL_1 server gives the following logs:

1
2
3
4
5
6
Apr 15 16:38:54 PSQL_1 etcd[1407644]: rejected connection from "192.168.1.2:49274" (error "remote error: tls: bad certificate", ServerName "PSQL_1")
Apr 15 16:38:54 PSQL_1 etcd[1407644]: rejected connection from "192.168.1.3:49294" (error "remote error: tls: bad certificate", ServerName "PSQL_1")
Apr 15 16:38:54 PSQL_1 etcd[1407644]: rejected connection from "192.168.1.2:49288" (error "remote error: tls: bad certificate", ServerName "PSQL_1")
Apr 15 16:38:55 PSQL_1 etcd[1407644]: rejected connection from "192.168.1.2:49316" (error "remote error: tls: bad certificate", ServerName "PSQL_1")
Apr 15 16:38:55 PSQL_1 etcd[1407644]: rejected connection from "192.168.1.3:49304" (error "remote error: tls: bad certificate", ServerName "PSQL_1")
Apr 15 16:38:55 PSQL_1 etcd[1407644]: rejected connection from "192.168.1.3:60958" (error "remote error: tls: bad certificate", ServerName "PSQL_1")

These logs indicate that the PSQL_2 and PSQL_3 nodes refused the connection from the PSQL_1 node, because the certificate of PSQL_1 is not valid in this context.

We therefore check the validity of the PSQL_1 node certificates by following the cluster certificate check procedure:

1
2
openssl x509 -in /etc/etcd/peer.crt -noout -enddate
openssl x509 -in /etc/etcd/ca.crt -noout -enddate

The output shows that the node certificate has expired:

1
2
notAfter=Dec 18 14:04:09 2025 GMT
notAfter=Dec 17 14:03:47 2030 GMT

In this case, you need to renew the PostgreSQL cluster certificates.

error listing data dir: /var/lib/etcd/cleanroom

If etcd does not have the right permissions on the /var/lib/etcd/cleanroom directory, the service refuses to start and writes the following logs:

1
2
3
4
5
6
Apr 10 14:15:26 PSQL_1 etcd[454701]: Go OS/Arch: linux/amd64
Apr 10 14:15:26 PSQL_1 etcd[454701]: setting maximum number of CPUs to 1, total number of available CPUs is 1
Apr 10 14:15:26 PSQL_1 etcd[454701]: error listing data dir: /var/lib/etcd/cleanroom
Apr 10 14:15:26 PSQL_1 systemd[1]: etcd.service: Main process exited, code=exited, status=1/FAILURE
Apr 10 14:15:26 PSQL_1 systemd[1]: etcd.service: Failed with result 'exit-code'.
Apr 10 14:15:26 PSQL_1 systemd[1]: Failed to start etcd.service - etcd - highly-available key value store

To solve this problem, assign the right permissions to the directory with the following command:

1
chown -R etcd:etcd /var/lib/etcd/cleanroom