Skip to content

Troubleshooting the PostgreSQL cluster

Info

The whole procedure must be followed as root. To elevate your privileges to root, you can use the su - command.

Warning

To follow this procedure, you must have carried out the additional configurations for patronictl.

Connection errors to the database of the CyberElements Bastion PostgreSQL cluster can have various causes. This section presents a general troubleshooting method, to be adapted to your case.

Checking the cluster status

Info

Checking the cluster status quickly locates the failing node or nodes.

You can check the cluster status with the following command, to be run on each node of the PostgreSQL cluster:

1
patronictl -c /etc/patroni/config.yml topology
Example of output on a working platform

1
2
3
4
5
6
7
+ Cluster: 15-cleanroomvault5 ------+---------+---------+----+-----------+
| Member           | Host           | Role    | State   | TL | Lag in MB |
+------------------+----------------+---------+---------+----+-----------+
| PSQL_3           | psql_3         | Leader  | running | 2  |           |
| + PSQL_1         | psql_1         | Replica | running | 2  |         0 |
| + PSQL_2         | psql_2         | Replica | running | 2  |         0 |
+------------------+----------------+---------+---------+----+-----------+
In the output above, everything indicates that the PostgreSQL cluster is working: the lag is 0 MB for all Replica nodes, and all nodes are in the list, with the running status. You need to get this output on all nodes to confirm that patroni and etcd are working.

Example of output when there is a communication problem between the nodes

1
2
3
4
5
6
7
+ Cluster: 15-cleanroomvault5 ------+---------+---------+----+-----------+
| Member           | Host           | Role    | State   | TL | Lag in MB |
+------------------+----------------+---------+---------+----+-----------+
| PSQL_3           | psql_3         | Leader  | running | 2  |           |
| + PSQL_1         | psql_1         | Replica | running | 2  |       230 |
| + PSQL_2         | psql_2         | Replica | running | 2  |         0 |
+------------------+----------------+---------+---------+----+-----------+
In the output above, the 230 MB lag is abnormal: it could indicate a communication problem between node 3 and node 1.
If the problem persists after several hours, this confirms it.

Example of output when a node is not properly started or stopped

1
2
3
4
5
6
7
+ Cluster: 15-cleanroomvault5 ------+---------+---------+----+-----------+
| Member           | Host           | Role    | State   | TL | Lag in MB |
+------------------+----------------+---------+---------+----+-----------+
| PSQL_3           | psql_3         | Leader  | running | 2  |           |
| + PSQL_1         | psql_1         | Replica | started | 2  |   unknown |
| + PSQL_2         | psql_2         | Replica | stopped | 2  |   unknown |
+------------------+----------------+---------+---------+----+-----------+
In the output above, the state (State column) indicates that node 1 is starting. If the started status persists for several minutes, or even hours, there may be a problem with patroni or etcd on that node.
Node 2 is in the stopped state: it is stopped. When a cluster node fails, this information is only displayed temporarily, before the node is removed from the list.

Example of output when etcd fails on the current node

1
2
3
2026-04-09 16:21:39,482 - WARNING - Retrying (Retry(total=1, connect=None, read=None, redirect=0, status=None)) after connection broken by 'NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7fcaa1618390>: Failed to establish a new connection: [Errno 111] Connection refused')': /version
2026-04-09 16:21:39,482 - WARNING - Retrying (Retry(total=0, connect=None, read=None, redirect=0, status=None)) after connection broken by 'NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7fcaa1618c10>: Failed to establish a new connection: [Errno 111] Connection refused')': /version
2026-04-09 16:21:39,483 - ERROR - Failed to get list of machines from https://PSQL_3:2379/v2: MaxRetryError("HTTPSConnectionPool(host='psql_3', port=2379): Max retries exceeded with url: /version (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7fcaa1619490>: Failed to establish a new connection: [Errno 111] Connection refused'))")
The output above indicates a connection error to etcd, which prevents retrieving the status of the PostgreSQL cluster.

Checking the status of the services

Once the failing node has been identified, you need to identify the malfunctioning service. The two main services used on the PostgreSQL cluster are patroni and etcd.

Info

The patroni service depends on the etcd service: if etcd is not working, patroni will not work either.

To identify the failing service, you can run the following commands:

1
2
systemctl status etcd
systemctl status patroni

These commands give the status of the service, along with its last 10 log lines.

Warning

A service with the active (running) status is not necessarily working: it can be running while writing error logs. It is therefore important to read the service logs to understand its actual state.

Example of output for etcd when the service is working
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
● etcd.service - etcd - highly-available key value store
   Loaded: loaded (/lib/systemd/system/etcd.service; enabled; preset: enabled)
   Active: active (running) since Thu 2026-04-09 16:40:01 CEST; 21h ago
   Docs: https://etcd.io/docs
           man:etcd
Main PID: 283425 (etcd)
   Tasks: 7 (limit: 2227)
   Memory: 73.2M
       CPU: 51min 9.223s
   CGroup: /system.slice/etcd.service
           └─283425 /usr/bin/etcd

Apr 10 09:26:31 PSQL_1 etcd[283425]: store.index: compact 1624026
Apr 10 09:26:31 PSQL_1 etcd[283425]: finished scheduled compaction at 1624026 (took 325.098µs)
Apr 10 10:26:31 PSQL_1 etcd[283425]: store.index: compact 1624386
Apr 10 10:26:31 PSQL_1 etcd[283425]: finished scheduled compaction at 1624386 (took 417.797µs)
Apr 10 11:26:31 PSQL_1 etcd[283425]: store.index: compact 1624746
Apr 10 11:26:31 PSQL_1 etcd[283425]: finished scheduled compaction at 1624746 (took 1.706489ms)
Apr 10 12:26:31 PSQL_1 etcd[283425]: store.index: compact 1625106
Apr 10 12:26:31 PSQL_1 etcd[283425]: finished scheduled compaction at 1625106 (took 300.098µs)
Apr 10 13:26:31 PSQL_1 etcd[283425]: store.index: compact 1625466
Apr 10 13:26:31 PSQL_1 etcd[283425]: finished scheduled compaction at 1625466 (took 311.398µs)
Example of output for etcd when the service is in error

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
× etcd.service - etcd - highly-available key value store
    Loaded: loaded (/lib/systemd/system/etcd.service; enabled; preset: enabled)
    Active: failed (Result: exit-code) since Fri 2026-04-10 14:15:26 CEST; 5s ago
Duration: 21h 35min 20.382s
    Docs: https://etcd.io/docs
            man:etcd
    Process: 454701 ExecStart=/usr/bin/etcd $DAEMON_ARGS (code=exited, status=1/FAILURE)
Main PID: 454701 (code=exited, status=1/FAILURE)
        CPU: 12ms

Apr 10 14:15:26 PSQL_1 etcd[454701]: [WARNING] Deprecated '--logger=capnslog' flag is set; use '--logger=zap' flag instead
Apr 10 14:15:26 PSQL_1 etcd[454701]: etcd Version: 3.4.23
Apr 10 14:15:26 PSQL_1 etcd[454701]: Git SHA: Not provided (use ./build instead of go build)
Apr 10 14:15:26 PSQL_1 etcd[454701]: Go Version: go1.19.8
Apr 10 14:15:26 PSQL_1 etcd[454701]: Go OS/Arch: linux/amd64
Apr 10 14:15:26 PSQL_1 etcd[454701]: setting maximum number of CPUs to 1, total number of available CPUs is 1
Apr 10 14:15:26 PSQL_1 etcd[454701]: error listing data dir: /var/lib/etcd/cleanroom
Apr 10 14:15:26 PSQL_1 systemd[1]: etcd.service: Main process exited, code=exited, status=1/FAILURE
Apr 10 14:15:26 PSQL_1 systemd[1]: etcd.service: Failed with result 'exit-code'.
Apr 10 14:15:26 PSQL_1 systemd[1]: Failed to start etcd.service - etcd - highly-available key value store.
In the output above, the etcd service has the failed status, and the error that prevented it from starting is the one on line 17: the service cannot read the content of the /var/lib/etcd/cleanroom directory. In this case, adjust the permissions on the directory.

Example of output for patroni when the service is working
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
● patroni.service - Runners to orchestrate a high-availability PostgreSQL
   Loaded: loaded (/lib/systemd/system/patroni.service; enabled; preset: enabled)
   Drop-In: /etc/systemd/system/patroni.service.d
           └─local.conf
   Active: active (running) since Thu 2026-04-09 16:40:06 CEST; 22h ago
   Process: 283510 ExecStartPre=/usr/bin/testetcd.py (code=exited, status=0/SUCCESS)
Main PID: 283517 (patroni)
   Tasks: 13 (limit: 2227)
   Memory: 56.1M
       CPU: 26min 54.918s
   CGroup: /system.slice/patroni.service
           ├─283517 /usr/bin/python3 /usr/bin/patroni /etc/patroni/config.yml
           ├─283537 /usr/lib/postgresql/15/bin/postgres -D /var/lib/postgresql/15/cleanroomvault5 --config-file=/etc/postgresql/15/cleanroomvault5/postgresql.conf"--listen_addresses=*" --port=5432 --cluster_name=15-cleanroomvault5 --wal_level=replica --hot_standby=on --max_connections=100 --max_wal_senders=10--max_prepared_transactions=0 --max_locks_per_transaction=64 --track_commit_timestamp=off --max_replication_slots=10 --max_worker_processes=8 --wal_log_hints=on
          ├─283538 "postgres: 15-cleanroomvault5: logger "
          ├─283540 "postgres: 15-cleanroomvault5: checkpointer "
           ├─283541 "postgres: 15-cleanroomvault5: background writer "
          ├─283542 "postgres: 15-cleanroomvault5: startup recovering 00000003000000000000000F"
          ├─283545 "postgres: 15-cleanroomvault5: walreceiver "
          └─283547 "postgres: 15-cleanroomvault5: postgres postgres [local] idle"
Apr 10 14:51:07 PSQL_1 patroni[283517]: 2026-04-10 14:51:07,381 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:51:17 PSQL_1 patroni[283517]: 2026-04-10 14:51:17,428 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:51:27 PSQL_1 patroni[283517]: 2026-04-10 14:51:27,381 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:51:37 PSQL_1 patroni[283517]: 2026-04-10 14:51:37,428 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:51:47 PSQL_1 patroni[283517]: 2026-04-10 14:51:47,381 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:51:57 PSQL_1 patroni[283517]: 2026-04-10 14:51:57,428 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:52:07 PSQL_1 patroni[283517]: 2026-04-10 14:52:07,381 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:52:17 PSQL_1 patroni[283517]: 2026-04-10 14:52:17,428 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:52:27 PSQL_1 patroni[283517]: 2026-04-10 14:52:27,381 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:52:37 PSQL_1 patroni[283517]: 2026-04-10 14:52:37,474 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Example of output for patroni when the service is in error

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
● patroni.service - Runners to orchestrate a high-availability PostgreSQL
    Loaded: loaded (/lib/systemd/system/patroni.service; enabled; preset: enabled)
    Drop-In: /etc/systemd/system/patroni.service.d
            └─local.conf
    Active: active (running) since Thu 2026-04-09 16:40:06 CEST; 23h ago
    Process: 283510 ExecStartPre=/usr/bin/testetcd.py (code=exited, status=0/SUCCESS)
Main PID: 283517 (patroni)
    Tasks: 12 (limit: 2227)
    Memory: 61.7M
        CPU: 28min 8.337s
    CGroup: /system.slice/patroni.service
            ├─283517 /usr/bin/python3 /usr/bin/patroni /etc/patroni/config.yml
            ├─283537 /usr/lib/postgresql/15/bin/postgres -D /var/lib/postgresql/15/cleanroomvault5 --config-file=/etc/postgresql/15/cleanroomvault5/postgresql.conf "--listen_addresses=*" --port=5432 --cluster_name=15-cleanroomvault5 --wal_level=replica --hot_standby=on --max_connections=100 --max_wal_senders=10 --max_prepared_transactions=0 --max_locks_per_transaction=64 --track_commit_timestamp=off --max_replication_slots=10 --max_worker_processes=8 --wal_log_hints=on
            ├─283538 "postgres: 15-cleanroomvault5: logger "
            ├─283540 "postgres: 15-cleanroomvault5: checkpointer "
            ├─283541 "postgres: 15-cleanroomvault5: background writer "
            ├─283542 "postgres: 15-cleanroomvault5: startup recovering 00000003000000000000000F"
            └─283547 "postgres: 15-cleanroomvault5: postgres postgres [local] idle"

Apr 10 15:52:37 PSQL_1 patroni[283517]: 2026-04-10 15:52:37,680 WARNING: Loop time exceeded, rescheduling immediately.
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:37,681 INFO: Lock owner: PSQL_2; I am PSQL_1
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,069 ERROR: Request to server https://PSQL_3:2379 failed: ReadTimeoutError("HTTPSConnectionPool(host='psql_3', port=2379): Read timed out. (read timeout=3.3331721344341836)")
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,069 INFO: Reconnection allowed, looking for another server.
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,069 INFO: Retrying on https://PSQL_1:2379
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,073 ERROR: Request to server https://PSQL_1:2379 failed: MaxRetryError("HTTPSConnectionPool(host='psql_1', port=2379): Max retries exceeded with url: /v3/lease/keepalive (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7fd77c411090>: Failed to establish a new connection: [Errno 111] Connection refused'))")
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,073 INFO: Reconnection allowed, looking for another server.
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,073 INFO: Retrying on https://PSQL_2:2379
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,074 ERROR: Request to server https://PSQL_2:2379 failed: MaxRetryError("HTTPSConnectionPool(host='psql_2', port=2379): Max retries exceeded with url: /v3/lease/keepalive (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7fd77c411090>: Failed to establish a new connection: [Errno 111] Connection refused'))")
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,074 INFO: Reconnection allowed, looking for another server.
In the case above, the service has the active (running) status without actually working: the logs show various connection errors to etcd, which patroni needs.
In this example, patroni is not necessarily at fault: since it depends on etcd, first fix the etcd problems, then check the status of patroni again.

Service troubleshooting procedures

Below is one troubleshooting procedure per service. If both etcd and patroni are malfunctioning, start with etcd.