Resolução de problemas do cluster PostgreSQL
Info
Todo o procedimento deve ser realizado como root. Para elevar os seus privilégios a root, pode utilizar o comando su -.
Os erros de ligação à base de dados do cluster PostgreSQL do CyberElements Bastion podem ter várias causas. Esta secção apresenta um método geral de resolução de problemas, a adaptar ao seu caso.
Verificar o estado do cluster
Info
Verificar o estado do cluster permite localizar rapidamente o nó ou os nós em falha.
Pode verificar o estado do cluster com o seguinte comando, a executar em cada nó do cluster PostgreSQL:
| patronictl -c /etc/patroni/config.yml topology
|
Exemplo de resultado numa plataforma funcional
| + Cluster: 15-cleanroomvault5 ------+---------+---------+----+-----------+
| Member | Host | Role | State | TL | Lag in MB |
+------------------+----------------+---------+---------+----+-----------+
| PSQL_3 | psql_3 | Leader | running | 2 | |
| + PSQL_1 | psql_1 | Replica | running | 2 | 0 |
| + PSQL_2 | psql_2 | Replica | running | 2 | 0 |
+------------------+----------------+---------+---------+----+-----------+
|
No resultado acima, tudo indica que o cluster PostgreSQL está a funcionar: o lag é de 0 MB em todos os nós Replica e todos os nós constam da lista, com o estado running. É necessário obter este resultado em todos os nós para confirmar que o patroni e o etcd estão a funcionar.
Exemplo de resultado quando existe um problema de comunicação entre os nós
| + Cluster: 15-cleanroomvault5 ------+---------+---------+----+-----------+
| Member | Host | Role | State | TL | Lag in MB |
+------------------+----------------+---------+---------+----+-----------+
| PSQL_3 | psql_3 | Leader | running | 2 | |
| + PSQL_1 | psql_1 | Replica | running | 2 | 230 |
| + PSQL_2 | psql_2 | Replica | running | 2 | 0 |
+------------------+----------------+---------+---------+----+-----------+
|
No resultado acima, o lag de 230 MB é anómalo: pode indicar um problema de comunicação entre o nó 3 e o nó 1.
Se o problema persistir após várias horas, tal confirma-o.
Exemplo de resultado quando um nó não foi corretamente iniciado ou parado
| + Cluster: 15-cleanroomvault5 ------+---------+---------+----+-----------+
| Member | Host | Role | State | TL | Lag in MB |
+------------------+----------------+---------+---------+----+-----------+
| PSQL_3 | psql_3 | Leader | running | 2 | |
| + PSQL_1 | psql_1 | Replica | started | 2 | unknown |
| + PSQL_2 | psql_2 | Replica | stopped | 2 | unknown |
+------------------+----------------+---------+---------+----+-----------+
|
No resultado acima, o estado (coluna State) indica que o nó 1 está a iniciar. Se o estado started persistir durante vários minutos, ou mesmo horas, pode existir um problema com o patroni ou o etcd nesse nó.
O nó 2 está no estado stopped: está parado. Quando um nó do cluster falha, esta informação só é apresentada temporariamente, antes de o nó ser removido da lista.
Exemplo de resultado quando o etcd falha no nó atual
| 2026-04-09 16:21:39,482 - WARNING - Retrying (Retry(total=1, connect=None, read=None, redirect=0, status=None)) after connection broken by 'NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7fcaa1618390>: Failed to establish a new connection: [Errno 111] Connection refused')': /version
2026-04-09 16:21:39,482 - WARNING - Retrying (Retry(total=0, connect=None, read=None, redirect=0, status=None)) after connection broken by 'NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7fcaa1618c10>: Failed to establish a new connection: [Errno 111] Connection refused')': /version
2026-04-09 16:21:39,483 - ERROR - Failed to get list of machines from https://PSQL_3:2379/v2: MaxRetryError("HTTPSConnectionPool(host='psql_3', port=2379): Max retries exceeded with url: /version (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7fcaa1619490>: Failed to establish a new connection: [Errno 111] Connection refused'))")
|
O resultado acima indica um erro de ligação ao etcd, que impede a obtenção do estado do cluster PostgreSQL.
Verificar o estado dos serviços
Depois de identificado o nó em falha, é necessário identificar o serviço que não está a funcionar. Os dois principais serviços utilizados no cluster PostgreSQL são o patroni e o etcd.
Info
O serviço patroni depende do serviço etcd: se o etcd não estiver a funcionar, o patroni também não funcionará.
Para identificar o serviço em falha, pode executar os seguintes comandos:
| systemctl status etcd
systemctl status patroni
|
Estes comandos indicam o estado do serviço, bem como as suas 10 últimas linhas de log.
Atenção
Um serviço com o estado active (running) não está necessariamente a funcionar: pode estar em execução e escrever logs de erro. Por isso, é importante ler os logs do serviço para compreender o seu estado real.
Exemplo de resultado para o etcd quando o serviço está a funcionar
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22 | ● etcd.service - etcd - highly-available key value store
Loaded: loaded (/lib/systemd/system/etcd.service; enabled; preset: enabled)
Active: active (running) since Thu 2026-04-09 16:40:01 CEST; 21h ago
Docs: https://etcd.io/docs
man:etcd
Main PID: 283425 (etcd)
Tasks: 7 (limit: 2227)
Memory: 73.2M
CPU: 51min 9.223s
CGroup: /system.slice/etcd.service
└─283425 /usr/bin/etcd
Apr 10 09:26:31 PSQL_1 etcd[283425]: store.index: compact 1624026
Apr 10 09:26:31 PSQL_1 etcd[283425]: finished scheduled compaction at 1624026 (took 325.098µs)
Apr 10 10:26:31 PSQL_1 etcd[283425]: store.index: compact 1624386
Apr 10 10:26:31 PSQL_1 etcd[283425]: finished scheduled compaction at 1624386 (took 417.797µs)
Apr 10 11:26:31 PSQL_1 etcd[283425]: store.index: compact 1624746
Apr 10 11:26:31 PSQL_1 etcd[283425]: finished scheduled compaction at 1624746 (took 1.706489ms)
Apr 10 12:26:31 PSQL_1 etcd[283425]: store.index: compact 1625106
Apr 10 12:26:31 PSQL_1 etcd[283425]: finished scheduled compaction at 1625106 (took 300.098µs)
Apr 10 13:26:31 PSQL_1 etcd[283425]: store.index: compact 1625466
Apr 10 13:26:31 PSQL_1 etcd[283425]: finished scheduled compaction at 1625466 (took 311.398µs)
|
Exemplo de resultado para o etcd quando o serviço está em erro
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20 | × etcd.service - etcd - highly-available key value store
Loaded: loaded (/lib/systemd/system/etcd.service; enabled; preset: enabled)
Active: failed (Result: exit-code) since Fri 2026-04-10 14:15:26 CEST; 5s ago
Duration: 21h 35min 20.382s
Docs: https://etcd.io/docs
man:etcd
Process: 454701 ExecStart=/usr/bin/etcd $DAEMON_ARGS (code=exited, status=1/FAILURE)
Main PID: 454701 (code=exited, status=1/FAILURE)
CPU: 12ms
Apr 10 14:15:26 PSQL_1 etcd[454701]: [WARNING] Deprecated '--logger=capnslog' flag is set; use '--logger=zap' flag instead
Apr 10 14:15:26 PSQL_1 etcd[454701]: etcd Version: 3.4.23
Apr 10 14:15:26 PSQL_1 etcd[454701]: Git SHA: Not provided (use ./build instead of go build)
Apr 10 14:15:26 PSQL_1 etcd[454701]: Go Version: go1.19.8
Apr 10 14:15:26 PSQL_1 etcd[454701]: Go OS/Arch: linux/amd64
Apr 10 14:15:26 PSQL_1 etcd[454701]: setting maximum number of CPUs to 1, total number of available CPUs is 1
Apr 10 14:15:26 PSQL_1 etcd[454701]: error listing data dir: /var/lib/etcd/cleanroom
Apr 10 14:15:26 PSQL_1 systemd[1]: etcd.service: Main process exited, code=exited, status=1/FAILURE
Apr 10 14:15:26 PSQL_1 systemd[1]: etcd.service: Failed with result 'exit-code'.
Apr 10 14:15:26 PSQL_1 systemd[1]: Failed to start etcd.service - etcd - highly-available key value store.
|
No resultado acima, o serviço etcd tem o estado failed e o erro que o impediu de iniciar é o da linha 17: o serviço não consegue ler o conteúdo do diretório /var/lib/etcd/cleanroom. Neste caso, ajuste as permissões do diretório.
Exemplo de resultado para o patroni quando o serviço está a funcionar
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29 | ● patroni.service - Runners to orchestrate a high-availability PostgreSQL
Loaded: loaded (/lib/systemd/system/patroni.service; enabled; preset: enabled)
Drop-In: /etc/systemd/system/patroni.service.d
└─local.conf
Active: active (running) since Thu 2026-04-09 16:40:06 CEST; 22h ago
Process: 283510 ExecStartPre=/usr/bin/testetcd.py (code=exited, status=0/SUCCESS)
Main PID: 283517 (patroni)
Tasks: 13 (limit: 2227)
Memory: 56.1M
CPU: 26min 54.918s
CGroup: /system.slice/patroni.service
├─283517 /usr/bin/python3 /usr/bin/patroni /etc/patroni/config.yml
├─283537 /usr/lib/postgresql/15/bin/postgres -D /var/lib/postgresql/15/cleanroomvault5 --config-file=/etc/postgresql/15/cleanroomvault5/postgresql.conf"--listen_addresses=*" --port=5432 --cluster_name=15-cleanroomvault5 --wal_level=replica --hot_standby=on --max_connections=100 --max_wal_senders=10--max_prepared_transactions=0 --max_locks_per_transaction=64 --track_commit_timestamp=off --max_replication_slots=10 --max_worker_processes=8 --wal_log_hints=on
├─283538 "postgres: 15-cleanroomvault5: logger "
├─283540 "postgres: 15-cleanroomvault5: checkpointer "
├─283541 "postgres: 15-cleanroomvault5: background writer "
├─283542 "postgres: 15-cleanroomvault5: startup recovering 00000003000000000000000F"
├─283545 "postgres: 15-cleanroomvault5: walreceiver "
└─283547 "postgres: 15-cleanroomvault5: postgres postgres [local] idle"
Apr 10 14:51:07 PSQL_1 patroni[283517]: 2026-04-10 14:51:07,381 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:51:17 PSQL_1 patroni[283517]: 2026-04-10 14:51:17,428 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:51:27 PSQL_1 patroni[283517]: 2026-04-10 14:51:27,381 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:51:37 PSQL_1 patroni[283517]: 2026-04-10 14:51:37,428 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:51:47 PSQL_1 patroni[283517]: 2026-04-10 14:51:47,381 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:51:57 PSQL_1 patroni[283517]: 2026-04-10 14:51:57,428 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:52:07 PSQL_1 patroni[283517]: 2026-04-10 14:52:07,381 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:52:17 PSQL_1 patroni[283517]: 2026-04-10 14:52:17,428 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:52:27 PSQL_1 patroni[283517]: 2026-04-10 14:52:27,381 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
Apr 10 14:52:37 PSQL_1 patroni[283517]: 2026-04-10 14:52:37,474 INFO: no action. I am (PSQL_1), a secondary, and following a leader (PSQL_2)
|
Exemplo de resultado para o patroni quando o serviço está em erro
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29 | ● patroni.service - Runners to orchestrate a high-availability PostgreSQL
Loaded: loaded (/lib/systemd/system/patroni.service; enabled; preset: enabled)
Drop-In: /etc/systemd/system/patroni.service.d
└─local.conf
Active: active (running) since Thu 2026-04-09 16:40:06 CEST; 23h ago
Process: 283510 ExecStartPre=/usr/bin/testetcd.py (code=exited, status=0/SUCCESS)
Main PID: 283517 (patroni)
Tasks: 12 (limit: 2227)
Memory: 61.7M
CPU: 28min 8.337s
CGroup: /system.slice/patroni.service
├─283517 /usr/bin/python3 /usr/bin/patroni /etc/patroni/config.yml
├─283537 /usr/lib/postgresql/15/bin/postgres -D /var/lib/postgresql/15/cleanroomvault5 --config-file=/etc/postgresql/15/cleanroomvault5/postgresql.conf "--listen_addresses=*" --port=5432 --cluster_name=15-cleanroomvault5 --wal_level=replica --hot_standby=on --max_connections=100 --max_wal_senders=10 --max_prepared_transactions=0 --max_locks_per_transaction=64 --track_commit_timestamp=off --max_replication_slots=10 --max_worker_processes=8 --wal_log_hints=on
├─283538 "postgres: 15-cleanroomvault5: logger "
├─283540 "postgres: 15-cleanroomvault5: checkpointer "
├─283541 "postgres: 15-cleanroomvault5: background writer "
├─283542 "postgres: 15-cleanroomvault5: startup recovering 00000003000000000000000F"
└─283547 "postgres: 15-cleanroomvault5: postgres postgres [local] idle"
Apr 10 15:52:37 PSQL_1 patroni[283517]: 2026-04-10 15:52:37,680 WARNING: Loop time exceeded, rescheduling immediately.
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:37,681 INFO: Lock owner: PSQL_2; I am PSQL_1
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,069 ERROR: Request to server https://PSQL_3:2379 failed: ReadTimeoutError("HTTPSConnectionPool(host='psql_3', port=2379): Read timed out. (read timeout=3.3331721344341836)")
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,069 INFO: Reconnection allowed, looking for another server.
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,069 INFO: Retrying on https://PSQL_1:2379
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,073 ERROR: Request to server https://PSQL_1:2379 failed: MaxRetryError("HTTPSConnectionPool(host='psql_1', port=2379): Max retries exceeded with url: /v3/lease/keepalive (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7fd77c411090>: Failed to establish a new connection: [Errno 111] Connection refused'))")
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,073 INFO: Reconnection allowed, looking for another server.
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,073 INFO: Retrying on https://PSQL_2:2379
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,074 ERROR: Request to server https://PSQL_2:2379 failed: MaxRetryError("HTTPSConnectionPool(host='psql_2', port=2379): Max retries exceeded with url: /v3/lease/keepalive (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x7fd77c411090>: Failed to establish a new connection: [Errno 111] Connection refused'))")
Apr 10 15:52:41 PSQL_1 patroni[283517]: 2026-04-10 15:52:41,074 INFO: Reconnection allowed, looking for another server.
|
No caso acima, o serviço tem o estado active (running) sem estar realmente a funcionar: os logs mostram vários erros de ligação ao etcd, de que o patroni necessita.
Neste exemplo, o patroni não é necessariamente o culpado: uma vez que depende do etcd, resolva primeiro os problemas do etcd e, em seguida, volte a verificar o estado do patroni.
Procedimentos de resolução de problemas dos serviços
Segue-se um procedimento de resolução de problemas por serviço. Se tanto o etcd como o patroni estiverem em falha, comece pelo etcd.