Observability, backups, and CDC
Diagnose metrics, logs, backups, restores, CDC apply, and webhook delivery.
An empty chart, disabled backup button, or empty webhook table can be correct for the current plan or state. Use the exact surface message and database profile before treating it as an outage.
Check the plan boundary first
| Surface | Shared Developer | Eligible dedicated database |
|---|---|---|
| Metrics | Project-level usage/operation metrics | Database and provider infrastructure metrics |
| Logs | Not available | Available for the active supported profile |
| Manual backups and restore | Not available | Available when role, provider metadata, workflow state, and backup compatibility checks pass |
| CDC configuration | Not available | Direct AMQP or Managed AMQP when the profile is eligible |
| Managed webhook history | Not available | Available for Managed AMQP deliveries |
Provider availability for customer deployments is determined by the current create catalog and active database profile. Do not infer availability from a provider name appearing in a generic capability description.
Before troubleshooting, record the organization slug, database name/UUID, plan, provider/region shown in Parix, profile state, selected time range, UTC time, exact UI message, and relevant workflow ID. Never include CDC, webhook, API, or provider secrets.
Metrics are empty or stop updating
Symptom. Charts contain no points, some host/TigerBeetle series are missing, or the last update does not advance.
Likely cause. The selected window has no activity, the profile is not active, you are viewing Shared Developer project metrics rather than host metrics, or provider guest/TigerBeetle telemetry is not yet available for this deployment.
Exact checks.
- Confirm the plan and profile state. Shared databases intentionally show project metrics; dedicated infrastructure metrics require an active profile.
- Change the range among 15 minutes, 1 hour, 3 hours, 6 hours, 12 hours, 24 hours, or 7 days/custom and note whether older points appear.
- Generate one safe read operation from Query, then wait for the metrics refresh and return focus to the page. Live metrics refresh about every 30 seconds and can also update from realtime invalidation.
- Read any telemetry notice on the page:
- a new GCP deployment can be waiting for guest telemetry, and TigerBeetle series can be waiting for the metrics bridge;
- an older AWS deployment can have partial runtime telemetry until it is reprovisioned/upgraded, and TigerBeetle series depend on its runtime metrics bridge.
- Distinguish a valid empty chart from an explicit loader/telemetry error.
Recovery. Select a range containing known activity and allow at least one refresh interval after a safe query. Finish any active provisioning workflow. If Parix displays a pending-telemetry notice, follow the offered upgrade/reprovision guidance or wait for activation rather than changing the application payload. Customers should not modify provider agents directly.
When to contact support. Escalate when an active eligible database has known operations but remains empty beyond multiple refreshes, or when a telemetry error persists. Include database UUID, provider/region, plan, missing chart names, selected UTC range, a safe test-operation time, screenshot, and exact message.
Logs show no rows or fail to load
Symptom. The Logs page says there are no logs for the range, or it displays a failure to load logs.
Likely cause. An empty result is valid when no matching events occurred. A loader error is a separate provider/query failure. Shared Developer does not expose the dedicated logs surface.
Exact checks.
- Confirm the database is on an eligible dedicated plan with an active profile.
- Read the empty/error copy exactly: No logs is different from Failed to load logs.
- Try 15 minutes, 1 hour, 6 hours, and 24 hours, using a UTC incident time to choose the window.
- For a live window with a returned cursor, leave the page open briefly; it refreshes about every five seconds and can receive realtime invalidation. A genuinely empty result has no cursor, so use realtime invalidation, Refresh, or a page reload instead of assuming periodic polling is active.
- Compare the incident time with provisioning/decommission state; a profile that did not exist at that time cannot have runtime logs.
Recovery. Widen or move the time range when the result is valid but empty. Refresh once after the profile becomes active. For an explicit loader failure, preserve the message and stop repeated polling; changing application code will not repair a provider log query.
When to contact support. Escalate a persistent Failed to load logs message or missing logs for a known event on an active profile. Include database UUID, provider/region, exact UTC range, UI message, and screenshot. Do not paste secrets or sensitive ledger payloads.
Manual backup is disabled or rejected
Symptom. Create backup is hidden/disabled, or a manual backup request is rejected before a workflow starts.
Likely cause. The plan/role is ineligible, the profile is not active, decommission is active, provider runtime metadata is incomplete, or another backup/conflicting workflow is running.
Exact checks.
- Confirm an eligible dedicated plan and owner/admin role; Shared Developer has no self-service backups.
- Confirm the profile is active and the decommission state is inactive.
- Check the Backups page and Cluster > Changes for an active backup, provision, restore, import, or decommission operation.
- Confirm provider/region and deployment metadata are populated on the database profile. Backup creation supports only a provider path offered for the active profile and fully configured in the current environment.
- Read the exact eligibility/error reason shown by the form.
Recovery. Wait for the conflicting workflow to finish, then refresh and submit one backup request. If profile/provider metadata is incomplete, do not guess or edit cloud resources; use the provider/profile recovery offered by Parix or contact support.
When to contact support. Escalate when an owner/admin cannot back up an active eligible profile with no conflicting work, or a backup workflow fails. Include database UUID, provider/region, backup/workflow ID, profile/decommission state, exact message, and UTC timeline.
Restore is disabled or rejected
Symptom. A completed backup cannot be selected/restored, confirmation is rejected, or the restore request reports an incompatibility.
Likely cause. Restore deliberately requires exact compatibility and a quiet database. The backup may lack artifacts/metadata, not match the active deployment, or target an unsupported provider/storage path.
Exact checks.
- Confirm owner/admin role, eligible dedicated plan, active profile, and inactive decommission state.
- Confirm the backup status is Completed and its required snapshot/file artifacts and restore metadata are present.
- Confirm the backup contains a complete replica set.
- Compare backup and active profile values exactly: region, node count, vCPU, memory, storage, network mode, development mode, and TigerBeetle version.
- Confirm no backup, provision, restore, import, or decommission workflow is active.
- Confirm provider/storage eligibility. GCP restore follows its supported snapshot path; for an existing AWS profile, restore is limited to the supported Local NVMe path—non-Local-NVMe AWS restore is rejected.
- Enter the database name exactly in the destructive confirmation field.
Recovery. Restore replaces the current database state, discards writes made after the selected recovery point, cannot merge data, and cannot be canceled once it starts. Review Backup and restore, confirm the recovery point, then select a completed compatible backup for the unchanged target shape and retry once after conflicting workflows finish. Do not bypass metadata, replica, or version checks. If the desired target shape differs, use a supported migration/new-database path from Provisioning and configuration rather than forcing an in-place restore.
When to contact support. Escalate when displayed metadata matches every eligibility check but restore remains unavailable, or a restore workflow fails. Include database UUID, backup ID, workflow ID, provider/storage backend, compatibility values, exact message, and UTC time. Do not attach backup artifacts unless given an approved secure transfer path.
CDC configuration is queued or in error
Symptom. CDC remains queued/running, reaches error, or the save response says changes were queued but the apply workflow failed.
Likely cause. Parix persisted the desired CDC configuration before the apply workflow was enqueued, or the provider workflow could not apply the validated settings. The saved configuration can therefore remain queued even when no apply workflow started.
Exact checks.
- Confirm dedicated plan, owner/admin role, and a profile that is not provisioning.
- Read the CDC status and exact save/apply message.
- Open Cluster > Changes and locate the CDC workflow by time/actor. Determine whether it is absent, Pending, In progress, Completed, or Failed.
- For Direct AMQP, recheck literal public IP, port, username, virtual host, exchange, routing key, TLS, and publish-confirm selections. Do not reveal the password.
- For Managed AMQP, confirm at least one active destination, a literal public-IP HTTPS URL on port 443, no embedded credentials/redirect dependency, and a secret supplied on first enable.
- Confirm the provider/profile is still one offered and supported for the active database.
Recovery. If the workflow is running, wait for it to finish. If the configuration is queued but no workflow exists, preserve the desired state and exact enqueue error; do not keep clicking apply. If a terminal error names an invalid/reachable setting, correct that one setting and submit one new apply request. Leaving an existing secret field blank retains its stored value.
When to contact support. Escalate a persisted queued configuration with no workflow, a non-progressing apply, or a terminal provider error after inputs are corrected. Include database UUID, CDC mode/status, destination names only, provider/region, workflow ID if present, exact message, and UTC timeline. Never include passwords or authentication/signing secrets.
See Webhooks and CDC for the configuration and delivery contract.
Managed webhook history is empty
Symptom. The Webhooks page has no rows, even though CDC configuration exists.
Likely cause. Managed AMQP is not active, no new CDC events entered the managed path, filters exclude the attempts, destinations are inactive, or retained history has aged/rolled out. Direct AMQP never creates managed webhook history.
Exact checks.
- Confirm the CDC mode is Managed AMQP and status is
active—not Direct AMQP, queued, running, disabled, or error. - Confirm at least one destination is active.
- Set status to All, destination to All, and range to the smallest interval containing a known event; then widen through 24 hours, 7 days, or 14 days.
- Select Latest to leave an older cursor page.
- Confirm a TigerBeetle change occurred after Managed AMQP became active.
Recovery. Activate/apply Managed AMQP first, wait for active, then generate only a safe intended test change and refresh Latest. Do not expect pre-activation events or Direct AMQP deliveries to appear. The UI's 14-day selector does not guarantee all events remain retained for that duration.
When to contact support. Escalate when a known post-activation event produces no attempt for any active destination. Include database UUID, CDC status, event time/ID if known, destination name, filter values, provider/region, and screenshot—never the destination secret.
Managed webhook delivery fails or repeats
Symptom. Attempts show Failed, the receiver sees timeouts/non-2xx responses, or a destination receives the same event more than once.
Likely cause. The endpoint violates egress policy, cannot be reached, exceeds the 15-second timeout, redirects, rejects authentication/signature, or returns non-2xx. Duplicates are expected under retry: when one destination fails, Parix retries the event as a whole, so destinations that already succeeded can receive it again.
Exact checks.
- Select the event in Webhooks and record event ID, attempt ID/time, destination, queue attempt, HTTP status, failure reason, and error text.
- Interpret the failure category:
egress_policy_error: URL is not an allowed public literal-IP HTTPS target;build_request_error: destination/auth request could not be constructed;network_error: connection/TLS/network failure;timeout: no accepted response within 15 seconds;http_error: redirect or non-2xx response.
- Search receiver logs by
x-parix-event-id, not only by request time. - For static-header auth, confirm the configured header name and current secret at the receiver.
- For HMAC, calculate
sha256=<hex>over<x-parix-signature-ts>.<raw-request-body>using the stored secret. Do not parse/re-serialize JSON before verification. - Confirm the receiver durably records the event ID before returning 2xx and treats a recognized completed ID as success.
Recovery. Correct the public literal-IP HTTPS endpoint, certificate/listener, authentication, HMAC calculation, or receiver latency/status. Make processing idempotent by event ID before re-enabling a failing destination. Return 2xx only after durable acceptance. Do not add redirect endpoints or assume exactly-once delivery.
When to contact support. Escalate when Parix records a network/build/egress failure that contradicts verified endpoint behavior, or attempts stop before the documented retry boundary without success. Include database UUID, destination name, event/attempt IDs, UTC times, queue attempt, HTTP status/failure category, and sanitized receiver logs. Never send endpoint secrets or sensitive payload contents.