HOME/BLOG/technology
technology

The PostgreSQL Upgrade That Became Boring on Purpose

Explore PostgreSQL upgrades, compatibility checks, rehearsal and safe cutovers with clear ownership and reliable rollback.

AUTHOR: Kedar•03 December 2025•5 MIN READ
The PostgreSQL Upgrade That Became Boring on Purpose

DATABASE MODERNISATION

The PostgreSQL Upgrade That Became Boring on Purpose

A migration diary about compatibility discovery, rehearsal and a rollback boundary everyone understood.

The production cutover finished before midnight, but the decisive work happened weeks earlier: an extension was replaced, a long-running query was fixed, disk headroom was increased and the application reconnect storm was reproduced in a rehearsal.

Start with one workload | Scale into a complete platform | Open Source Freedom | Enterprise-Grade Operations.

Published | December 2025

The work before the maintenance window

The production cutover finished before midnight, but the decisive work happened weeks earlier: an extension was replaced, a long-running query was fixed, disk headroom was increased and the application reconnect storm was reproduced in a rehearsal.

Version upgrades fail when teams treat the binary switch as the project. The real upgrade includes data, extensions, configuration, replicas, drivers, connection pools and the period during which rollback remains possible.

The Upgrade Window Starts Months Before Cutover.

The turning point

The team made the route visible: discover, choose in-place or parallel, rehearse with representative data, cut over behind gates, validate business journeys, then close the rollback window deliberately.

The team used OSS Manager to make planned PostgreSQL minor and major version upgrades reviewable and repeatable. The engine still owned its native behaviour; the control plane made intent, prerequisites, execution and evidence visible to the people responsible for the service.

What had to become explicit

  • The service boundary: OSS Manager selects an in-place or parallel major-upgrade strategy after compatibility discovery, rehearses it with production-like data, controls cutover, and preserves a tested rollback boundary.
  • The operational controls: version path, extension compatibility, backup and restore proof, pg_upgrade checks, replica strategy, connection drain, cutover gate, rollback deadline, and application tests.
  • The capacity conversation: database and index size, free disk, link or copy mode, WAL and backup throughput, replica rebuild time, validation duration, and rollback storage.
  • The human boundary: who may observe, who may approve, who may execute, and who decides whether the application is ready.

The architecture that changed the conversation

The diagram is deliberately centred on the decision the team had to make. It is not a product inventory. It shows where authority sits, what crosses the boundary, and where a failed assumption must stop the workflow.


The practical design choices

  • Compatibility. Stop condition: Unsupported extension or driver. Decision owner: DBA + application.
  • Cutover. Stop condition: Validation or timing exceeds limit. Decision owner: Change authority.
  • Post-upgrade. Stop condition: Critical journey regresses. Decision owner: Business owner.
  • Rehearsal. Stop condition: Runtime exceeds the approved window. Decision owner: Database and platform leads.
  • Rollback closure. Stop condition: Business proof remains incomplete. Decision owner: Change authority and service owner.

These choices are intentionally small enough to review and test. They keep the architecture tied to operating behaviour instead of allowing a visually impressive diagram to hide unclear ownership.

What OSS Manager changed - and what it did not

OSS Manager brought discovery, planning, guarded execution and normalised status into one path for PostgreSQL. It did not replace the engine's correctness model or the application's responsibility for data semantics.

  • Before change, validate versions, hosts, identities, artifacts, topology and the recovery boundary without mutating the service.
  • During change, persist progress, expose stop conditions and prevent a partial result from being mistaken for success.
  • After change, observe upgrade duration; service availability; replication state; query latency; error rate; WAL growth; extension load; connection recovery; post-upgrade bloat and statistics and run representative application journeys, not only process checks.
  • For recovery, protect the last known-good authority and require an explicit decision before promotion, rollback or destructive cleanup.

The failure modes worth rehearsing

The test plan should make room for unsupported upgrade path, incompatible extension, insufficient disk, failed pg_upgrade check, partial replica rebuild, connection storm, query regression, and expired rollback window. The purpose is not to produce a longer checklist. It is to learn whether operators can recognise the failure and choose the safe next action while the evidence is incomplete.

A human operating model

L1 operators need a plain-language answer to what is healthy, what is delayed and whether a change is in progress. Platform engineers need topology and evidence. Application owners need to know what users will experience. A useful platform connects those views without giving every person the same privileges.

A quiet cutover was not luck. It was the visible result of moving uncertainty out of the maintenance window.

Where this pattern earns its place

Retail

This pattern is relevant where teams need campaign peaks, catalogue freshness, transaction continuity and rapid rollback. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.

Telecom

This pattern is relevant where teams need high event volume, distributed operations, identity boundaries and service assurance. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.

Gaming

This pattern is relevant where teams need latency-sensitive sessions, leaderboards, event streams and regional player experience. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.

The lesson we would carry into the next project

A quiet cutover was not luck. It was the visible result of moving uncertainty out of the maintenance window.

The strongest open-source platforms are not the ones with the most automation. They are the ones where automation makes ownership, risk and recovery easier for people to understand.

Start with one workload. Prove the operating model. Then scale it into a complete platform.