Scope and evidence

Source review
Cloud scope
AWS / Azure / GCP
Availability
Varies by feature
Exercise status
Exercises not run

Conditions and limits

  • Lakebase GA was verified separately for AWS and Azure; GCP remains Beta in the opened overview and supports us-east4, us-central1 and europe-west3.
  • Implementation settings use AWS project-based Lakebase documentation. Older provisioned instances and cloud-specific feature availability require their own checks.
  • AWS LTAP Direct Writes is GA for new synced tables on Postgres 16/17/18; Lakebase CDF is Public Preview. Account rollout and deployed regions were not tested.
  • The design, exercises and recovery plan are unexecuted; no latency, throughput, cost or failover measurement is claimed.

A recommendation and a reservation need different owners

A fictional returns application shows an agent an order's return-risk score, then lets a support operator reserve a refund. Yesterday's analytical model may suggest reviewing the return. The reservation must still prevent two operators from authorizing the same refund. A fresh prediction and a correct transaction answer different questions.

The Lakebase overview describes a managed Postgres database integrated with Databricks. This lesson uses that transactional role to design the reservation, while lakehouse data supplies derived context. Prerequisites are a basic understanding of table grain, primary keys, Delta tables and the platform boundaries in the earlier lessons. No account is needed for the paper exercise. Creating projects, running synchronization or testing failover is future work, not something this article has done.

Scroll horizontally to read the diagram.

Original schematic: Unity Catalog features flow to Lakebase serving tables; an application reads context and writes separate operational tables; optional LTAP Direct Writes bypass compute for bulk loads; a change feed emits operational history.

Original conceptual diagram, created for this lesson. It is a relationship map, not a product screenshot, measured architecture or promise of availability. Numbered paths are explained below; the table of cloud and feature status governs their scope.

Check the generation, cloud and feature before the walkthrough

“Lakebase is available” is too coarse for a design review. These claims were read on October 4, 2026:

Scope Verified state Design boundary
Project-based Lakebase on AWS GA entry dated January 22, 2026 Use current project/branch/endpoint guides; older provisioned instances have a different contract.
Lakebase on Azure GA entry dated March 2, 2026 An Azure release is checked separately from an AWS release.
Lakebase on GCP Overview says Beta as of June 15 The opened project guide lists us-east4, us-central1, europe-west3; the project uses the workspace region.
AWS Lakebase Search GA entry dated September 18, 2026 Search is a distinct capability; this exercise does not enable extensions or evaluate ranking.
AWS LTAP Direct Writes GA entry dated October 1, 2026 Postgres 16/17/18; opt in while creating a synced table.
AWS Lakebase CDF Current guide labels Public Preview Explicit preview enablement and destination requirements apply.

The dated entries come from the AWS release notes and Azure release notes; GCP state and regions come from its overview and project guide. CDF's scope comes from its feature guide. Releases can be staged. None of these documents proves that a particular account has received a feature, and AWS feature status is not asserted as GCP parity.

The relevant historical change is the AWS project-based generation: Beta in October 2025 and Public Preview in December 2025 preceded the 2026 GA. The previous-month review window is September 4–October 4, 2026, Asia/Tokyo calendar dates. Search GA and Direct Writes GA fall inside it. Dated entries establish those changes; they do not establish the publication date of the whole document. Source publication dates remain unknown when not established. Entries after October 4 are excluded.

For a visual walkthrough, Introduction to Lakebase: OLTP for Data Apps and AI Agents is published by Databricks (@Databricks) on December 23, 2025 and linked from the official product page. Use it as supplementary orientation; current feature guides determine this lesson's 2026 settings and availability. The video is linked with attribution; its images are not reproduced here.

Four paths, four meanings

Path 1: serve an analytical result. A Unity Catalog source such as analytics.returns.risk_by_order supplies a synced serving table. The synced-table guide describes a managed Lakebase copy and advises read queries on that destination. Snapshot refreshes the full data; Triggered and Continuous use source changes after the initial load. Pick a mode using the application's acceptable data age and source capabilities.

In our design, risk_by_order has one row per order and a model_version plus computed_at. The application can display “score computed at 09:00” instead of presenting any successful database query as a current model judgment. The writer remains the analytics pipeline. An operator's decision goes in a different table.

Path 2: commit operational state. refund_reservations is application-owned Postgres data. Its rows answer which request reserved which refund, who decided, and whether the reservation was released. A synced risk score may inform that choice; it does not own the reservation. This separation lets an analytical job change its recommendation without overwriting a business action.

Path 3: expose changes downstream. Lakebase CDF produces Unity Catalog managed Delta change history from Postgres writes. Its setup requires a suitable destination catalog, privileges and REPLICA IDENTITY FULL on source tables. Default-storage destinations are unsupported. The application commits its reservation first; a downstream BI model interprets emitted changes separately. A row-history feed is neither the reservation command nor a completed business metric.

Path 4: change how a bulk load reaches storage. LTAP is an architecture delivered through individual capabilities. Direct Writes changes the bulk-load route, not who owns the data. The sync guide applies it to initial loads in all modes and full Snapshot refreshes. Later Triggered/Continuous increments keep their change-feed path. It cannot be toggled onto an existing synced table; migration planning would be necessary, and no deletion is part of this exercise.

Keep authorization attached to the path. The LTAP guide distinguishes Unity Catalog analytical governance from Postgres roles and privileges for transactional access. A Unity Catalog grant should not be assumed to grant an application permission to update a Postgres reservation. Define the required reads and writes for each identity, then verify both permission boundaries.

Model the action before selecting compute

These are original teaching tables, not Databricks schemas or a deployed example:

Dataset One row means Identity and authority
risk_by_order One current analytical risk result per order Analytics owns order_id, score, model version and computation time.
refund_reservations One refund request and its current state Application owns non-null request_id, order_id, amount, currency and state version.
refund_decisions One recorded business state transition Application owns decision_id, actor, reason, decision time and the score/model version used.
BI refund fact One accepted reservation or agreed event grain Modeling owner defines which states count, restatements and currency treatment.

If an order can legitimately have multiple partial refunds, uniqueness on order_id alone rejects valid requests. If the business permits only one full refund, allowing arbitrary different request_id values without an order-level rule permits duplicates. The correct key follows the business invariant; it is not selected from the easiest column to index.

For this exercise, the invariant is one full-refund reservation per order, with one durable result per caller request. Retrying the same request with a changed amount must be rejected as a conflicting reuse of identity. A non-null request key, a stored request fingerprint and a durable outcome make this rule inspectable. PostgreSQL's constraints guide establishes how unique and primary-key constraints behave; the fingerprint and outcome policy here are an original application design.

An application check followed later by an insert is not the whole concurrency contract. In an unexecuted paper schedule, operators A and B both read “unreserved.” A reserves, then B writes using its earlier observation. The design must decide where the invariant is enforced when operations overlap. For a simple predetermined row, consider a conditional state transition or row locking with a short transaction. For a multi-row rule, consider the appropriate isolation and retry policy. PostgreSQL 17 isolation documentation explains statement snapshots at Read Committed and whole-transaction retries after serialization failure. Stronger isolation does not make retry handling optional.

Record the reservation and its decision evidence together. Treat any external payment as a separate operation with its own idempotency contract. A database transaction cannot prove an external refund happened exactly once. On a lost response after commit, return the stored result for the same request; blindly repeating the payment would turn a recoverable connection error into a business error.

Settings change capacity and failure behavior

The following settings are from AWS project-based guides, checked October 4. They are configuration limits, not recommended production values.

Setting Verified effect or limit Question it forces in this design
Minimum / maximum CU Autoscaling up to 64 CU; maximum − minimum ≤ 16 CU. Can the working set and first burst fit the minimum? What happens if load exceeds the maximum?
Scale-to-zero Eligible maximum is 32 CU or less; timeout 60 seconds–7 days, default 24 hours. Can the client tolerate reactivation and rebuild session state?
High availability One primary plus 1–3 secondaries; scale-to-zero is unavailable. Is reserved failover capacity worth keeping active for this workflow?
Connection capacity Autoscaling limits depend on the smaller of maximum CU and 8 × minimum CU; administrative connections are reserved separately. What is the total pool budget across application instances and other consumers?

These limits are supported respectively by autoscaling, scale-to-zero, HA and compute management. Changing configured CU boundaries may interrupt connections, even though automatic scaling within an established range does not require restarts. Suspending compute resets session context. HA covers compute failure within one region; existing connections need to reconnect after failover.

For an original configuration check, 2–8 CU spans 6 CU, while 2–32 spans 30 CU and violates the documented range limit. An 8–40 CU range fails both that span check and the scale-to-zero maximum check. This arithmetic establishes eligibility only. It provides no latency estimate or guarantee that a refund service will cope with a burst.

Suppose a hypothetical deployment has six application instances, each proposing a pool of 20 connections. That is 120 proposed client connections before migration tools and other consumers. Check the actual endpoint limit, then set one shared budget; autoscaling the web tier without a pool budget can multiply database pressure. These numbers are invented for planning and have not been applied to Lakebase.

When a reservation becomes slow, first separate lock waits, query work, pool waiting and connection recovery. Increasing maximum CU cannot resolve a conflicting business rule, and shortening idle timeout can increase reactivation frequency for intermittent traffic. Change one hypothesis at a time and record the relevant evidence.

Setup, alternatives and the cost ledger

An eventual sandbox setup would begin with the workspace cloud, region and generation. Create an isolated project/branch/database only after confirming availability and authorized costs. Define application roles separately from the synchronization identity; prepare the source schema and source-change support; choose the serving key and sync mode; create the serving table; then connect through the correct branch endpoint using the approved authentication method. The compute guide distinguishes endpoint resource names from connection hostnames, so do not build a hostname by guessing from a display name. No credentials or connection string are provided here.

Start with a written workload and compare actual alternatives:

Workload Candidate Tradeoff to evaluate
Large historical scan, aggregate or BI report Existing analytical tables and SQL compute Keeps the analytical model central; does not by itself implement the refund transaction.
Application lookup using derived data plus transactional state Lakebase serving table plus separate operational tables Convenient co-location of context and decisions; adds synchronization freshness, Postgres permission and operational capacity decisions.
Existing operational database already meets the requirement Retain it; design the necessary analytical exchange Avoids an unnecessary migration; compare integration ownership and duplicate-data handling.
Repeated read-only result with an accepted stale window An application cache, if the workload permits Can reduce database reads; cannot become authoritative for conflicting refund writes.

These are design candidates, not a measured ranking. Before adopting a new service, ask whether the present database and a clearer data contract solve the problem.

Keep a cost ledger containing active compute by branch and replica, stored data and indexes, retained backup/history or snapshots, synchronization work, and applicable transfer/connectivity charges. Some entries may belong to a different Databricks or cloud billing surface; verify the selected deployment's actual meter. Lakebase pricing explicitly separates pricing display from regional availability. The AWS release history records snapshot storage billing from June 1, 2026. Suspending compute does not mean every retained resource is free. No price quotation or cost comparison has been calculated here.

An offline fixture that makes failures visible

On paper, prepare two orders and these fictitious requests. Keep analytical score time, database commit time and downstream observation time in separate columns.

Input Expected business result in this exercise Evidence to retain
req-A, order 10, amount 60, first attempt One reservation Request identity, agreed payload, committed state and decision record
req-A, same payload, response lost then retried Same stored outcome No second reservation or external payment
req-A, amount changed to 80 Conflict Rejected payload mismatch tied to the original request
req-B, order 10, overlaps req-A Only one full-refund reservation succeeds Enforced order invariant; loser receives a defined result
Order 20 score changes after its reservation Historical action remains attributable Score/model version used at decision time and current score remain distinguishable

This fixture is unexecuted. Walk through both possible winners for overlapping requests. Next, remove a client response after commit, insert a stale score, and delay downstream change processing. Explain which state each component knows and which UI message is justified. “The sync is green” is not evidence that a refund was paid; “the app retried” is not evidence that it created two reservations.

A later authorized test should freeze cloud, region, Postgres version, feature state, schema, role grants, pool limits, CU bounds and sync settings. Test duplicate requests, conflicting payloads, two concurrent reservations, reconnect behavior and delayed downstream processing in isolated data. Observe p50/p95/p99 request latency, lock waits, pool wait, source-to-serving age, CDC backlog and the actual cost meters. Verify stored results after reconnect and inspect permission failures with a read-only identity. HA or fault injection needs its own bounded test plan; none has been run.

The useful mental model is the ownership map: analytics writes the recommendation, the application commits the reservation, and downstream modeling explains the committed history. Compute and synchronization settings tune those paths; they do not choose the business truth for you.

Sources

Publication dates belong to the source; access dates record when it was checked. Community observations are separate from official statements.

01
Official documentationDatabricks: Lakebase Postgres on AWS ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
02
Official documentationDatabricks: Lakebase release notes on AWS ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
03
Official documentationMicrosoft Learn: Lakebase release notes on Azure ↗learn.microsoft.comPublished: Unknown · Accessed: 2026-10-04
04
Official documentationDatabricks: Lakebase Postgres on Google Cloud ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
05
Official documentationDatabricks: Manage projects on Google Cloud ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
06
Official documentationDatabricks: Serve lakehouse data with synced tables ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
07
Official documentationDatabricks: LTAP architecture ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
08
Official documentationDatabricks: Lakebase Change Data Feed ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
09
Official documentationDatabricks: Lakebase autoscaling ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
10
Official documentationDatabricks: Lakebase scale to zero ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
11
Official documentationDatabricks: Lakebase high availability ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
12
Official documentationDatabricks: Manage Lakebase computes ↗docs.databricks.comPublished: Unknown · Accessed: 2026-10-04
13
Official documentationDatabricks: Lakebase pricing ↗www.databricks.comPublished: Unknown · Accessed: 2026-10-04
14
Official documentationPostgreSQL 17: Transaction isolation ↗www.postgresql.orgPublished: Unknown · Accessed: 2026-10-04
15
Official documentationPostgreSQL 17: Constraints ↗www.postgresql.orgPublished: Unknown · Accessed: 2026-10-04
16
Official videoDatabricks: Introduction to Lakebase: OLTP for Data Apps and AI Agents ↗www.youtube.comPublished: 2025-12-23 · Accessed: 2026-10-04
17
Official documentationDatabricks: Lakebase product page and official demos ↗www.databricks.comPublished: Unknown · Accessed: 2026-10-04
Saved in this browser only.