Perspective
The Real Switching Cost in Modern Data Platforms
Open table formats solved data movement, not management.
July 24, 2026 · Perspective · Leon Liang
Most comparisons of Databricks, AWS, and Google Cloud are framed as a bake-off: which platform is fastest, which is cheapest, or who has the better AI story. This framing persists right up until you try to source the actual numbers.
There is a more useful question, and the last two years have finally made it answerable. All three platforms now support the same open table formats, meaning your data is genuinely portable in a way it wasn’t in 2023.
What remains non-portable is the governance model wrapped around that data. This is where the real switching cost lives. It is rarely marketed because it is the one thing you cannot take with you.
The current landscape
Terminology changes quickly. To avoid stale vocabulary, here is what these platforms actually are today:
Databricks is a multicloud platform operating across all three hyperscalers. Its current stack includes Delta Lake and Photon at the base, Unity Catalog for governance, and Lakeflow for ingestion and orchestration (which integrated Delta Live Tables as Lakeflow Declarative Pipelines and Workflows as Lakeflow Jobs on 12 June 2025). It also features Databricks SQL, Mosaic AI, AI/BI Genie, and Lakebase—a serverless Postgres born from the Neon acquisition.
AWS has repositioned around SageMaker Unified Studio, which reached GA on 13 March 2025. Crucially, Unified Studio “brings together functionality” from EMR, Glue, Athena, Redshift, Bedrock, and SageMaker AI; it does not replace them. These services still exist and are billed separately. DataZone also remains the underlying substrate for SageMaker’s domain and governance model.
Google Cloud centers on BigQuery, using BigLake for open-format tables. Its catalog layer has undergone frequent renaming: Data Catalog became Dataplex Universal Catalog, which was renamed again to Knowledge Catalog on 10 April 2026. Any comparison mentioning “Dataplex” is likely outdated.
A period of rapid consolidation
The last fourteen months of platform strategy
-
AWS ships SageMaker Unified Studio
One workspace over EMR, Glue, Athena, Redshift and Bedrock.
-
Databricks acquires Neon
Becomes Lakebase. Ghodsi cites agents provisioning most Neon databases.
-
Lakeflow reaches GA
Delta Live Tables and Workflows fold into one brand.
-
Google renames Dataplex Universal Catalog to Knowledge Catalog
-
Databricks moves on security, agreeing to acquire Panther
-
Databricks raises at a $188B valuation
Coatue-led, per WSJ and Reuters. Up from $134B in December.
Where the lock-in actually lives
By mid-2026, all three vendors supported the Apache Iceberg REST Catalog protocol. AWS introduced Glue catalog federation for remote Iceberg catalogs in November 2025, allowing queries of external catalogs (including Unity Catalog) without moving data. Google’s BigLake metastore and Databricks’ open-sourced Unity Catalog (June 2024) follow the same path.
Cross-engine reads now work. The format war is over; open formats won.
However, the protocol does not transport:
- Row and column-level permissions
- Masking policies
- Data lineage
- Credential vending
- Audit logs
Each vendor manages these on top of the shared metadata layer, and they do not map one-to-one. Unity Catalog ACLs, Lake Formation tag-based access control, and Knowledge Catalog policies are overlapping but distinct models.
Consequently, migrating between platforms is no longer a data migration—it is a governance reimplementation. This is significantly harder to estimate because it isn’t a volume problem. You cannot calculate a timeline based on terabytes; you must re-derive and prove exactly who is allowed to see what.
Note one asymmetry: Databricks-managed Iceberg tables require Unity Catalog. Write support for external engines into these tables remains in “public preview” rather than GA in Databricks’ own documentation. Across the board, read-out is better supported than write-in.
Why you should distrust benchmarks
The 2021 TPC-DS incident illustrates the pattern for most vendor benchmarks.
Databricks announced an audited TPC-DS record (which was legitimate), but separately published an unaudited, TPC-DS-derived comparison claiming massive advantages over Snowflake. Snowflake’s founders responded by calling it a “marketing stunt lacking integrity.” They re-ran the test using a warehouse half the size Databricks had claimed was necessary, reporting a cost of $267 against the $1,791 Databricks had attributed to them.
Two parties, one workload, a sevenfold difference
What Snowflake's cost was said to be, by each interested party.
| Item | Reported cost for the same workload |
|---|---|
The TPC’s fair-use policy explicitly states that derived benchmarks are not comparable to official results. Yet, the “Databricks is many times cheaper” narrative still circulates, tracing back to this disowned comparison rather than the audited record.
Currently, there is no independent, methodologically rigorous benchmark of Databricks against BigQuery and Redshift at realistic scale. Fivetran provides the closest effort, though they are transparent about their limits (1TB scale, sequential single-stream execution, no tuning) and are themselves a vendor in the ecosystem.
The structural impossibility of pricing comparisons
Each platform meters different units.
- Databricks bills DBUs (starting at $0.15 for data engineering, $0.22 for warehousing, and $0.40 for interactive workloads), plus a separate CU rate for Lakebase. In classic deployments, this is on top of the cloud infrastructure costs you pay the provider.
- Redshift Serverless bills RPU-hours at $0.375.
- Athena bills $5 per TB scanned.
- BigQuery offers per-TB-scanned or slot-based capacity, with a choice between logical and physical storage billing.
There is no exchange rate between a DBU, an RPU, and a slot. A DBU is a proxy for compute on a specific instance; a slot is a unit of parallel query capacity decoupled from storage. Any comparison that produces a single “winner” has made assumptions about workload shape, concurrency, and tuning that essentially pre-determine the conclusion.
We searched for a credible, independently funded TCO study. None exist. Most pricing comparison sites are marketing-adjacent and disclose no methodology.
One structural detail to remember: in classic deployments, a Databricks bill is effectively two invoices—the DBU charge and the underlying cloud spend. This is the most common cost complaint from teams new to the platform.
A practical decision framework
Stop asking which platform is “best.” Ask these questions instead:
- Where does your data already live, and what is the egress cost to move it? This usually outweighs any feature advantage.
- Which governance model can your team actually operate? Unity Catalog is the most coherent option if you are all-in on Databricks, but it also assumes you will stay there.
- What is the shape of your workload? Heavy Spark-based transformation and ML align naturally with Databricks. SQL-first analytics on data already in BigQuery has little reason to move. These aren’t benchmark claims; they are observations on where friction is lowest.
- Who is going to run it on Tuesday? A platform your team can manage at 2am is superior to one that scored better in a whitepaper written by someone who isn’t on call.
Aeolus view — Startups and scale-ups often overthink this. Your data is likely already in one cloud and your identity provider is already there. The marginal capability difference between these platforms won’t determine if your AI initiative succeeds. What will determine success is whether your metrics are defined once and your pipelines are trustworthy. This work is platform-independent. Pick the one your team can run, keep your tables in Iceberg to ensure portability, and spend your energy on the governance model. That is the only part you would actually have to rebuild.
A note on the numbers
Because Databricks is private, its figures are self-reported. Its December 2025 announcement cited a revenue run rate above $4.8 billion (55% YoY growth), over 20,000 organizations, and a $134 billion valuation. Later reports from CNBC (June 2026) and other outlets cite a $6.9 billion run rate and a $188 billion valuation; while these sources are credible, the data remains company-disclosed rather than filed.
For a more transparent comparison, refer to audited public filings:
- Snowflake: $1.33 billion product revenue (quarter ending 30 April 2026), up 34%, with 126% net revenue retention.
- AWS: $37.6 billion in Q1 2026, up 28%.
- Google Cloud: $24.8 billion for Q2 2026, up 82%.
Finally, be wary of statistics claiming “X percent of enterprises have adopted a lakehouse.” Every such figure we traced originated from vendor-commissioned surveys. There is no neutral adoption data available.
Where this leaves you
The “winner” table you were looking for would be built on unsourcable numbers. Instead, we can say this with confidence: the data format question is settled and portable, the governance question is not, and the cost question cannot be answered generically.
If you are choosing a platform or suspect you chose the wrong one, we can help you evaluate the situation. If your data isn’t ready for this yet, we’ll help you get there.
Want a second opinion on your data stack?
Every Aeolus engagement starts with a fixed-fee data & AI-readiness audit — a short, low-risk first step before any larger build.
Book a data & AI-readiness audit