Data Warehouse vs Lakehouse vs Data Mart: What Mid-Market Businesses Actually Need

The moment a business decides its reporting needs a proper foundation, three vendors arrive with three different nouns: one proposes a data warehouse, one a lakehouse, one says a mart will do. All three demos look similar. The prices do not.
Meanwhile, the estate already has an accidental answer: a Power BI model that has quietly grown into a warehouse-shaped object - integrating six sources, holding years of history, governed by nobody - and taking four hours to refresh.
This guide untangles the three tiers, maps the platforms onto them, and gives the decision signals that say which one your business actually needs.
Key Takeaways
The tiers differ by job: marts serve a subject area, warehouses integrate and govern, lakehouses converge structured and raw at scale.
The signals are measurable: source count, volume, history needs - and the four-hour-refresh threshold where models need foundations.
Stage the path: mart to warehouse to lakehouse without rebuilds, sized to the next three years rather than the vendor's quota.
What Does Each Tier Actually Do?
A data mart is a curated slice for one subject area - sales, finance, operations - small, fast and cheap. A warehouse is the governed integration layer where multiple sources are conformed into one modelled truth with history. A lakehouse stores everything - structured tables and raw files - in open formats one engine family can both govern and analyse. The tiers are jobs, not sizes.
The confusion arises because each tier can imitate the others briefly: a mart can absorb a second subject area, a warehouse can store files awkwardly, a Power BI model can impersonate a mart until it cannot. The imitation works right up until the workload reveals the mismatch - usually as cost, refresh time or governance gaps.
Where Do Fabric, Snowflake and BigQuery Fit?
All three are platforms on which any tier can be built - the tier is your architecture, the platform is where it runs. Microsoft Fabric leans lakehouse-first with OneLake at the centre and warehouse semantics on top; Snowflake grew from warehouse excellence toward lake workloads; BigQuery is a serverless warehouse whose economics suit spiky analytical loads and whose ecosystem anchors Google-centric estates.
The selection logic for mid-market Australian businesses is usually gravitational rather than technical: estates living in Microsoft 365 and Power BI minimise friction on Microsoft Fabric; marketing-heavy estates already exporting GA4 to BigQuery extend what they have; genuinely multi-cloud or data-sharing-heavy businesses lean Snowflake. The wrong move is choosing the platform before the tier - buying scale you do not need on the strength of a demo.

What Are the Signals That Say Which Tier You Need?
Count sources, measure volume, and check the refresh clock. The signals are unglamorously concrete:
One to three sources, one subject, gigabytes: a mart - or honestly, a well-modelled Power BI semantic model with proper foundations.
Four-plus sources needing conformed definitions, history that must survive system changes, multiple consuming teams: a warehouse. The tell is reconciliation pain - when "customer" and "product" need one ruling across systems, that ruling needs a home.
Large or semi-structured volumes (events, telemetry, files), data science alongside BI, or refresh windows that defeat copying: a lakehouse, where data lands once in open format and engines come to it.
The universal threshold: a Power BI dataset taking hours to refresh is a model doing a foundation's job - the signal that the integration work needs to move upstream regardless of which tier it moves to.
What Is the Accidental Warehouse - and Why Is It a Trap?
It is the Power BI model that became the integration layer by accretion: source after source merged in Power Query, business rules embedded in measures, history retained because deleting felt risky. It works - until it is the slowest, least governed, most irreplaceable object in the company, owned by whoever last edited it.
The trap is that it postpones the foundation decision while compounding its cost: every month adds logic that must eventually migrate, and the model's fragility taxes every change. The escape is not dramatic - it is promoting the model's integration logic upstream into a real tier, table by table, while the reports keep running. Recognising the accidental warehouse early is half the value of an architecture review; most businesses are surprised to learn they already built one.
How Do You Stage the Path Without Rebuilds?
Build each tier as the durable layer beneath the previous one, never as a replacement beside it. The mart's tables become the warehouse's first subject area; the warehouse's modelled layer becomes tables in the lakehouse when scale demands it. Open formats and a dimensional design are what make the staging possible - the modelling survives every platform move even when the storage changes.
The sequencing that works for mid-market: first, lift the accidental warehouse's logic into a governed transformation layer (the reports notice nothing except faster refreshes); second, conform the shared dimensions - customer, product, calendar - because they are the asset every future consumer reuses; third, expand by subject area in priority order, each one a bounded project with a named consumer. Sized this way, the foundation arrives in funded slices rather than a monolithic programme - and each slice pays back before the next begins.
What Should You Refuse to Buy?
Scale without a workload, and platforms without a tier decision. The lakehouse pitched to a business with three sources and clean gigabytes is capacity rented for prestige; the warehouse programme scoped as a two-year monolith is risk dressed as rigour. Equally, refuse the opposite economy - the fourth year of working around the accidental warehouse costs more than the foundation it is avoiding.
The honest sizing question is three years long: where will sources, volume and consumers be in three years, and what is the smallest tier that carries that - built in slices that can each justify themselves? Vendors size to their quota; you size to your trajectory.
How Long Does Each Tier Take to Stand Up?
A focused data mart - one domain, one or two sources, a dimensional model and the reports on top - is a four-to-eight-week build, which is why it is the right first move for most mid-market organisations. A genuine warehouse spanning core systems is a quarter-by-quarter programme: a domain per quarter, each shipping usable reporting, rather than a single eighteen-month big bang.
The schedule risk is rarely the technology - provisioning Fabric, Snowflake or BigQuery takes a day. It is source access, definitional agreement and data quality discovery, which is why the staged path matters: each delivered domain retires its share of those risks while producing value, instead of accumulating them all behind one distant go-live date.
The Tier Is a Decision, Not a Destiny
Warehouse, lakehouse and mart are answers to different questions - integration, scale and focus - and most mid-market businesses need them in sequence, not in competition. Read your own signals, name the accidental warehouse if you have one, and stage the path so each layer becomes the foundation of the next.
The architecture that results is unfashionably boring: the right tier, one size larger than today, built in slices. Boring is what foundations are supposed to be.



