“The reports of my death are greatly exaggerated,” said Mark Twain, upon reading his own obituary. Data warehouses have existed in some form for at least 30 years and continue to play an important role in managing data for all kinds of organisations. And yet, with each new development in data management technology, people (us included) are liable to ask: “are data warehouses dead?” Like Mark Twain, we think the data warehouse’s death has surely been greatly exaggerated, and that it’s therefore worthwhile investigating its would-be murderers. In doing so we’ve concluded that — unlike Mark Twain — the data warehouse is very much alive and well.
What can you typically expect from a data warehouse?
A data warehouse is a centralised repository for:
- Aggregating data from multiple systems and sources.
- Storing data in a cleansed, transformed, and catalogued form.
- Staging data in a way that makes it easy to run queries, reports, dashboards and other kinds of data analytics.
- Running queries, reports, dashboards and other kinds of data analytics without placing additional load on operational systems.
- Maintaining historical data that operational systems do not.
While we see the data warehouse as a key part of a modern data management solution, there are competing technologies that have, at times, purported to do it. In this, the nerdiest of all games of Cluedo, the wannabe assassins include:
- The data lake and its lovechild with the data warehouse, the data lakehouse.
- Data virtualisation.
- Data fabric.
Data lakes and lakehouses
An important step in getting data into a data warehouse is transforming it into an ideal format. Sometime in the early 2010s, data analysts looked at the effort involved in doing this and thought to themselves, “how about we don’t?” Thus was the data lake born. Though they still aggregate data in a central location, data lakes store data in its native format, that is, the format that it took at its source.
Data Warehouse and Data Lake

Data lakes, like data warehouses, aggregate data from multiple systems and sources, but with substantially less effort. However, as they do not force data into a schema or establish the relationships between disparate datasets, it is more difficult to generate reports and dashboards from a data lake than a data warehouse.
Data lakes can also promote a philosophy towards data ingestion best described as “just bung it in”. Because it is easy to load new data into a data lake, users tend to load more and more data without worrying about how it should be organised or future usefulness. This can lead to data quality and accessibility issues, and a data lake turning into what is commonly referred to as a data swamp.
For these reasons the data lake never killed the data warehouse. Instead, they became complementary components of an ideal data management solution. In fact, the data lake and data warehouse spawned a child with potentially oedipal tendencies: the data lakehouse.
The data lakehouse, as marketed by companies like Databricks and AWS, add data warehouse capabilities to data lake object storage. They add a metadata layer to a data lake that, in theory, results in a data store that can handle both structured and unstructured data, and supports the kind of queries, reporting, and cataloguing that data warehouses do so well. Databricks encourages organising lakehouses according to a “medallion architecture” with three layers: bronze, silver, and gold. The structure and quality of data should improve from raw in the bronze layer, to aggregated and clean in the gold layer.
Data Lakehouse

Despite their promise, data lakehouses remain a relatively new architectural approach. Best practices are still being developed, and there are concerns about vendor lock-in and the availability of resources with related skillsets. So, for now at least, their claim to have killed the data warehouse feels weak.
Data virtualisation
Data virtualisation attempts to do many of the same things as a data warehouse (or indeed a data lakehouse), but without aggregating data in a centralised repository. The idea is to give users an interface that allows them to interact with the data as if it were all in one place, while behind the scenes it remains spread across disparate source systems and databases.
Doing things this way offers a few advantages over data warehousing:
- It’s relatively quick to deploy because it doesn’t involve building complex data integrations.
- It presents users with near-real time data, whereas data warehouses are only as current as the last data load (often daily).
- It adapts easily when data sources change.
At this point you could be forgiven for asking, is that it then? Has data virtualisation murdered the data warehouse? But no, of course it hasn’t. Because data virtualisation relies on continuous access to data sources, using it to run large and complex queries places heavy loads on operational systems. This can lead to latency and disruptions for staff and customers.
Once again, data virtualisation makes most sense as a complement to the data warehouse. They offer a potentially quicker and cheaper way to generate ad hoc and operational reporting, while more complex reporting and analytics is left to data warehouses.
Data Fabric
At the time of writing, data fabric is more a buzzword than a practical approach to data management. In some respects, data fabric builds on data virtualisation. Data virtualisation relies on metadata about the systems and sources in which data resides. As users find new ways to combine data using data virtualisation, metadata can be captured about the kinds of queries, views and reports they create. What defines a data fabric is the use of this metadata to identify and automate data management activities that users might find useful.
Imagine something like this:
- Staff at Novigi often use a data virtualisation tool to combine data from Slack (our instant messaging system) with data from Tactic (our desk booking system) and data from the Bureau of Meteorology to determine the perfect timing for sharing a hilarious meme with their coworkers.
- Artificial intelligence identifies that Novigi staff are combining these data sources regularly to address this very important business problem.
- Automation tooling automatically builds the data integration and transformations to make this combined data available in the data warehouse.
- Employee satisfaction at Novigi reaches record highs.
While this sounds amazing in theory, we are not aware of any organisation that has managed to build this capability at an enterprise scale. Data fabric may end up being the next big thing in data management, but for now its main role is to give data management professionals something to talk about at conferences. Needless to say, it hasn’t killed the data warehouse.
Long live the data warehouse
Despite the many attempts on its life, the data warehouse isn’t going anywhere anytime soon. New approaches to data management have so far proved complementary to data warehousing rather than supplanting it. Artificial intelligence may shake things up, but at present we think it’s more likely to make some of the challenges inherent in data warehousing — like difficult data integrations — less challenging. Organisations considering implementing a data warehouse can do so without worrying too much that they’ll regret it in the foreseeable future.
This article was produced as part of The Quarterly – Q1 FY25
For more information about anything you’ve read here, or if you have a more general inquiry, please contact us.
Key Contributors:

Kevin Fernandez is General Manager, Market Strategy and Propositions at Novigi, and is based in the Melbourne office.

Sophie Bowen-James is an analyst in the Market Strategy and Propositions team at Novigi, and is based in the Sydney office.
