September 28, 2006

Relational data warehouse Expansion (or Explosion) Ratios

One of the least understood aspects of data warehouse technology is what may be called the

Expansion Ratio = (Total disk space used, except for mirroring) / (Size of the base database).

This is similar to the explosion ratio discussed in the OLAP Report’s justly famous discussion of database explosion, but I’m going with my own terminology because I don’t want to be tied to their precise terminology, nor to their technical focus. Expansion Ratios are hotly debated, with some figures being:

Teradata claims an Expansion Ratio of 8-9X for Oracle, 6X for DB2 (open system version), and 2.5X for Teradata. The underlying source is data warehouses they’ve replaced, so there may be a bias toward out-of-control warehouses on the part of their competitors.
An anonymous appliance vendor exec said to me off the top of his head that Oracle has 6-8X Expansion Ratios.
Oracle’s TPC-H submissions in the largest size range (10 terabytes) have 9.7-10.5X Expansion Ratios, if I’m reading the TPCs correctly.
Oracle cites a survey of 8 customers with 10-60 Tb database size in which the Expansion Ratio works out to 1.6X. (More on this anomalous result below.)

I don’t have actual figures from Netezza and DATallegro, but I imagine they’d come out lower than 2X, possibly well below.

Categories: Data warehouse appliances, Data warehousing, Database compression, DATAllegro, IBM and DB2, Netezza, Oracle, Teradata

9 Comments

September 27, 2006

Logless, lockless Netezza more carefully explained

I talked at length with Bill Blake and Doug Johnson of Netezza today. (Bill is exactly the guy I complained of previously having had my access cut off to.) One takeaway was a clarification of their approach to transactions, which sounds even cooler than I first thought. It’s actually not a new idea; they just timestamp rows with CreateIDs and DeleteIDs, then exploit those to the hilt. Actually, it seems like this approach would be interesting in OTLP as well, although I’m not aware of it being used in any of the more successful OLTP DBMS systems. (Yes, this is an open invitation to fans of less-established DBMS products to tell me of their virtues, preferably in a flame-free manner.)
Read more

Categories: Data warehouse appliances, Netezza

5 Comments

September 27, 2006

Oracle and Microsoft in data warehousing

Most of my recent data warehouse engine research has been with the specialists. But over the past couple of days I caught up with Oracle and Microsoft (IBM is scheduled for Friday). In at least three ways, it makes sense to lump those vendors together, and contrast them with the newer data warehouse appliance startups:

Shared-everything architecture
End-to-end solution story
OLTP industrial-strengthness carried over to data warehousing

In other ways, of course, their positions are greatly different. Oracle may have a full order-of-magnitude lead on Microsoft in warehouse sizes, for example, and has a broad range of advanced features that Microsoft either hasn’t matched yet, or else just released in SQL Server 2005. Microsoft was earlier in pushing DBA ease as a major product design emphasis, although Oracle has played vigorous catch-up in Oracle10g.

Categories: Data warehouse appliances, DATAllegro, EAI, EII, ETL, ELT, ETLT, IBM and DB2, Microsoft and SQL*Server, Netezza, Oracle, Parallelization, Teradata

1 Comment

September 24, 2006

More on data warehouse architecture choices

The very name of this blog comes from the kind of “horses for courses” data store strategy implied by my recent post on different kinds of data warehouse uses. A number of other commentators have recently made similar points, although they may not agree with every detail. For example, William McKnight pretty much makes the pure DBMS2 argument, pointing out that a partially virtual warehouse is often superior to a fully centralized physical one. And Andy Hayler of Kalido says pretty much the same thing, although he strongly calls out his difference in emphasis from William’s view.

A tip of the hat to Mark Rittman for pointing me to those two and others.

Categories: Data warehouse appliances, EAI, EII, ETL, ELT, ETLT, Theory and architecture

Data warehouse and mart uses – a tentative taxonomy

I’ve been posting a lot recently about the diverse database technologies used to support data warehousing. With the marketplace supporting such a broad range of architectures, it seems clear that a lot of those architectures actually deserve to thrive, presumable each in a different kind of usage scenario. So in this post I’ll take a pass at dividing up use cases for data warehouses, and suggesting which kinds of data warehouse management technologies might do the best job of supporting them. To start with, I’ve divided things into a number of buckets:

Pinpoint data lookup
Constrained query and reporting
Cube-filling calculations
Hardcore tabular data crunching
Text and media search
Specialty areas, such as relationship analytics

Categories: Data warehouse appliances, Data warehousing, DATAllegro, IBM and DB2, MOLAP, Netezza, Teradata

1 Comment

September 22, 2006

My blogs stopped working through IE!

EDIT: Now they seem to be working again, with no action on my part and no known software updates through the whole process. Go figure. I do not know WordPress well enough to guess just exactly what had to have been broken and then fixed at my hosting provider to have caused these effects.

As of this writing, my blogs (DBMS2, the Monash Report, Text Technologies, and Software Memories) are all working in Firefox, and the top page of each is working in IE, but the rest of the pages/links are NOT working in IE. (But www.monash.com, a non-Wordpress site on the same host, is still working through IE.) Naturallly, I’m addressing this problem as fast as I can. I imagine the fix will involve some sort of a reinstall and/or theme change, which could alter the blogs’ look-and-feel, maybe not for the better (especially at first). I apologize for the inconvenience!

Categories: About this blog

1 Comment

September 22, 2006

Competitive issues in data warehouse ease of administration

The last person I spoke with at the Netezza conference on Tuesday was a customer/presenter that the company had picked out for me. One thing he said baffled me — he claimed that Netezza was a real appliance vendor, but DATallegro wasn’t, presumably due to administrability issues. Now, it wasn’t clear to me that he’d ever evaluated DATallegro, so I didn’t take this too seriously, but still the exchange brought into focus the great differences between data warehouse products in the area of administration. For example:

Netezza has no indices at all. And no caches. And the hardware is preconfigured. This all makes administration pretty simple.
DATallegro has almost no indices, and also has preconfigured hardware. But it has some partitioning, optionally.
Teradata also has preconfigured hardware. It does have indices, but rather simple ones. Plus it has join indices. And it has a few more configuration options in other areas (e.g., block size) than the other appliance vendors. (Yes, I count Teradata among the appliances.)
If you go through all the fuss of installing SAP’s applications and BI technology anyway, the incremental administration of just SAP BI Accelerator is pretty light.
Oracle and IBM have mammothly complex indexing options, but have put large amounts of work into tools to lessen the resulting administrative burden.
IBM offers preconfigured hardware units to simplify some installation issues.
Come to think of it, I don’t really know how hard it is to administer columnar systems (e.g., Sybase IQ).

Categories: Data warehouse appliances, Data warehousing, DATAllegro, Greenplum, IBM and DB2, Netezza, Oracle, SAP AG, Teradata

3 Comments

September 20, 2006

SAP’s BI Accelerator

I wrote about SAP’s BI Accelerator quite a bit in my white paper on memory-centric data management, but otherwise I seem not to have posted much about it here. In essence, it’s a product that’s all RAM-based, and generally geared for multi-hundred-gigabyte data marts. The basic design is a compression-heavy column-based architecture, evolved from SAP’s text-indexing technology TREX. Like data warehouse appliances, it eschews indexing, relying instead on blazingly fast table scans.

I asked Lothar Schubert of SAP how BIA was doing in the market in its early going. This was his response:

Categories: Analytic technologies, Business intelligence, Data warehouse appliances, Data warehousing, Database compression, Memory-centric data management, SAP AG

8 Comments

September 20, 2006

Myths about DATallegro, Ingres, open source, etc.

Sometimes, when one talks to a company about a close competitor, what one hears may not be 100% strictly accurate. Yesterday, I more than once heard claims that sounded oddly like “DATallegro has to open source whatever software it develops.” Today, DATallegro CEO Stuart Frost clarified as follows:

• DATallegro has no (little?) legal obligation to open source anything. Even the version of Ingres they use is not the GPL one.
• They do give a few enhancements back to Ingres (via open source?) rather than maintain them themselves.
• The whole MPP technology is proprietary, in every sense of “proprietary.” (For example, they use a whole different optimizer than Ingres’s. I’ve forgotten whether the Ingres optimizer is also left in place.)

Categories: Actian and Ingres, Data warehouse appliances, DATAllegro, Memory-centric data management, Open source

1 Comment

September 20, 2006

Teradata vs. the new appliance vendors, technically

Todd Walter and Randy Lea of Teradata gave generously of their time today, ducking out of their user conference, and shared their take on issues we’ve been discussing here recently. Overall, Teradata response to the data warehouse appliance guys is essentially: “Well, those may be fine for specific queries, or for data marts, but in true blended enterprise data warehouse workloads we’re superior, including in performance.”

Specific takeaways included:

Categories: Data warehouse appliances, DATAllegro, Netezza, Teradata

4 Comments

Monash Research blogs

DBMS 2 covers database management, analytics, and related technologies.
Text Technologies covers text mining, search, and social software.
Strategic Messaging analyzes marketing and messaging strategy.
The Monash Report examines technology and public policy issues.
Software Memories recounts the history of the software industry.

User consulting

Building a short list? Refining your strategic plan? We can help.

Vendor advisory

We tell vendors what's happening -- and, more important, what they should do about it.

Monash Research highlights

Learn about white papers, webcasts, and blog highlights, by RSS or email.

Links
- Monash Research
- White Papers
Admin
- Log in

Relational data warehouse Expansion (or Explosion) Ratios

Logless, lockless Netezza more carefully explained

Oracle and Microsoft in data warehousing

More on data warehouse architecture choices

Data warehouse and mart uses – a tentative taxonomy

My blogs stopped working through IE!

Competitive issues in data warehouse ease of administration

SAP’s BI Accelerator

Myths about DATallegro, Ingres, open source, etc.

Teradata vs. the new appliance vendors, technically

Search our blogs and white papers

Monash Research blogs

User consulting

Vendor advisory

Monash Research highlights

Recent posts

Categories

Date archives

Admin