A Survey and Classification of Storage Deduplication Systems

Pereira, José Orlando; Paulo, João

Citation:: Paulo J, Pereira JO. 2014. A Survey and Classification of Storage Deduplication Systems. ACM Computing Surveys. 47(1):1–30.

Abstract:

The automatic elimination of duplicate data in a storage system commonly known as deduplication is increasingly accepted as an effective technique to reduce storage costs. Thus, it has been applied to different storage types, including archives and backups, primary storage, within solid state disks, and even to random access memory. Although the general approach to deduplication is shared by all storage types, each poses specific challenges and leads to different trade-offs and solutions. This diversity is often misunderstood, thus underestimating the relevance of new research and development.

The first contribution of this paper is a classification of deduplication systems according to six criteria that correspond to key design decisions: granularity, locality, timing, indexing, technique, and scope.
This classification identifies and describes the different approaches used for each of them. As a second contribution, we describe which combinations of these design decisions have been proposed and found more useful for challenges in each storage type. Finally, outstanding research challenges and unexplored design points are identified and discussed.

Citation Key:

PP14b

DOI:

10.1145/2611778

Preview	Attachment	Size
	pp14b.pdf	555.8 KB

José Orlando Pereira

Distributed Systems