Designing an Efficient Deduplication Algorithm for Audio Files in Cloud Storage

Authors

DOI:

https://doi.org/10.5281/zenodo.19347954

Keywords:

Deduplication, Hash Table, MD6, Audio Files, Cloud Storage

Abstract

Duplicate data is a major problem in big data storage systems because it consumes additional space and affects data organization, management, and processing, while ideal systems are expected to use storage capacity efficiently. To address this, hash algorithms are used to generate hash keys for files, so that identical files share the same key. However, different files may sometimes receive the same key, which is known as a collision, and the likelihood of this depends on the key length: the longer the key, the lower the probability of collision. When a file is uploaded to cloud storage, its hash key is compared with the keys already stored, but as the volume of data grows, the search time increases. This paper proposes a file-level deduplication technique for audio data in cloud storage. The technique relies on building a hash table with multiple indexes according to the audio file format (uncompressed, lossy compressed, lossless compressed), where a separate index is assigned to each format. The MD6 algorithm, which generates a 512-bit key, is used to reduce the probability of collision.

Downloads

Download data is not yet available.

References

[1] I. A. T. Hashem, I. Yaqoob, N. B. Anuar, S. Mokhtar, A. Gani, and S. U. Khan, “The rise of ‘big data’ on cloud computing: Review and open research issues,” Information Systems, vol. 47, pp. 98–115, 2015, doi: 10.1016/j.is.2014.07.006.

[2] N. Sharma, A. V. Krishna Prasad, and V. Kakulapati, “Data deduplication techniques for big data storage systems,” Int. J. Innov. Technol. Explor. Eng. (IJITEE), vol. 8, no. 10, pp. 1145–1150, 2019, doi: 10.35940/ijitee.J9129.0881019.

[3] M. Muniswamaiah, T. Agerwala, and C. Tappert, “Big data in cloud computing review and opportunities,” Int. J. Comput. Sci. Inf. Technol. (IJCSIT), vol. 11, no. 4, pp. 43–57, Aug. 2019, doi: 10.5121/ijcsit.2019.11404.

[4] V. Schmitt and J. Jordaan, “Establishing the validity of MD5 and SHA-1 hashing in digital forensic practice in light of recent research demonstrating cryptographic weaknesses in these algorithms,” Int. J. Comput. Appl., vol. 68, no. 23, pp. 40–43, Apr. 2013, doi: 10.5120/11723-7433.

[5] M. Eichlseder, F. Mendel, and M. Schläffer, “Branching heuristics in differential collision search with applications to SHA-512,” in Fast Software Encryption (FSE 2014), 2014, pp. 473–488, doi: 10.1007/978-3-662-46706-0_24.

[6] R. L. Rivest, “The MD6 hash function: A proposal to NIST for SHA-3,” Massachusetts Institute of Technology, Cambridge, MA, USA, Tech. Rep., 2008.

[7] “Audio file format,” Wikipedia, accessed Feb. 23, 2026.

[8] N. A. Naveen and V. Ravi, “Client side deduplication scheme for secured data storage in cloud environments,” Int. J. Eng. Res. Technol. (IJERT), vol. 4, no. 5, pp. 1465–1467, 2015.

[9] V. S. R. and D. K. Singh, “Secure deduplication techniques: A study,” Int. J. Comput. Appl., vol. 137, no. 8, pp. 41–43, 2016, doi: 10.5120/ijca2016908874.

[10] P. Prajapati, P. Shah, A. Ganatra, and S. Patel, “Efficient cross user client side data deduplication in Hadoop,” J. Comput., vol. 12, no. 4, pp. 362–370, 2017, doi: 10.17706/JCP.12.4.362-370.

[11] I. Vaidya and R. Nath, “An improved de-duplication technique for small files in Hadoop,” Int. Res. J. Eng. Technol. (IRJET), vol. 4, no. 7, pp. 2040–2045, 2017.

[12] M. R. Hudagi and S. A. Urabinahatti, “Efficient deduplication using Hadoop,” Int. J. Latest Trends Eng. Technol., vol. 10, no. 3, pp. 236–238, 2018.

[13] Q. He, G. Bian, B. Shao, and W. Zhang, “Research on multifeature data routing strategy in deduplication,” Scientific Programming, vol. 2020, Art. no. 8869237, pp. 1–11, 2020, doi: 10.1155/2020/8869237.

[14] N. Sharma and A. V. Krishna Prasad, “File-level deduplication by using text files—Hive integration,” in Proc. Int. Conf. Comput. Commun. Informatics (ICCCI), 2021, pp. 1–6.

[15] G. Sujatha and J. Retna Raj, “Improving the efficiency of deduplication process by dedicated hash table for each digital data type in cloud storage system,” Webology, vol. 18, Special Issue on Artificial Intelligence in Cloud Computing, pp. 288–301, 2021.

Downloads

Published

2023-12-03

Issue

Section

Articles – Volume 1 Number 1

Categories

How to Cite

[1]
A. . Zakzouk and H. Hasan, “Designing an Efficient Deduplication Algorithm for Audio Files in Cloud Storage”, J.W.P.U, vol. 1, no. 1, pp. 162–172, Dec. 2023, doi: 10.5281/zenodo.19347954.

Similar Articles

11-20 of 26

You may also start an advanced similarity search for this article.