Designing an Efficient Deduplication Algorithm for Audio Files in Cloud Storage

Authors

DOI:

https://doi.org/10.5281/zenodo.19347954

Keywords:

Deduplication, Hash Table, MD6, Audio Files, Cloud Storage

Abstract

Duplicate data poses a significant challenge in big data storage systems as it consumes storage space, affecting data organization, management, and processing. To solvethis problem, hashalgorithms are used to generate hashkeys for files. However, as theamount of data stored in the cloud increases, the search and matching process takes longer. Additionally, hashkeys can match different files, known as collisions, which are related to the length of the hashkey. The longer the key, the less likely collisions will occur.In this paper, we present a technique for eliminating duplicate data at the file level to reduce storage of duplicate audio data in the cloud storage system. The proposed technique aims to reduce the search time for hashvalues by creatinga reduction table with multiple indexes. These indexes are designed based on the audio file format. Therefore, the hashtable includes multiple indexes, each for a specific format. To minimize the probabilityof collisions, MD6 algorithm is used, which produces a key with a length of 512 bits.

Downloads

Download data is not yet available.

References

[1] I. A. T. Hashem, I. Yaqoob, N. B. Anuar, S. Mokhtar, A. Gani, and S. U. Khan, “The rise of ‘big data’ on cloud computing: Review and open research issues,” Information Systems, vol. 47, pp. 98–115, 2015, doi: 10.1016/j.is.2014.07.006.

[2] N. Sharma, A. V. Krishna Prasad, and V. Kakulapati, “Data deduplication techniques for big data storage systems,” Int. J. Innov. Technol. Explor. Eng. (IJITEE), vol. 8, no. 10, pp. 1145–1150, 2019, doi: 10.35940/ijitee.J9129.0881019.

[3] M. Muniswamaiah, T. Agerwala, and C. Tappert, “Big data in cloud computing review and opportunities,” Int. J. Comput. Sci. Inf. Technol. (IJCSIT), vol. 11, no. 4, pp. 43–57, Aug. 2019, doi: 10.5121/ijcsit.2019.11404.

[4] V. Schmitt and J. Jordaan, “Establishing the validity of MD5 and SHA-1 hashing in digital forensic practice in light of recent research demonstrating cryptographic weaknesses in these algorithms,” Int. J. Comput. Appl., vol. 68, no. 23, pp. 40–43, Apr. 2013, doi: 10.5120/11723-7433.

[5] M. Eichlseder, F. Mendel, and M. Schläffer, “Branching heuristics in differential collision search with applications to SHA-512,” in Fast Software Encryption (FSE 2014), 2014, pp. 473–488, doi: 10.1007/978-3-662-46706-0_24.

[6] R. L. Rivest, “The MD6 hash function: A proposal to NIST for SHA-3,” Massachusetts Institute of Technology, Cambridge, MA, USA, Tech. Rep., 2008.

[7] “Audio file format,” Wikipedia, accessed Feb. 23, 2026.

[8] N. A. Naveen and V. Ravi, “Client side deduplication scheme for secured data storage in cloud environments,” Int. J. Eng. Res. Technol. (IJERT), vol. 4, no. 5, pp. 1465–1467, 2015.

[9] V. S. R. and D. K. Singh, “Secure deduplication techniques: A study,” Int. J. Comput. Appl., vol. 137, no. 8, pp. 41–43, 2016, doi: 10.5120/ijca2016908874.

[10] P. Prajapati, P. Shah, A. Ganatra, and S. Patel, “Efficient cross user client side data deduplication in Hadoop,” J. Comput., vol. 12, no. 4, pp. 362–370, 2017, doi: 10.17706/JCP.12.4.362-370.

[11] I. Vaidya and R. Nath, “An improved de-duplication technique for small files in Hadoop,” Int. Res. J. Eng. Technol. (IRJET), vol. 4, no. 7, pp. 2040–2045, 2017.

[12] M. R. Hudagi and S. A. Urabinahatti, “Efficient deduplication using Hadoop,” Int. J. Latest Trends Eng. Technol., vol. 10, no. 3, pp. 236–238, 2018.

[13] Q. He, G. Bian, B. Shao, and W. Zhang, “Research on multifeature data routing strategy in deduplication,” Scientific Programming, vol. 2020, Art. no. 8869237, pp. 1–11, 2020, doi: 10.1155/2020/8869237.

[14] N. Sharma and A. V. Krishna Prasad, “File-level deduplication by using text files—Hive integration,” in Proc. Int. Conf. Comput. Commun. Informatics (ICCCI), 2021, pp. 1–6.

[15] G. Sujatha and J. Retna Raj, “Improving the efficiency of deduplication process by dedicated hash table for each digital data type in cloud storage system,” Webology, vol. 18, Special Issue on Artificial Intelligence in Cloud Computing, pp. 288–301, 2021.

Downloads

Published

2023-12-03

Issue

Section

Articles – Volume 1 Number 1

Categories

How to Cite

[1]
A. . Zakzouk and H. Hasan, “Designing an Efficient Deduplication Algorithm for Audio Files in Cloud Storage”, J.W.P.U, vol. 1, no. 1, pp. 162–172, Dec. 2023, doi: 10.5281/zenodo.19347954.

Similar Articles

1-10 of 26

You may also start an advanced similarity search for this article.