TY - GEN
T1 - Decoding the Granular Puzzle of Macromolecules
T2 - 2024 IEEE International Conference on Big Data, BigData 2024
AU - Małysiak-Mrozek, Bozena
AU - Pawłowicz, Paulina
AU - Sunderam, Vaidy
AU - Hung, Che Lun
AU - Kwiecien, Andrzej
AU - Mrozek, Dariusz
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - Proteins are complex biological information granules that play a crucial role in various cellular processes within living organisms. Processing 3D protein structures, which are the most informative from the biological point of view, is both intricate and time-consuming. In particular, performing 3D protein structure searches against large protein datasets involves identifying similarities and conducting structural alignments across numerous molecules (granules). This task demands advanced methods for matching identical and similar regions within protein structures and substantial computational resources to handle large collections of macromolecular data efficiently. In this paper, we present our parallel implementation of scalable 3D structural alignment on the Apache Spark big data platform. We describe a customized approach that leverages Spark data transformations within the data processing pipeline for the alignment process. Our experimental results demonstrate that this solution, tightly integrated with the Spark processing model, is both efficient and scalable, even with the increasing volume of protein structure data.
AB - Proteins are complex biological information granules that play a crucial role in various cellular processes within living organisms. Processing 3D protein structures, which are the most informative from the biological point of view, is both intricate and time-consuming. In particular, performing 3D protein structure searches against large protein datasets involves identifying similarities and conducting structural alignments across numerous molecules (granules). This task demands advanced methods for matching identical and similar regions within protein structures and substantial computational resources to handle large collections of macromolecular data efficiently. In this paper, we present our parallel implementation of scalable 3D structural alignment on the Apache Spark big data platform. We describe a customized approach that leverages Spark data transformations within the data processing pipeline for the alignment process. Our experimental results demonstrate that this solution, tightly integrated with the Spark processing model, is both efficient and scalable, even with the increasing volume of protein structure data.
KW - 3D structures
KW - Apache Spark
KW - Big Data
KW - alignment
KW - information granules
KW - proteins
UR - https://www.scopus.com/pages/publications/85218066416
U2 - 10.1109/BigData62323.2024.10825877
DO - 10.1109/BigData62323.2024.10825877
M3 - Conference contribution
AN - SCOPUS:85218066416
T3 - Proceedings - 2024 IEEE International Conference on Big Data, BigData 2024
SP - 8317
EP - 8324
BT - Proceedings - 2024 IEEE International Conference on Big Data, BigData 2024
A2 - Ding, Wei
A2 - Lu, Chang-Tien
A2 - Wang, Fusheng
A2 - Di, Liping
A2 - Wu, Kesheng
A2 - Huan, Jun
A2 - Nambiar, Raghu
A2 - Li, Jundong
A2 - Ilievski, Filip
A2 - Baeza-Yates, Ricardo
A2 - Hu, Xiaohua
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 15 December 2024 through 18 December 2024
ER -