Logo image
Proteintox: A multifaceted machine learning strategy for identifying cardiotoxic, neurotoxic, and enterotoxic proteins
Journal article

Proteintox: A multifaceted machine learning strategy for identifying cardiotoxic, neurotoxic, and enterotoxic proteins

Pradnya Kamble, Anju Sharma, Aritra Banerjee, Shubham Pandey, Shubham K Pandey and Prabha Garg
Computational toxicology, Vol.36, p.100390
12/01/2025

Abstract

Life Sciences & Biomedicine Science & Technology Toxicology
Accurate prediction of protein toxicity is paramount in various fields, ranging from pharmaceutical drug development to environmental risk assessment, as it allows for early identification and mitigation of potentially harmful effects associated with protein exposure. Cardiotoxicity, enterotoxicity, and neurotoxicity are critical concerns that demand rigorous assessment during the early stages of drug development. This study addresses the need for accurate prediction models to identify proteins and peptides with potential cardiotoxic, enterotoxic, or neurotoxic effects. By leveraging machine learning (ML) techniques (support vector machine (SVM), random forest (RF), k-nearest neighbour (kNN)), and comprehensive datasets encompassing a wide range of molecular features, robust prediction models were developed to reliably classify proteins and peptides based on their potential toxicity profiles. The models integrate diverse features, including amino acid composition (Compo), conjoint-triads (CTriad), composition-transition-distribution (CTD), and physicochemical n-gram properties (PnGT) derived from protein primary sequences, enabling holistic analysis of the toxicity potential of the molecules. Various models were developed using isolated feature sets and combinations of four feature sets. The RF model consistently outperforms the other models in toxicity prediction, with the Compo + CTriad + CTD feature set being recommended because of its ability to capture intricate molecular interactions and structural details. The proposed model, Proteintox, balances detailed structural insights with practicalities, enhancing its ability to assess impacts involving molecular interactions. It delivers high accuracy, sensitivity, and specificity across all testing scenarios while remaining computationally efficient and interpretable. The study also highlights the significance of selecting appropriate feature sets to enhance model performance without increasing complexity, demonstrating that adding more features does not always translate to improved predictive ability. The significance of this work lies in its potential to streamline the drug discovery process by providing early toxicity predictions, thus reducing the reliance on costly and time-consuming experimental assays. The data and source code are available at https://github.com/PGlab-NIPER/Proteintox.
url
Article Landing PageView

Metrics

Details

Logo image