OAK

Enhanced toxicity prediction of PFAS in aquatic species via AI-Based docking

Metadata Downloads
Author(s)
Yongwan Lee
Type
Thesis
Degree
Master
Department
공과대학 환경·에너지공학과
Advisor
Kim, Sang Don
Abstract
With the rapid expansion of global chemical industries, regulatory frameworks such as EU REACH and US EPA TSCA demand extensive toxicological profiles for chemical substances. However, conventional in vivo animal testing suffers from structural bottlenecks due to ethical concerns, alongside temporal and economic constraints. While traditional in silico alternatives like QSAR and read-across offer high-throughput capabilities, they lack molecular-level mechanistic insights and are highly constrained by the availability of structurally similar training data. To address these challenges, reverse molecular docking leveraging artificial intelligence and comprehensive structural databases has re-emerged as a pivotal tool for mechanism-based toxicity prediction.
The primary objective of this study is to establish and validate a highly efficient, Proteome-wide reverse molecular docking pre-screening framework in zebrafish (Danio rerio). Redefining the paradigm of molecular docking, the core novelty of this pipeline lies not in pinpointing a single, absolute target protein, but in acting as a robust prioritization filter that compresses thousands of proteomic candidates down to a high-value top 5% list prior to laboratory testing. This strategic compression fundamentally circumvents the statistical dilution of toxicological signals often encountered during broad pathway enrichment analyses.
The workflow initiated with the construction of an Exploratory Target Library, curating 86,691 zebrafish whole-proteome entries from the UniProtKB database through a systematic 5-step filtering protocol to mitigate computational overhead and structural noise. The retrieved AlphaFold 3D coordinates (.pdb) were preprocessed into .pdbqt format by integrating essential hydrogen atoms and partial charges. Reverse docking simulations were thoroughly compared between the empirical scoring-based AutoDock Vina and the deep-learning-based GNINA software. The early target recognition and enrichment capabilities of the framework were rigorously quantified using the EF5% and the Boltzmann-Enhanced Discrimination of Receiver Operating Characteristic (BEDROC) metrics, with the exponential parameter(α) set to 32.2 to focus 80% of the statistical weight on the top 5% tier.
The methodology was pre-validated using four reference chemicals with well-characterized toxicological profiles: Bisphenol A (BPA), Triclosan, Rotenone, and Valproic acid (VPA). GNINA significantly outperformed AutoDock Vina, effectively concentrating known toxic targets within the top 5% tier and yielding high EF5% values ranging from 4 to 8 for multi-target compounds (BPA, Triclosan, and Rotenone). In contrast, VPA exhibited a lower EF5% of 1.7. This variance was structurally and pharmacologically defended; VPA's small, flexible fatty acid backbone lacking aromatic rings encourages non-specific insertion into shallow pockets, thereby elevating false-positive background noise. Furthermore, the lack of dynamic flexibility and co-factor data — specifically the crucial zinc ion active coordinates inside Histone Deacetylase (HDAC) — in static AlphaFold structures inherently limited binding energy calculations for specific metalloproteins. Nevertheless, VPA's enrichment remained statistically superior to random selection.
The validated prioritization framework was subsequently applied to per- and polyfluoroalkyl substances (PFAS: PFBA, PFOS, and PFNA), which are critical unregulated emerging contaminants. Pathway enrichment analyses derived exclusively from the prioritized top 5% target list successfully predicted complex organ- and system-level toxicity pathways, including lipid metabolism disruption and oxidative stress. While computational limitations regarding rigid protein libraries persist, they can be effectively complemented by future coupling with Molecular Dynamics (MD) simulations and targeted in vitro cell-line-based high-throughput assays.
In conclusion, this study offers a resource-efficient guideline that bridges the gap between massive chemical inventories and limited empirical screening capacities. By precisely isolating high-value core targets and preventing blind laboratory testing, this framework establishes a practical foundation for New Approach Methodologies (NAMs), shifting the toxicity assessment paradigm toward target-centric virtual screening.
URI
https://scholar.gist.ac.kr/handle/local/34506
Fulltext
http://gist.dcollection.net/common/orgView/200001019215
Alternative Author(s)
이용완
Appears in Collections:
Department of Environment and Energy Engineering > 3. Theses(Master)
공개 및 라이선스
  • 공개 구분공개
파일 목록
  • 관련 파일이 존재하지 않습니다.

Items in Repository are protected by copyright, with all rights reserved, unless otherwise indicated.