Identification of a natural product inhibitor of SARS-CoV-2 Mpro through integrated virtual screening, DFT, and molecular dynamics simulations
Data files
Sep 09, 2026 version files 136.26 MB
-
data1.zip
136.11 MB
-
README.md
24.34 KB
-
Virtual_screening_ligand_library_data.zip
131.82 KB
Abstract
This dataset contains computational data from a virtual screening procedure to identify SARS-CoV-2 Mpro inhibitors. Approximately 220,000 natural product compounds from ZINC15 were screened using ADMET filtering, molecular docking against Mpro (PDB ID: 6LU7), and docking against hMAT1A and CYP450 isoforms to evaluate hepatotoxicity and drug clearance. Top compounds were subjected to DFT calculations, followed by MD simulations with post-MD analyzes of the top lead molecule (Lig14) (COM trajectory extraction, sub-diffusive behavior, MM-GBSA, and alchemical solvation free-energy calculations). Mutational analysis was performed on relevant mutants (G143S, G143D, G143N, C145Y, E166K) to assess resistance profiles. The dataset contains (1) filtered ligand library details; (2) DFT-optimized structured files for the top four compounds; (3) MD trajectory files for the lead molecule (Lig14); and (4) mutational analysis (PDB files and sequences) for wild-type and mutant Mpro complexes. This resource facilitates method validation, comparative analysis, and future computational drug discovery research.
Dataset DOI: 10.5061/dryad.76hdr7tbm
Corresponding Author: Narayan Prasad Adhikari, narayan.adhikari@cdp.tu.edu.np
Associated Publication: Scientific Reports
1. Summary
This dataset provides the computational data generated during a study to identify potential inhibitors of the SARS-CoV-2 main protease (Mpro) from a natural product library. The study integrates multiple computational techniques: virtual screening with ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) filtering, molecular docking, Density Functional Theory (DFT) calculations, Molecular Dynamics (MD) simulations, and mutational analysis. The data presented here allow for the reproduction of the figures and tables in the associated publication and can be reused for further computational drug discovery studies.
2. Description of the Data and File Structure
The dataset is provided in two compressed files named 'data1.zip' and 'Virtual_screening_ligand_library_data.zip' . When unzipped, 'data1.zip' creates a main directory called data1, which includes four primary sub-folders. Each folder is organized to correspond to specific figures, analyses, or stages of the computational pipeline. Below is a detailed breakdown of the contents of each folder. The 'Virtual_screening_ligand_library_data.zip' contains the information regarding the ligand library molecules.
Inside 'data1.zip’ folder :
Folder: 1. Figure_data
This folder contains the data files used to generate figures of the publications.
- Fig1_optimized_coordinate_files/:
o Description: Contains DFT-optimized structure files (in .txt format) for the top four ligands (Lig1,Lig2,Lig4 , and Lig14). These are the optimized 3D coordinates obtained after Density functional Theory (DFT) calculations.
o File Format: Plain text with four columns and a header row.
o File Structure:
- Line 1: Ligand name (e.g., Lig1, Lig2, Lig4, Lig14)
- Line 2: Charge and multiplicity (e.g., "0 1" =0: neutral charge, 1: singlet state )
- Line 3: Header row: "Atoms X Y Z"
- Remaining lines: Each line represents an atom type (e.g., C,H,O,N) followed by its X,Y,Z, coordinates in Angstroms (Å).
o Variables:
- Atoms: Chemical element symbol (C, H, O, N)
- X: Cartesian X-coordinate in Angstroms (Å)
- Y: Cartesian Y-coordinate in Angstroms (Å)
- Z: Cartesian Z-coordinate in Angstroms (Å)
o Interpretation: These coordinates describe the lowest-energy (optimized) 3D structure of each ligand in its ground state. The data can be utilized for molecular visualization, as input for studies, including quantum chemical computations.
o Software: Viewable in any text editor. Recommended for visualization: Avogadro (free/open) .
Fig2_complex_representation/: Includes a .pdb file for the PyMOL representations of protein Mpro-Lig14 complex.
File Format: .pdb Files (Protein Data Bank Format):
Variables in .pbd files:
-
CRYST1 record: Describes the crystallographic unit cell dimensions (a,b,c in Å) , and angles (alpha, beta, gamma in degrees)
-
ATOM records: Contain atomic coordinates with the following variables:
(1) Atom serial number : unique atom identifier expressed in integer (e.g., 1,2,3,..)
(2) Atom name : Chemical element symbol (e.g., N, C, O, H)
(3) Residue name : Amino acid three-letter code (e.g., SER, GLY, ALA)
(4) Chain identifier : Chain ID (e.g., P - protein chain)
(5) Residue sequence number: Position of the residue in the protein sequence (e.g., 1,2,3, ..)
(6) X, Y, Z Cartesian coordinates in Angstroms (Å) (e.g., -22.758 -5.168 -26.981)
(7) Occupancy : Electron density occupancy (0.00 or 1.00)
(8) Temperature factor: B-factor (thermal mobility) in Ų (e.g. 0.00)
Note: Extra columns (e.g., “PROA ” ) are software-specific annotations (protein segment identifier ) which can be ignored
Example:
Column: (1) (2) (3) (4) (5) (6) (6) (6) (7) (8)
ATOM 1 N SER P 1 -22.758 -5.168 -26.981 0.00 0.00
Interpretation: The coordinates define the 3D structure of the molecule, which can be visualized in PyMOl,,VMD or any molecular visualization software .
Fig3_solvation/:
Description
- Includes sub-folder ‘data_xvg_format’ that contains. xvg files with the raw energy derivative (dU/dλ) data for Lig14, generated from alchemical solvation free energy calculations at lambda (λ) states using GROMACS software.
o Purpose: Used to obtain the solvation free energy (ΔG_solv) of Lig14 by integrating dU/dλ values over the λ coordinate.
o File Format: Plain text (GROMACS .xvg format).
o File Structure:
- Lines starting with '#' or '@' contain metadata and plotting instructions
- Data lines contain 29 columns of numerical values.
o Variables (Column 1 and 2 are the most important):
-Column 1: Time (ps)
-Column 2: Total dU/dλ (kJ/mol [λ]⁻¹) ( Main value for integration)
- Column 3: mass-lambda (contribution from mass perturbation)
- Column 4: coul-lambda (contribution from Columbic/electrostatic interactions)
- Column 5: vdw-lambda (contributions from van der Waals interactions)
-Column 6: bonded-lambda (contribution from bonded interactions)
-Column 7: restraint-lambda (contribution from restraints)
- Columns 8-28 : Cumulative ΔU values at different λ states (21 columns)
- Column 29: pV term (pressure × volume in kJ/mol)
o How to use these data:
1. View in any text editor to inspect raw data
2. Extract column 2 (total dU/dλ ) for each λ-state
3. Average dU/dλ over time for each λ-state (Python)
4. Plot average dU/dλ vs. λ and integrate to get ΔG_solv (kJ/mol)
o Software:
-View: Any text editor
- Plot: Python
Fig4_subdiffusive_nature_Lig14/:
It contains data to analyze the sub-diffusive nature of Lig14 within the biological environment, which includes two files:
File 1: Center_of_mass_trajectory_data_Lig14.txt :
This text file contains the extracted center of mass (COM) coordinates of the Lig14 ligand over the course of the MD trajectory.
File Format: Plain text with three columns representing X, Y, and Z coordinates in Angstroms (Å).
Data Details:
§ Each row represents the COM position at a specific time point.
§ The data was extracted at regular intervals of 10 ns between successive frames (stride = 100 frames, assuming 100 ps per saved frame)
§ This trajectory spans the entire simulation length (approximately 300 ns).
Source: It was extracted from trajectory file (.dcd file) of the Mpro-Lig14 complex MD simulation using the provided TCL script.
File 2; msd_extract.tcl
Description: A TCL script used within the Visual Molecular Dynamics (VMD) software to extract the COM trajectory data from the MD simulation files.
· Fig5_swissadme_Lig14/: Contains the SMILES string file in .csv format of the lead compound Lig14.
· Purpose: To provide the chemical structure of the lead compound Lig14 to upload to server like SWISSADME to get chemical properties of the molecule
o Software: Any text editor can open the file.
Folder: 2. Additional_data
· MD_input_Lig14/ and MD_input_mpro_protein/:
o Description: These folders contain input structure files for NAMD (Nano scale molecular dynamics) simulations.
- MD_input_Lig14/: Contains 'input_complex.pdb' for the Mpro-Lig14 complex (protein + ligand).
- MD_input_mpro_protein/: Contains the ‘input_mpro.pdb’ file for the Mpro protein in its apo (unbound) form.
o File Format: .pdb (standard Protein Data Bank format -variable descriptions are the same as in the "Figure_data/Fig2_complex_representation" section).
o Note: Extra columns (e.g., "PROA N") are software-specific annotations (protein segment identifier and element symbol) which can be ignored.
o Contents: Atomic coordinates for the Mpro-Lig14 complex (bound) and apo Mpro (unbound) form.
· Software: NAMD (simulation), VMD/PyMOL (visualization).
· MMGBSA_Lig14/:
o Description: Contains a .dat file with MM-GBSA calculated binding free energy (ΔG_bind) data for the Lig14-Mpro complex, obtained from the last 100ns of the MD simulation.
o File Format: Plain text (.dat file) with two columns.
o Variables:
- Column 1: Time (ns) - simulation time in nanoseconds
- Column 2: ΔG_bind (kcal/mol) – binding free energy at each time point
o Interpretation: The average of the (ΔG_bind) values over the entire 100 ns trajectory gives the final estimated binding free energy of the Lig14-Mpro complex. More negative values indicate stronger binding affinity.
o To use:
1. Open in any text editor or spreadsheet software
2. Calculate the average of column 2 to obtain final (ΔG_bind) .
· Molecules_for_virtual_screenings/:
o Description: Contains .csv files of the molecules utilized in the virtual
screening campaign.
o Files:
-File 1: Initial_Full_molecules_for_screenings_Natural_products.csv: Contains all
224,205 molecules obtained from the natural-products subset from the ZINC database with columns: ZINC ID and SMILES.
-File 2: After_ADMET_molecules.csv: Contains the 28,421 molecules that passed initial ADMET filtering It contains the columns: Molecule No., ZINC ID, and SMILES.
o File Format: .csv (Comma-Separated Values) - viewable in any text editor or spreadsheet software.
o Variables:
- Molecule No.: Serial number assigned to each molecule (only in the
filtered file ‘After_ADMET_molecules.csv’)
- ZINC ID: Unique identifier for each molecule from the ZINC database
- SMILES: Simplified molecular-input line-entry system string describing
the chemical structure
o How to use :
1. Open in Excel or any text editor.
2. Search for specific molecules using ZINC IDs
3. Use SMILES strings for further computational studies (e.g., docking,
similarity searching)
o Software: Any text editor , including Excel , and Python.
Folder: 3. Supplementary_figure_data
This folder contains raw data for supplementary figures (S1-S9), supporting analyses of molecular interactions, stability, and energetics. The following are the sub-folders:
· Figure_S1/: .pdb files for protein-ligand complexes (Lig1, Lig2, Lig4, and Lig14 with protein ‘Mpro’)
Purpose: To visualize 2D/3D interactions.
File Format: .pdb. -variable descriptions are the same as in the "Figure_data/Fig2_complex_representation" section.
Software: PyMOL, Discovery studio
Figure_S2/: .csv files with RMSD data for Lig1-Mpro (fig. S2 a. Lig1_rmsd_data), Lig2-Mpro(fig. S2 b. Lig2_rmsd_data), Lig4-Mpro(fig. S2 c. Lig4_rmsd_data), and Lig14-Mpro (fig. S2 d. Lg14_rmsd_data ) complexes after MD simulations.
Purpose: To assess structural stability over time.
Format: Time (ns) | RMSD (Å)
Software: Excel, Python, R
· Figure_S3/, Figure_S6/, Figure_S7/, Figure_S8/, and Figure_S9/:
These folders contain .txt files with DFT-optimized structures (same structure format as Fig1_optimized_coordinate_files/).
Purpose: Provide optimized 3D geometries for amino acids, ligands, and frequency/NMR calculations.
Contents optimized geometries of the amino acids with representations:
- Figure_S3/: His, Glu, Thr , and ligand Lig14
- Figure_S6/: Cys, Tyr
- Figure_S7/: Ser, Gly, Asn, and Asp
- Figure_S8/: Glu, and Lys
- Figure_S9/: Contains optimized structure of Lig14 for Frequency and NMR calculation
Software: Any text editor (view), Avogadro (for visualization ).
· Figure_S4/: Contains four .csv files analyzing binding energetics and interactions of the Lig14-Mpro complexes including SASA, hydrogen bonding and NAMD energies.
Purpose: To quantify the binding interface, hydrogen bonding, and energy components of the Lig14-Mpro complex over the MD simulation trajectory.
- File1: fig.S4a.lig14bsa_data.csv:
Description: Buried surface area (BSA) data for the Lig14-Mpro complex over time.
Format: Time (ns) | Contact Area (Ų)
Variables:
-Time (ns): Simulation time in nanoseconds
-Contact area (Ų) : The surface area of the protein that becomes buried upon ligand binding. Calculated as (SASA_protein + SASA_ligand – SASA_complex)/2. Higher values indicate greater protein-ligand contact.
- File 2: FigS4a_supporting_data_bsa.csv:
Description: Supporting SASA calculations file for BSA analysis
Format: SASA_protein (Ų) | SASA_HETA(Ų) | SASA_COMPLEX (Ų) | Contact_area(BSA) (Ų),
Variables:
-SASA_protein (Ų) : Solvent Accessible Surface Area of the protein along (Mpro)
-SASA_HETA (Ų) : Solvent Accessible surface Area of the ligand alone (Lig14)
-SASA_COMPLEX (Ų): Solvent Accessible Surface Area of the protein-ligand complex
-Contact_area (BSA) (Ų) : Burried surface area
Interpretation: These values provide the raw SASA data used to calculate BSA
File 3: fig.S4b.lig14hbond_data.csv:
Description: Hydrogen bond evolution data over the simulation.
Format: Time (ns) | No. of_Hbond
Variables:
-Time (ns): Simulation time in nanoseconds
-No. of _Hbond: Number of hydrogen bonds formed between Lig14 and Mpro at each time point.
Interpretation: Hydrogen bonds are key interactions stabilizing protein-ligand complexes. This data shows how many hydrogen bonds exits over time, indicating binding stability.
File 4: fig.S4c_namd_energy_data.csv:
Description: NAMD energy components for the Lig14-Mpro complex
Format: Time (ns) | Electrostatic Energies (Kcal/mol) | VdW Energies (Kcal/mol) | Nonbond Energies (Kcal/mol)
Variables :
Time (ns): Simulation time in nanoseconds
Electrostatic (kcal/mol): Coulombic/electrostatic interaction energy between protein and ligand
Vdw(Kcal/mol): van der Waals interaction energy
-Nonbond (Kcal/mol): Total non-bonded energy = Electrostatic+ VdW
Interpretation: More negative energy values indicate stronger interactions. This data helps to understand the energetic basis of ligand bindings.
Software: Excel, Python, R.
· Figure_S5/: .csv files with RMSF and RMSD data for complex and apo forms of Lig14-Mpro.
Purpose: Compare flexibility(RMSF) and stability (RMSD) of complex vs. apo protein.
-File 1: fig.S5b.Lig14_rmsd_of_apo_and_complex_data.csv
Description: RMSD data comparing the apo and complex forms over time.
Format: Time_Mpro (ns) | Mpro_RMSD (Å) | Time_Complex (ns) | Complex_RMSD (Å)
Variables:
-Time_Mpro (ns): Simulation time for the apo protein
- Mpro_RMSD (Å) : RMSD of the apo protein relative to its initial structure
- Time_Complex (ns) : Simulation time for the complex
Interpretation: Lower and more stable RMSD values indicate a more stable structure. Comparing apo vs. complex RMSD reveals whether ligand binding stabilizes or destabilizes the protein.
-File 2: FigS5a_Lig14_apo_complex_rmsf_data.csv:
Description: RMSF data comparing residue flexibility in apo vs complex forms.
Format: Residue No. | Mpro RMSF (Å) | Lig14 RMSF (Å)
Variables:
-Residue No. : Amino acids residue number in the Mpro sequence
-Mpro RMSF (Å) : RMSF of each residue in the apo protein.
-Lig14 RMSF (Å) : RMSF of each residue in the complex (when Lig14 is bound)
Interpretation: RMSF measures atomic flexibility . Higher values indicate more flexible regions. Comparing apo vs. complex RMSF reveals which residues become more rigid or flexible upon ligand binding.
Software: Excel, Python, R .
Folder: 4. Mutational_analysis_data
This folder contains data for Mpro’s mutational analysis . It contains a folder ‘mutated_proteins’ and a .txt file.
· Folder : mutated_proteins/: Contains .pdb files for mutated Mpro proteins.
Mutants are G143S, G143D, G143N, C145Y, E166K, D187N.pdb, A191T, G302D, G302N, G302S,
M49I.pdb, and M165I.pdb.
Purpose: To study the structural and binding effects of mutations on Mpro
Software: PyMOL, VMD.
· File: all_mutant_details_sequences_fasta.txt: A FASTA file containing the sequence
information for all mutations.
Format: Standard FASTA format (plain text).
Purpose: To provide sequence data for mutational analysis and comparison.
Software: Any text editor, or sequence analysis tools.
Inside Virtual_screening_ligand_library_data.zip folder
Inside Virtual_screening_ligand_library_data.zip folder:
Virtual_screening_ligand_library_data/
This folder documents the ligand library details and molecular docking results.
Folder 1. Ligand_library/
· SMILES_ligand_library_molecules/:
File: Ligand_smile_zincid.csv: Contains SMILES strings and ZINC IDs for 4,929 molecules that passed initial filtering (ADMET, hMAT1A, and CYP3A4).
Note: hMAT1A = human Methionine Adenosyltransferase 1A (toxicity screening);
CYP3A4 = Cytochrome P450 3A4 (drug-drug interaction screening).
Format: Molecule No. |ZINC ID| SMILES
- Purpose: Provides structures of filtered molecules for further screening.
- Software: Excel, Python or any text editor.
· Docking_result_ligand_library/:
-Contains docking results (binding energy scores) against target protein CYP3A4 in .csv files
named 'Partition1', 'Partition2', etc.
- Format: Ligand | Binding Energy (kcal/mol)
- Purpose: To evaluate drug-drug interactions risk via CYP3A4 binding
- Software: Excel, Python, R.
Folder 2. Target_docking/
· Target_docking_results.csv:
Contains primary results against SARS-CoV-2 Mpro (PDB ID: 6LU7).
Format: Ligands | Binding Energy (kcal/mol)
- Purpose: To identify top hit compounds based on binding affinity to Mpro
-Software: Excel, Python, R
Scheme of data archive
Dryad Repository
│
├── data1.zip
│ │
│ ├── Figure_data/
│ │ ├── Fig1_optimized_coordinate_files/ # .txt files
│ │ ├── Fig2_complex_representation/ # .pdb file
│ │ ├── Fig3_solvation/
│ │ │ └── data_xvg_format/ # .xvg files
│ │ ├── Fig4_subdiffusive_nature_Lig14/ # .tcl, .txt files
│ │ └── Fig5_swissAdme_Lig14/ # .csv file
│ │
│ ├── Additional_data/
│ │ ├── MD_input_Lig14/ # .pdb file
│ │ ├── MD_input_mpro_protein/ # .pdb file
│ │ ├── MMGBSA_Lig14/ # .dat file
│ │ └── Molecules_for_virtual_screenings/ # .csv files
│ │
│ ├── Supplementary_figure_data/
│ │ ├── Figure_S1/ # .pdb files
│ │ ├── Figure_S2/ # .csv files
│ │ ├── Figure_S3/ # .txt files
│ │ ├── Figure_S4/ # .csv files
│ │ ├── Figure_S5/ # .csv files
│ │ ├── Figure_S6/ # .txt files
│ │ ├── Figure_S7/ # .txt files
│ │ ├── Figure_S8/ # .txt files
│ │ └── Figure_S9/ # .txt files
│ │
│ └── Mutational_analysis_data/
│ ├── mutated_proteins/ # .pdb files
│ └── all_mutant_details_sequences_fasta # .txt file
│
└── Virtual_screening_ligand_library_data.zip
│
└── Virtual_screening_ligand_library_data/
│
├── Ligand_library/
│ ├── SMILES_ligand_library_molecules/ # .csv file
│ └── Docking_result_ligand_library/ # .csv files
│
└── Target_docking/ # .csv files
3. Usage Notes
All files can be viewed using any text editor. Recommended free and open-source software.
· Visualization: VMD, PyMOl, Avogadro
· Simulation: NAMD , GROMACS
· Plotting/Analysis: Python, R, Excel , Grace
· Chemical Processing: OpenBabel, PyRx (Python Prescription-virtual screening software)
4. Abbreviations
ADMET: Absorption, Distribution, Metabolism, Excretion, Toxicity
BSA: Buried Surface Area
COM : Center of Mass
CYP3A4 : Cytochrome P450 3A4
DFT : Density Functional Theory
dU/dλ : Potential energy derivative with respect to λ
GROMACS: Molecular dynamics software
hMAT1A: Human Methionine Adenosyltransferase 1A
MD: Molecular Dynamics
MM-GBSA: Molecular Mechanics with Generalized Born and Surface Area Solvation
Mpro: Main protease
NAMD: Nanoscale Molecular Dynamics
NMR: Nuclear Magnetic Resonance
PDB: Protein Data Bank
pV: Pressure × Volume term
SASA: Solvent Accessible Surface Area
SMILES: Simplified Molecular – Input Line-Entry System
TCL: Tool Command Language
VdW: van der Waals
VMD: Visual Molecular Dynamics
ΔG_bind: Binding free energy (kcal/mol)
ΔG_solv: Solvation free energy (kJ/mol)
5. Sharing/Access Information
This dataset is archived on Dryad. The associated publication is available from [https://www.nature.com/srep/].
