Data from: Physics-informed deep learning for plasmonic sensing of nanoscale protein dynamics in solution
Data files
Aug 29, 2025 version files 7.70 MB
-
README.md
3.12 KB
-
sourcedata_CNN.zip
1.18 MB
-
sourcedata_Fig.1B_C_D.xlsx
55.03 KB
-
sourcedata_Fig.2A_C_D_E_f.xlsx
184.19 KB
-
sourcedata_Fig.3B_C_D.xlsx
15.92 KB
-
sourcedata_Fig.4A_B_C_E_F.xlsx
67.71 KB
-
sourcedata_Fig.S10B_D_F.xlsx
35.78 KB
-
sourcedata_Fig.S11.xlsx
1.30 MB
-
sourcedata_Fig.S12.xlsx
28.86 KB
-
sourcedata_Fig.S13.xlsx
32.07 KB
-
sourcedata_Fig.S14A_B.xlsx
275.39 KB
-
sourcedata_Fig.S15A_B_C_D.xlsx
37.05 KB
-
sourcedata_Fig.S16.xlsx
1.30 MB
-
sourcedata_Fig.S17A_B.xlsx
279.67 KB
-
sourcedata_Fig.S18A_B.xlsx
32.44 KB
-
sourcedata_Fig.S19A_B.xlsx
77.05 KB
-
sourcedata_Fig.S1A_B_C_D_E.xlsx
119.05 KB
-
sourcedata_Fig.S20A_B_C_D.xlsx
216.25 KB
-
sourcedata_Fig.S21A_B_C_D.xlsx
47.37 KB
-
sourcedata_Fig.S22A_B.xlsx
68.54 KB
-
sourcedata_Fig.S23A_B.xlsx
116.54 KB
-
sourcedata_Fig.S3A_B_C.xlsx
46.49 KB
-
sourcedata_Fig.S4A_B_C_D_E_F_G_H_I.xlsx
220.04 KB
-
sourcedata_Fig.S5A.xlsx
10.23 KB
-
sourcedata_Fig.S6A_B_C.xlsx
165.50 KB
-
sourcedata_Fig.S7B_D.xlsx
31.98 KB
-
sourcedata_Fig.S8A_B.xlsx
41.66 KB
-
sourcedata_Fig.S9B.xlsx
31.81 KB
-
sourcedata_FigS24A_B_C_D_E_F_G_H_I.xlsx
1.68 MB
Abstract
Quantifying nanoscale protein secondary structure in aqueous solutions is crucial for understanding protein interactions and dynamics. Deep learning models are adept at predicting protein secondary structures but their ability to model them in aqueous solutions is hindered by data constraints. Here, we present a mid-infrared plasmonic sensor integrated with a synthesized complex-frequency waves (s-CFW) informed convolutional neural network (CNN) to address these limitations. Our sensor enables direct probing of the amide I band in sub-10 nm proteins. By employing s-CFW to amplify target spectral features, the developed physics-informed CNN achieves a mean relative error of less than 0.1 in predicting secondary structure percentages—over twice as accurate as a pristine CNN. Our method enables in-situ and real-time quantification of subtle conformational changes during protein assembly, thereby addressing the issue of data scarcity that currently hinders the development of advanced deep learning models for predicting protein dynamics and interactions in physiological environments.
Dataset DOI: 10.5061/dryad.qjq2bvqtd
Open Source datasets for the manuscript titled "Physics-informed Deep Learning for Plasmonic Sensing of Nanoscale Protein Dynamics in Solution".
This data consists of the raw data of the images and the code involved in the article. It contains raw data from nano-fabriction experiments to simulations, and the CNN models.
Description of the data and file structure
We describe our raw data in the order in which each image appears in the main text. The final dataset includes:
- Experimental spectra (results of infrared spectrum measurement of SNFs)
- Simulated spectra (simulated data of proteins under different parameters)
- Spectra after data augmentation (with noise addition and s-CFW processing)
- Data generated by the CNN model (prediction results and loss values during the training process)
- Device and sample characterization data (AFM characterization results)
- A compressed package of CNN datasets (all processing code)
Raw data naming format: sourcedata_FigureNumber_Subnumber.xlsx
Comprehensive examples:
- sourcedata_Fig.1B_C_D.xlsx
sourcedata - indicates this is a source data to plot figures
Fig.1 - indicated this is the source data from Figure 1
B_C_D - the origin data of figure subnumber B, C, and D - sourcedata_Fig.S1A_B_C_D_E.xlsx
sourcedata - indicates this is a source data to plot figures
Fig.S1 - indicated this is the source data from Supplymentry Figure S1
A_B_C_D_E - the origin data of figure subnumber A, B, C, D, and E
Note: The datasets are listed vertically (or column-wise) within our sourcedata.xlsx files. For data processing scenarios (e.g., spreadsheet analysis, code reading), "column-wise arrangement" specifically refers to data distribution in columns (as opposed to "row-wise" for horizontal distribution), which should be noted when importing or parsing the data.
The folder "sourcedata_CNN.zip" contains all source code files used in this study.
Within this directory, the file "README_CNN.docx" provides detailed explanations of each Python script (.py), including:
- the specific function of each script
- the purpose of the test files
- descriptions of the model-generated weight parameters
Sharing/Access Information
All data is available here. If anything appears to be missing, please contact the corresponding author (Chenchen Wu).
Code/software
All our CNN model codes are run and computed using Python.
Please refer to GitHub for the complete code with train/test/prediction dataset: https://github.com/syY6777/Predicting-protein-secondary-structure.git
Human subjects data
We state that our received explicit consent from our participants to publish the de-identified data in the public domain.
Human subjects data
We state that our received explicit consent from our participants to publish the de-identified data in the public domain.
