Data from: PMSeeker: A Scheme for paternity marker set mining
Data files
Dec 25, 2021 version files 14.07 MB
-
peer_review_data.zip
14.07 MB
Aug 06, 2026 version files 5.05 MB
-
pmseeker_data.tar.gz
5.04 MB
-
README.md
7.11 KB
Abstract
The parentage assignment is a genetic test that analyzes genetic characteristics (mostly molecular markers) to identify whether two individuals have a parent-child relationship. It is commonly employed in judicial identification and in the breeding of economic species including a number of fishes. Because a single marker's discriminability is restricted, numerous markers are typically utilized to produce an accurate result. Obviously, having an excessive number of redundant markers wastes time and resources, and adequate approaches are required to screen reduced and efficient parentage marker sets (PMS). This study established a scheme to screen low-redundancy PMS using the exhaustive algorithm and greedy algorithm. When screening PMS, the greedy algorithm selects markers based on the parental dispersity index (PDI), a uniquely defined metric that outperforms polymorphic information content (PIC) and probability of exclusion (PE). With the conjunctive use of the two algorithms, nonredundant PMSs were found for more than 99.7% of solvable cases in three groups of random sample experiments in this study. Then a low-redundancy PMS can be composed using two or more of these nonredundant PMSs. This scheme effectively reduces the number of markers in PMS, so conserving people and experimental resources and laying the groundwork for the widespread implementation of parentage assignment technology in economic species breeding.
Dataset DOI: 10.5061/dryad.qfttdz0jh
Description of the data and file structure
This package (pmseeker_data.tar.gz) provides the data needed to reproduce the results presented in the following paper: Biology, 2025, 13(2): 100 (doi.org/10.3390/biology13020100).
It is an update to a previous version, which was uploaded before the manuscript was submitted. However, while the manuscript was being revised, the data changed but were not updated on Dryad. This update represents the final version.
Files and variables (column names)
The whole package: pmseeker_data.tar.gz
Description: The package contains three directories that hold the datasets and results corresponding to the first three sections in the 'Results' part of the paper. The following tree lists the content of the first two layers:
pmseeker_data_20260803/
├── Result_3_1
│ ├── greedy_time.txt
│ └── time_consumming.txt
├── Result_3_2
│ ├── 20210909_CPE.txt
│ ├── 20210909_greedy_exhaus2_info.txt
│ ├── 20210909_greedy_exhaus2_number.txt
│ ├── 20210909_greedy_exhaus2.txt
│ ├── 20210909_nosex_greedy_exhaus2_info.txt
│ ├── 20210909_nosex_greedy_exhaus2_number.txt
│ ├── 20210909_nosex_greedy_exhaus2.txt
│ ├── 20210909_nosex.pdi_info2.txt
│ ├── 20210909.pdi_info2.txt
│ ├── rank.nosex.Pe_num.txt
│ ├── rank.nosex.Pic_num.txt
│ ├── rank.Pe_num.txt
│ ├── rank.Pic_num.txt
│ ├── README.txt
│ └── simu_data
└── Result_3_3
├── CERVUS_peer_review
└── COLONY_peer_review
Result_3_1/greedy_time.txt:
- par_num: Number of parents.
- mar_num: Number of candidate marks.
- time: Time consuming for greedy algorithm.
Result_3_1/time_consumming.txt:
- par_num: Number of parents.
- can_num: Number of candidate marks.
- time: Time consuming (s).
- type: algorithm used.
Result_3_2/simu_data:
Simulation of different combination of microhaplotypes markers, with each file corresponding to one simulation, and "ID" is the parent ID, ".A" is the genotype of the first allele of specific marker and ".B" is the other.
Result_3_2/20210909.pdi_info2.txt:
- Marker_ID: ID of the marker.
- PDI, PIC, PE: PDI, PIC and PE value of the selected marker within the simulation dataset. The indicators were calculated according to Supplementary Notes of the paper.
Result_3_2/20210909_CPE.txt:
- 1st column: Simulations that have PMS tickout.
- 2st column: CPE value of the PMS generated by the algorithm listed in column 3.
Result_3_2/20210909_greedy_exhaus2.txt:
- Simu_round: Simulation round.
- Time_Greedy: Time for greedy algorithm.
- Time_exhaus: Time for exhaustive algorithm.
- Number_Greedy: Number of PMS detected by greedy algorithm.
- Number_exhaus: Number of PMS detected by exhaustive algorithm.
- Info: "Both_failed" refers to failure to detect PMS by both algorithm; "Greedy_contained" means PMS from greedy algorithm is contained in the exhaustive algorithm's result; "Greedy_partly_contained" means PMS from both algorithms are not the same, but there are some overlaps; "Exhaus" means exhaustive algorithm could figure out more PMSs than greedy algorithm but with the same number.
- greedy_PE: PE value for markers in each PMS detected by greedy algorithm, with the last number referring to the mean value of the indicator.
- exhaus_PE: PE value for markers in each PMS detected by exhaustive algorithm, with the last number referring to the mean value of the indicator.
- greedy_PIC: PIC value for markers in each PMS detected by greedy algorithm, with the last number referring to the mean value of the indicator.
- exhaus_PIC: PIC value for markers in each PMS detected by exhaustive algorithm, with the last number referring to the mean value of the indicator.
Result_3_2/20210909_greedy_exhaus2_info.txt:
- ROUND_NUM: Simulations that contain PMSs, indexed from 1.
- CPE: Mean PE value of the indicator (not CPE value).
- II: Mean PIC value of the indicator (not CPE value).
- ALGO: Algorithm used.
Result_3_2/20210909_greedy_exhaus2_number.txt:
- ROUND_NUM: Simulations that contain PMSs, indexed from 1.
- GREEDY_num: Number of markers within one PMS detected by greedy algorithm.
- EXHAUS_num: Number of markers within one PMS detected by exhaustive algorithm.
Result_3_2/rank.Pe_num.txt, Result_3_2/rank.Pic_num.txt:
Such files have no matched results in the paper.
Result_3_2/nosex:
Same format and information as those without "nosex".
Result_3_3/
In this directory, file names with "MAF/maf1" represent the SNP subsets whose MAF > 0.1, "MAF/maf4" for those MAF > 0.4, "MAF/maf45" for those > 0.45 and "MAF/maf475" for those > 0.475. Last number "1" for PMS-1, "2" for PMS-2 and "12" for PMS1 + PMS2. Files with "msat" stand for the SSR markers. "rlt1_1" stands for "S1-1", the first PMS from SSR marker set, "no14" stands for "S1-2" where the first marker (LOCUS 14) was removed, and "2sets" stands for "S1-1+2".
Result_3_3/CERVUS_peer_review:
*.txt:
Input of CERVUS, which contains the genotypes of markers in PMS.
*.out.csv:
Original results from CERVUS.
*.mm:
Output of PMSeeker, where each line represents a marker in the PMS (in the order from the greedy algorithm), and if multiple markers are in the same line, the markers can substitute with each other.
*.log:
Log file when running PMSeeker.
check.*.log:
Log file for checking the parentage assignment results according to CERUS's results named with "*.out.csv". The file lists the reasons why some offspring have been assigned the wrong parent pairs.
Result_3_3/COLONY_peer_review:
msat:
SSR-based input (*.csv) and analysis results of COLONY (COLONY_data).
*.colony.log:
Log file for COLONY.
*.txt.dir (subfolders):
- MAF**.txt.***: Input of COLONY.
- COLONY_data: Analysis results of COLONY.
- *.dat: The input command of COLONY.
- *.ParentPair: The file was used to calculate the accuracy.
*.txt.check.log:
Summary of COLONY results, with failures further classified as "EMPTY ASSIGNMENTS" or "WRONG ASSIGNMENT". Note: although MGW_1399 is excluded from the analysis, it was still included in all_wrong when computing ACCURACY = 17 - all_wrong. Since MGW_1399 should not count as a wrong assignment, the correct formula is 18 - all_wrong, so the true number of correctly assigned offspring is the reported value + 1 if MGW_1399 has been marked as "WRONG".
Sharing/Access information
Data for Result_3_3 was derived from Andrews et al., 2018 (https://doi.org/10.1111/1755-0998.12910).
Changes after Dec 25, 2021:
The previous data package had been uploaded before the manuscript was submitted. However, while the manuscript was being revised, the data was changed but not updated on Datadryad. This update represents the final version.
- Xia, Lei; Shi, Mijuan; Li, Heng et al. (2024). PMSeeker: A Scheme Based on the Greedy Algorithm and the Exhaustive Algorithm to Screen Low-Redundancy Marker Sets for Large-Scale Parentage Assignment with Full Parental Genotyping. Biology. https://doi.org/10.3390/biology13020100
