Additive effects of multiple loci underlie fruit sweetness variation in melon
Data files
Sep 11, 2026 version files 132.08 MB
-
melon_sweetness_dataset.zip
132.05 MB
-
README.md
32.73 KB
Abstract
This dataset supports a manuscript currently under review.
This dataset contains processed population genomic, transcriptomic, phenotypic, and variant call data supporting the analyses presented in the manuscript.
The dataset includes:
- Genome-wide population genetic statistics (FST, XP-CLR, LD decay, local LD structure, allele frequencies)
- Transcriptome data (gene expression count matrices, WGCNA module assignments for all genes, and differential expression statistics for sugar metabolism and transport-related genes)
- Phenotypic measurements (TSS)
- Variant call data (raw and filtered VCF files)
Filtered high-confidence variants (PASS) were used for marker development and functional analyses. All VCF files are provided in bgzip-compressed format with tabix indexing.
Dataset DOI: 10.5061/dryad.tqjq2bwdh
Overview
This dataset contains processed data supporting the population genomic, transcriptomic, and phenotypic analyses presented in the associated manuscript, “Additive multilocus architecture underlies fruit sweetness variation in melon.” The dataset is intended to be understandable and reusable independently of the manuscript.
The dataset is organized into four directories:
Population_genomics/: population genomic summary statistics and local analyses around the tss9.1 and tss10.1 TSS-QTL regions.Transcriptome/: gene-level RNA-seq read counts and WGCNA-related datasets.Phenotype/: TSS measurements used for WGCNA and module–trait correlation analyses.Variants/: VCF files containing variant calls from whole-genome resequencing of the parental lines.
Unless otherwise noted, tabular files are tab-delimited text files (.tsv). VCF files are bgzip-compressed and accompanied by tabix index files (.tbi).
Raw whole-genome resequencing and RNA-seq data are available through NCBI SRA under BioProject PRJNA1442044. Population genomic analyses were based on publicly available melon resequencing data from Liu et al. (2020).
Experimental and analytical context
The study investigated the genetic architecture of total soluble solids (TSS), used as a measure of fruit sweetness, in melon. A recombinant inbred line (RIL) population derived from the elite sweet melon cultivars ‘B2’ and ‘Pearl’ was evaluated over four cultivation years. The transcriptomic datasets comprise fruit samples from ‘B2’ and ‘Pearl’ collected at 0, 2, 15, 29, 36, 43, and 55 days after pollination (DAP), with three biological replicates per cultivar and developmental stage.
Four stable TSS QTLs were identified by multi-year QTL analysis. Subsequent sequence and transcriptomic analyses focused on tss9.1 and tss10.1. Population genomic analyses around these two regions compared improved and landrace melon groups derived from the public resequencing dataset of Liu et al. (2020).
1. Population genomic data
These data were derived from publicly available melon resequencing data from Liu et al. (2020) and processed as described in the associated manuscript.
For the population genomic analyses, 30 improved cultivars and 30 traditional landraces were selected for each of C. melo ssp. melo and ssp. agrestis, giving 60 accessions per population category and 120 accessions in total. Subsequent analyses compared the combined Improved (IMP_M + IMP_A) and Landrace (LDR_M + LDR_A) groups.
Population genomic analyses were performed using the DHL92 reference genome (CM3.6.1). Genomic coordinates in the population genomic files follow this reference assembly.
1.1 FST (population differentiation)
Files
FST_chr09.tsvFST_chr10.tsv
These files contain window-based FST estimates between the Improved and Landrace groups, calculated using the Weir and Cockerham estimator implemented in BCFtools (+weir-fst-pop). The data were used to evaluate population differentiation around the TSS-QTL regions.
Columns
| Column | Description |
|---|---|
CHROM |
Chromosome identifier. |
BIN_START |
Start coordinate of the genomic window (bp). |
BIN_END |
End coordinate of the genomic window (bp). |
N_VARIANTS |
Number of SNPs included in the window. |
WEIGHTED_FST |
Weighted FST estimate for the window. |
MEAN_FST |
Mean FST estimate for the window. |
Parameters
- Window size: 20 kb.
- Step size: 20 kb.
1.2 XP-CLR (cross-population composite likelihood ratio)
Files
XPCLR_chr09.tsvXPCLR_chr10.tsv
These files contain XP-CLR statistics calculated to detect genomic regions showing evidence of differentiated selection between the Landrace and Improved groups. The data were used to evaluate selection signals around tss9.1 and tss10.1.
Columns
| Column | Description |
|---|---|
id |
Identifier for the genomic window. |
chrom |
Chromosome identifier. |
start |
Start coordinate of the genomic window (bp). |
stop |
End coordinate of the genomic window (bp). |
pos_start |
Physical position of the first SNP considered in the window (bp). |
pos_stop |
Physical position of the last SNP considered in the window (bp). |
modelL |
Composite likelihood under the selection model. |
nullL |
Composite likelihood under the null model. |
sel_coef |
Estimated selection coefficient from the XP-CLR calculation. |
nSNPs |
Number of SNPs in the window used by the XP-CLR calculation. |
nSNPs_avail |
Number of SNPs available for the XP-CLR calculation. |
xpclr |
XP-CLR score. |
xpclr_norm |
Normalized XP-CLR score. |
Parameters
- Additional MAF filter: MAF >= 0.05.
- Window size: 20 kb.
- Step size: 1 kb.
--minsnps 3;--maxsnps 300;--size 20000;--step 1000.
XP-CLR values were interpreted relative to genome-wide background levels.
1.3 Linkage disequilibrium (LD)
LD decay
Files
LD_decay_chr09_improved.tsvLD_decay_chr09_landrace.tsvLD_decay_chr10_improved.tsvLD_decay_chr10_landrace.tsv
These files contain LD-decay data for the indicated chromosome and population group. Each row contains a physical-distance value and the corresponding mean pairwise LD measured as r². The files do not contain column headers.
File structure
| Position | Description |
|---|---|
| Column 1 | Physical distance between SNP pairs (bp). |
| Column 2 | Mean pairwise linkage disequilibrium (r²) at the corresponding physical distance. |
The filename identifies the chromosome (chr09 or chr10) and population group (improved or landrace).
LD decay was estimated separately for the Landrace and Improved populations using PLINK. Pairwise LD was calculated using --r2, --ld-window 99999, --ld-window-kb 500, and --ld-window-r2 0. LD decay distance was defined as the physical distance at which mean pairwise r² decayed to 0.2. No LD pruning was applied to the genome-wide LD dataset.
Local LD block lengths
Files
Local_LD_length_tss9.1_improved.tsvLocal_LD_length_tss9.1_landrace.tsvLocal_LD_length_tss10.1_improved.tsvLocal_LD_length_tss10.1_landrace.tsv
These files contain local LD-block information around the indicated TSS-QTL region. LD block lengths were calculated for the regions shown in the local LD analysis. LD blocks longer than 3 kb were rare and were not shown in the corresponding figure.
Columns
| Column | Description |
|---|---|
CHR |
Chromosome number. |
BP1 |
Physical position of the first SNP in the LD block (bp). |
BP2 |
Physical position of the last SNP in the LD block (bp). |
KB |
Physical length of the LD block in kilobases. |
NSNPS |
Number of SNPs included in the LD block. |
SNPS |
SNP identifiers included in the LD block, separated by `.` |
length |
LD-block length in kilobases as reported in the supplied block-length output. |
group |
Population group, when present (Improved or Landrace). |
The group column is present in the tss9.1 files and is not present in the tss10.1 files.
Local LD matrices
Files
Local_LD_matrix_tss9.1_improved.tsvLocal_LD_matrix_tss9.1_improved_matrix_label.tsvLocal_LD_matrix_tss9.1_landrace.tsvLocal_LD_matrix_tss9.1_landrace_matrix_label.tsvLocal_LD_matrix_tss10.1_improved.tsvLocal_LD_matrix_tss10.1_improved_matrix_label.tsvLocal_LD_matrix_tss10.1_landrace.tsvLocal_LD_matrix_tss10.1_landrace_matrix_label.tsv
The Local_LD_matrix_*.tsv files contain pairwise LD (r²) matrices used to visualize local LD structure around the candidate QTL regions in Fig. 4.
Local LD matrices were calculated in PLINK after SNP pruning using --indep-pairwise 50 5 0.8 to reduce marker redundancy for visualization. LD heatmaps were generated in R using ggplot2.
The matrix files do not contain column headers. They are square matrices in which rows and columns correspond to the SNPs listed, in the same order, in the corresponding _matrix_label.tsv file.
Matrix structure
- Each row corresponds to one SNP.
- Each column corresponds to one SNP.
- The row/column order is defined by the corresponding
_matrix_label.tsvfile. - Matrix entries are pairwise LD values (
r²). - The diagonal represents self-comparisons and is expected to be 1.
- Missing matrix entries are represented as
nanin the supplied files.
The corresponding _matrix_label.tsv files contain one SNP identifier per line and do not contain a column header.
SNP dataset summary
The principal SNP datasets used for population genomic analyses were:
| Dataset | Filtering / pruning | Purpose | Chr. 9 | Chr. 10 |
|---|---|---|---|---|
| Raw SNP set | BCFtools call | Variant discovery | 766,970 | 913,093 |
| Population structure | QUAL >= 30, DP >= 5; --max-missing 0.8; MAF >= 0.05; PLINK --indep-pairwise 50 10 0.2 |
Population structure | 85,254 | 100,402 |
| Genome-wide LD | QUAL >= 30, DP >= 5; no LD pruning | LD decay | 597,530 | 720,155 |
| FST | QUAL >= 30, DP >= 5; --max-missing 0.8; 0–5 Mb |
FST | 124,660 | 92,206 |
| XP-CLR | QUAL >= 30, DP >= 5; --max-missing 0.8; MAF >= 0.05 |
XP-CLR | 68,061 | 43,611 |
| Local LD | QUAL >= 30, DP >= 5; --max-missing 0.8; PLINK --indep-pairwise 50 5 0.8 |
Local LD heatmap; Δ allele frequency | 2,423 | 1,154 |
SNPs used for Δ allele frequency estimation were located within the genomic intervals 2200–2700 kb on chromosome 9 (including tss9.1) and 0–300 kb on chromosome 10 (including tss10.1).
1.4 Allele frequency
Files
Allele_frequency_tss9.1_improved.tsvAllele_frequency_tss9.1_landrace.tsvAllele_frequency_tss10.1_improved.tsvAllele_frequency_tss10.1_landrace.tsv
These files contain site-level allele-frequency summaries for SNPs in the tss9.1 and tss10.1 regions for the indicated population group. Each file contains 60 taxa.
Columns
| Column | Description |
|---|---|
Site Number |
Sequential site number in the input dataset. |
Site Name |
SNP/site identifier. |
Chromosome |
Chromosome number. |
Physical Position |
Physical genomic position (bp). |
Number of Taxa |
Number of taxa included at the site. |
Ref |
Reference allele. |
Alt |
Alternate allele. |
Major Allele |
Allele with the highest observed frequency. |
Major Allele Gametes |
Number of observed gametes carrying the major allele. |
Major Allele Proportion |
Proportion of observed gametes carrying the major allele. |
Major Allele Frequency |
Frequency of the major allele. |
Minor Allele |
Allele with the second-highest observed frequency. |
Minor Allele Gametes |
Number of observed gametes carrying the minor allele. |
Minor Allele Proportion |
Proportion of observed gametes carrying the minor allele. |
Minor Allele Frequency |
Frequency of the minor allele. |
Allele 3 |
Third observed allele, when present. |
Allele 3 Gametes |
Number of observed gametes carrying allele 3. |
Allele 3 Proportion |
Proportion of observed gametes carrying allele 3. |
Allele 3 Frequency |
Frequency of allele 3. |
Allele 4 |
Fourth observed allele, when present. |
Allele 4 Gametes |
Number of observed gametes carrying allele 4. |
Allele 4 Proportion |
Proportion of observed gametes carrying allele 4. |
Allele 4 Frequency |
Frequency of allele 4. |
Allele 5 |
Fifth observed allele, when present. |
Allele 5 Gametes |
Number of observed gametes carrying allele 5. |
Allele 5 Proportion |
Proportion of observed gametes carrying allele 5. |
Allele 5 Frequency |
Frequency of allele 5. |
Allele 6 |
Sixth observed allele, when present. |
Allele 6 Gametes |
Number of observed gametes carrying allele 6. |
Allele 6 Proportion |
Proportion of observed gametes carrying allele 6. |
Allele 6 Frequency |
Frequency of allele 6. |
Gametes Missing |
Number of missing gametes at the site. |
Proportion Missing |
Proportion of missing gametes at the site. |
Number Heterozygous |
Number of heterozygous taxa at the site. |
Proportion Heterozygous |
Proportion of heterozygous taxa. |
Inbreeding Coefficient |
Inbreeding coefficient calculated for the site. |
Inbreeding Coefficient Scaled by Missing |
Inbreeding coefficient adjusted/scaled according to missing data. |
NA indicates that the corresponding additional allele was not observed or was not applicable at that site. nan is used only in the local LD matrix files to represent missing matrix entries. TBD in the allele-frequency files indicates that the corresponding inbreeding-coefficient value was not determined in the supplied output.
2. Transcriptome data
2.1 Gene-level read counts
File
Gene_count_matrix.tsv
This file contains gene-level RNA-seq read counts for ‘B2’ and ‘Pearl’. Raw reads were quality-filtered using fastp v0.23.4 and aligned to the Harukei-3 reference genome using HISAT2 v2.2.1. Gene-level read counts were generated using featureCounts v2.0.6 with -s 2 -p -C -Q 30 -t exon and the Harukei-3 gene annotation. Genes with fewer than 50 total reads across all samples were excluded from downstream analyses. Differential expression analysis was performed using DESeq2 v1.42.1, with each developmental stage compared against 0 DAP within each accession. DEGs were defined as FDR < 0.01 and |log2FC| >= 0.5.
The dataset includes fruit samples collected at 0, 2, 15, 29, 36, 43, and 55 DAP, with three biological replicates per cultivar and developmental stage.
The first line of the file is featureCounts program metadata. The second line contains the column headers, followed by the gene-level count data.
Columns
| Column group | Description |
|---|---|
Geneid |
Gene identifier. |
Chr |
Chromosome or genomic feature location assigned by featureCounts. |
Start |
Start coordinate of the annotated genomic feature (bp). |
End |
End coordinate of the annotated genomic feature (bp). |
Strand |
Strand information for the annotated genomic feature. |
Length |
Feature length used by featureCounts (bp). |
| Sample columns | Raw gene-level read counts for each RNA-seq sample. |
Sample columns follow the format Cultivar_DAP_replicate_sorted.bam. For example, B2_15D_1_sorted.bam represents ‘B2’, 15 DAP, biological replicate 1. The sample columns cover both cultivars, all seven developmental stages, and three biological replicates per stage.
2.2 WGCNA module assignments
Files
Supplementary_Data_1_WGCNA_module_assignment_Pearl.tsvSupplementary_Data_2_WGCNA_module_assignment_B2.tsv
These files provide WGCNA module assignments for genes used in network construction for ‘Pearl’ and ‘B2’, respectively. WGCNA was performed separately for ‘Pearl’ and ‘B2’ using WGCNA v1.73 in R. Gene expression values were transformed using variance-stabilizing transformation (VST; blind = FALSE). Expression data from 15, 29, 36, 43, and 55 DAP were used for network construction.
Genes used for WGCNA network construction were defined as those with FDR < 0.01 in either the 43 DAP vs 15 DAP or 55 DAP vs 15 DAP comparison. No fold-change threshold was applied to this WGCNA gene-selection criterion. A signed network was constructed. Soft-thresholding powers were selected using an approximate scale-free topology criterion of R² > 0.8, resulting in β = 12 for ‘Pearl’ and β = 16 for ‘B2’. Modules were identified with a minimum module size of 30 genes and a merge cut height of 0.25. Module eigengenes were correlated with TSS using Pearson’s correlation coefficient, and P values were adjusted using the Benjamini–Hochberg method; adjusted P < 0.05 was considered statistically significant.
Columns
| Column | Description |
|---|---|
module.gene |
Gene identifier. |
mergedColors |
WGCNA module assigned to the gene after module merging. |
The module assignments correspond to the co-expression networks used in Fig. 3A.
2.3 Starch and sucrose metabolism genes
File
Supplementary_Data_3_WGCNA_starch_sucrose_genes.tsv
This file contains gene-level information for starch and sucrose metabolism-related genes underlying Fig. 3A, including WGCNA module assignments, expression statistics, differential-expression statistics, and expression values across developmental stages.
Columns
| Column group | Description |
|---|---|
Gene_id |
Melon gene identifier. |
Entrezgene_ID |
NCBI Entrez Gene identifier, when available. |
Gene_description |
Functional gene description. |
WGCNA module (B2) |
WGCNA module assigned to the gene in ‘B2’. |
baseMean 43vs15 B2 |
DESeq2 baseMean expression value for the 43 DAP versus 15 DAP comparison in ‘B2’. |
log2 fold change 43vs15 B2 |
log2 fold change for 43 DAP relative to 15 DAP in ‘B2’. |
FDR 43vs15 B2 |
FDR-adjusted significance value for the 43 DAP versus 15 DAP comparison in ‘B2’. |
baseMean 55vs15 B2 |
DESeq2 baseMean expression value for the 55 DAP versus 15 DAP comparison in ‘B2’. |
log2 fold change 55vs15 B2 |
log2 fold change for 55 DAP relative to 15 DAP in ‘B2’. |
FDR 55vs15 B2 |
FDR-adjusted significance value for the 55 DAP versus 15 DAP comparison in ‘B2’. |
WGCNA module (Pearl) |
WGCNA module assigned to the gene in ‘Pearl’. |
baseMean 43vs15 Pearl |
DESeq2 baseMean expression value for the 43 DAP versus 15 DAP comparison in ‘Pearl’. |
log2 fold change 43vs15 Pearl |
log2 fold change for 43 DAP relative to 15 DAP in ‘Pearl’. |
FDR 43vs15 Pearl |
FDR-adjusted significance value for the 43 DAP versus 15 DAP comparison in ‘Pearl’. |
baseMean 55vs15 Pearl |
DESeq2 baseMean expression value for the 55 DAP versus 15 DAP comparison in ‘Pearl’. |
log2 fold change 55vs15 Pearl |
log2 fold change for 55 DAP relative to 15 DAP in ‘Pearl’. |
FDR 55vs15 Pearl |
FDR-adjusted significance value for the 55 DAP versus 15 DAP comparison in ‘Pearl’. |
B2 [DAP] [replicate] |
VST-transformed gene expression value for the indicated ‘B2’ sample. |
Pearl [DAP] [replicate] |
VST-transformed gene expression value for the indicated ‘Pearl’ sample. |
For the sample-expression columns, [DAP] is one of 0D, 2D, 15D, 29D, 36D, 43D, or 55D, and [replicate] is biological replicate 1, 2, or 3.
For differential-expression analyses, log2 fold-change values represent expression at 43 or 55 DAP relative to 15 DAP. Differentially expressed genes (DEGs) were defined as FDR < 0.01 and |log2FC| ≥ 0.5. This DEG definition differs from the WGCNA gene-selection criterion, which was based on FDR significance without a fold-change threshold.
2.4 SWEET family genes
File
Supplementary_Data_4_WGCNA_SWEET_genes.tsv
This file contains gene-level information for SWEET family genes underlying Fig. 3A, including Arabidopsis homolog identifiers, WGCNA module assignments, expression statistics, differential-expression statistics, and expression values across developmental stages.
Columns
| Column group | Description |
|---|---|
Gene ID |
Melon gene identifier. |
Entrezgene ID |
NCBI Entrez Gene identifier, when available. |
Arabidopsis homolog (TAIR ID) |
Arabidopsis homolog identifier from The Arabidopsis Information Resource (TAIR), when available. |
WGCNA module (b2) |
WGCNA module assigned to the gene in ‘B2’. |
baseMean 43vs15 b2 |
DESeq2 baseMean expression value for 43 DAP versus 15 DAP in ‘B2’. |
log2 fold change 43vs15 b2 |
log2 fold change for 43 DAP relative to 15 DAP in ‘B2’. |
FDR 43vs15 b2 |
FDR-adjusted significance value for 43 DAP versus 15 DAP in ‘B2’. |
baseMean 55vs15 b2 |
DESeq2 baseMean expression value for 55 DAP versus 15 DAP in ‘B2’. |
log2 fold change 55vs15 b2 |
log2 fold change for 55 DAP relative to 15 DAP in ‘B2’. |
FDR 55vs15 b2 |
FDR-adjusted significance value for 55 DAP versus 15 DAP in ‘B2’. |
WGCNA module (Pearl) |
WGCNA module assigned to the gene in ‘Pearl’. |
baseMean 43vs15 Pearl |
DESeq2 baseMean expression value for 43 DAP versus 15 DAP in ‘Pearl’. |
log2 fold change 43vs15 Pearl |
log2 fold change for 43 DAP relative to 15 DAP in ‘Pearl’. |
FDR 43vs15 Pearl |
FDR-adjusted significance value for 43 DAP versus 15 DAP in ‘Pearl’. |
baseMean 55vs15 Pearl |
DESeq2 baseMean expression value for 55 DAP versus 15 DAP in ‘Pearl’. |
log2 fold change 55vs15 Pearl |
log2 fold change for 55 DAP relative to 15 DAP in ‘Pearl’. |
FDR 55vs15 Pearl |
FDR-adjusted significance value for 55 DAP versus 15 DAP in ‘Pearl’. |
B2 [DAP] [replicate] |
VST-transformed gene expression value for the indicated ‘B2’ sample. |
Pearl [DAP] [replicate] |
VST-transformed gene expression value for the indicated ‘Pearl’ sample. |
The differential-expression and missing-value conventions are the same as those described for Supplementary_Data_3_WGCNA_starch_sucrose_genes.tsv.
3. Phenotypic data
TSS measurements
File
TSS_data.tsv
This file contains total soluble solids (TSS) measurements used as phenotypic trait data for WGCNA module–trait correlation analyses.
Columns
| Column | Description |
|---|---|
Material_id |
Sample/material identifier. |
Accession |
Cultivar/accession identifier as provided in the original data file. |
DAP |
Days after pollination. |
Replication |
Biological replicate number. |
TSS |
Total soluble solids measured in °Brix. |
The TSS dataset contains measurements from ‘B2’ and ‘Pearl’ across the developmental stages represented in the file. Missing values, if present, are retained as missing rather than inferred or imputed.
4. Variant call data
Files
variants_raw.vcf.gzvariants_raw.vcf.gz.tbivariants_snps_PASS.vcf.gzvariants_snps_PASS.vcf.gz.tbivariants_indels_PASS.vcf.gzvariants_indels_PASS.vcf.gz.tbi
These files contain variant calls generated from whole-genome resequencing of the parental lines ‘B2’ and ‘Pearl’. The VCF files are VCF version 4.2, bgzip-compressed, and indexed with tabix.
The two samples in the VCF files are P15-3b and P85-3a.
4.1 VCF file types
variants_raw.vcf.gz: raw, unfiltered variant calls containing SNPs and indels before the final PASS filtering.variants_snps_PASS.vcf.gz: SNP variants that passed the applied GATK hard-filtering criteria.variants_indels_PASS.vcf.gz: indel variants that passed the applied GATK hard-filtering criteria.
4.2 VCF fields
Standard VCF columns are:
| Field | Description |
|---|---|
#CHROM |
Chromosome or contig identifier. |
POS |
One-based genomic position of the variant. |
ID |
Variant identifier, when available. |
REF |
Reference allele. |
ALT |
Alternate allele(s). |
QUAL |
Variant quality score. |
FILTER |
Variant filter status. PASS indicates that the variant passed the applied filters. |
INFO |
Additional site-level information. |
FORMAT |
Format definition for sample genotype fields. |
| Sample fields | Genotype and supporting information for the indicated parental sample. |
The supplied VCFs include standard genotype and read-support fields such as GT, AD, DP, GQ, MIN_DP, PGT, PID, and PL. The VCF header provides their formal definitions.
4.3 Filtering
Raw variants were hard-filtered using GATK VariantFiltration. For SNPs:
- QD < 2.0
- QUAL < 30.0
- SOR > 4.0
- FS > 60.0
- MQ < 40.0
- MQRankSum < -12.5
- ReadPosRankSum < -8.0
For insertions and deletions (InDels):
- QD < 2.0
- QUAL < 30.0
- FS > 200.0
- SOR > 10.0
- ReadPosRankSum < -20.0
Only variants labeled PASS were retained in the corresponding PASS files. For marker development, high-confidence homozygous variants with a minimum read depth of 10 were selected from the PASS variants using VCFtools v0.1.16.
The .tbi files are tabix indexes corresponding to their respective compressed VCF files and are provided to facilitate region-based queries.
Genomic coordinates and chromosome naming conventions follow the reference genome used for the variant-calling analysis.
5. Missing values and data conventions
NAin the allele-frequency files indicates that the corresponding additional allele was not observed or was not applicable at that site.nanin local LD matrix files represents a missing matrix value.TBDin the allele-frequency files indicates that the corresponding inbreeding-coefficient value was not determined in the supplied output.- No values were intentionally inferred or imputed to replace missing observations in the TSS dataset.
- Empty cells in WGCNA-derived datasets indicate that the corresponding statistic or module assignment was unavailable/not applicable for that cultivar or comparison. In particular, an empty WGCNA module or statistic should not be interpreted as zero.
6. Abbreviations and terminology
- B2: parental melon cultivar ‘B2’.
- Pearl: parental melon cultivar ‘Pearl’.
- DAP: days after pollination.
- TSS: total soluble solids.
- °Brix: unit used for TSS measurements.
- QTL: quantitative trait locus.
- FST: fixation index measuring population differentiation.
- XP-CLR: cross-population composite likelihood ratio.
- LD: linkage disequilibrium.
- r²: squared correlation coefficient used as the measure of pairwise LD.
- SNP: single-nucleotide polymorphism.
- VCF: variant call format.
- WGCNA: weighted gene co-expression network analysis.
- DESeq2: differential-expression analysis software used for RNA-seq data.
- DEG: differentially expressed gene.
- FDR: false discovery rate.
- log2FC: log2 fold change.
- TAIR: The Arabidopsis Information Resource.
- GATK: Genome Analysis Toolkit.
- BCFtools: utilities for processing variant call format and genomic variant data.
- SRA: Sequence Read Archive.
- NCBI: National Center for Biotechnology Information.
- VST: variance-stabilizing transformation.
7. Figure and analysis correspondence
- Fig. 3A: WGCNA-derived datasets in
Transcriptome/Supplementary_Data_1_WGCNA_module_assignment_Pearl.tsv,Supplementary_Data_2_WGCNA_module_assignment_B2.tsv,Supplementary_Data_3_WGCNA_starch_sucrose_genes.tsv, andSupplementary_Data_4_WGCNA_SWEET_genes.tsvprovide the gene-level information used to show the distribution of carbohydrate metabolism and SWEET family genes across co-expression modules. - Fig. 4A:
Population_genomics/LD_decay_*.tsvfiles provide the genome-wide LD-decay data for Improved and Landrace populations on chromosomes 9 and 10. LD decay distance was evaluated at r² = 0.2. - Fig. 4B:
Local_LD_matrix_*.tsvand corresponding_matrix_label.tsvfiles provide the local LD values and SNP ordering used to reconstruct the local LD visualizations. The allele-frequency files provide site-level allele-frequency information around the candidate loci. - Fig. 4C:
Local_LD_length_*.tsvfiles provide LD-block length information used for the LD-block length distributions. - Fig. 4D:
FST_chr*.tsvandXPCLR_chr*.tsvprovide the FST and XP-CLR datasets used to evaluate population differentiation and localized selection signals. Phenotype/TSS_data.tsvprovides TSS phenotypic values used in WGCNA module–trait correlation analyses.- The VCF files provide the parental variant data used for variant characterization and downstream analyses.
8. References
Liu S, Gao P, Zhu Q, Zhu Z, Liu H, Wang X, Weng Y, Gao M, Luan F. 2020. Resequencing of 297 melon accessions reveals the genomic history of improvement and loci related to fruit traits in melon. Plant Biotechnology Journal.
Yano R, et al. 2020. Genome sequence and expression atlas of the melon cultivar Harukei-3. [Reference as cited in the associated manuscript.]
9. Data access
Raw sequencing data are available through the NCBI SRA under BioProject PRJNA1442044.
The processed datasets are intended to support independent reuse and reproduction of the analyses described above. Users should consult the associated manuscript for the complete experimental design and interpretation.
10. Contact
For questions regarding the data, please contact:
Katsunori TANAKA
Faculty of Agriculture and Life Science, Hirosaki University
k-tana3@hirosaki-u.ac.jp
