WGS Metadata Attributes

Fields that are collected for WGS data, available at dataset.metadata.<attribute>  

* indicates a required field

Attribute Type Description Allowable Values
version * Version of the schema to use when validating this metadata. '1'
description Free-text description of this assay.  
donor_id HuBMAP Display ID of the donor of the assayed tissue.  
tissue_id HuBMAP Display ID of the assayed tissue.  
execution_datetime Start date and time of assay, typically a date-time stamped folder generated by the acquisition instrument. YYYY-MM-DD hh:mm, where YYYY is the year, MM is the month with leading 0s, and DD is the day with leading 0s, hh is the hour with leading zeros, mm are the minutes with leading zeros.  
protocols_io_doi DOI for protocols.io referring to the protocol for this assay.  
operator Name of the person responsible for executing the assay.  
operator_email Email address for the operator.  
pi Name of the principal investigator responsible for the data.  
pi_email Email address for the principal investigator.  
assay_category Each assay is placed into one of the following 4 general categories: generation of images of microscopic entities, identification & quantitation of molecules by mass spectrometry, imaging mass spectrometry, and determination of nucleotide sequence. 'sequence'
assay_type The specific type of assay being executed. 'WGS'
analyte_class Analytes are the target molecules being measured with the assay. 'DNA'
is_targeted Specifies whether or not a specific molecule(s) is/are targeted for detection/measurement by the assay. 'Yes' 'No'
acquisition_instrument_vendor An acquisition instrument is the device that contains the signal detection hardware and signal processing software. Assays generate signals such as light of various intensities or color or signals representing the molecular mass.  
acquisition_instrument_model Manufacturers of an acquisition instrument may offer various versions (models) of that instrument with different features or sensitivities. Differences in features or sensitivities may be relevant to processing or interpretation of the data.  
gdna_fragmentation_quality_assurance Is the gDNA integrity good enough for WGS? This is usually checked through running a gel. 'Pass' 'Fail'
dna_assay_input_value Amount of DNA input into library preparation  
dna_assay_input_unit Units of DNA input into library preparation 'ug'
library_construction_method * Describes DNA library preparation kit. Modality of isolating gDNA, Fragmentation and generating sequencing libraries.  
library_construction_protocols_io_doi A link to the protocol document containing the library construction method (including version) that was used.  
library_layout State whether the library was generated for single-end or paired end sequencing. 'single-end' 'paired-end'
library_adapter_sequence The adapter sequence to be used for adapter trimming starting with the 5’ end. (eg. 5-ATCCTGAGAA)  
library_final_yield Total amount of library after final pcr amplification step  
library_final_yield_unit Total units of library after final pcr amplification step 'ng'
library_average_fragment_size * Average size in basepairs (bp) of sequencing library fragments estimated via gel electrophoresis or bioanalyzer/tapestation.  
sequencing_reagent_kit Reagent kit used for sequencing  
sequencing_read_format Slash-delimited list of the number of sequencing cycles for, for example, Read1, i7 index, i5 index, and Read2.  
sequencing_read_percent_q30 Q30 is the weighted average of all the reads (e.g. # bases UMI * q30 UMI + # bases R2 * q30 R2 + …)  
sequencing_phix_percent Percent PhiX loaded to the run  
contributors_path Relative path to file with ORCID IDs for contributors for this dataset.  
data_path Relative path to file or directory with instrument data. Downstream processing will depend on filename extension conventions.  

 

Deprecated Attributes

* indicates a field that was previously required

Attribute Type Description Allowable Values
donor_id HuBMAP Display ID of the donor of the assayed tissue.  
tissue_id HuBMAP Display ID of the assayed tissue.  
execution_datetime Start date and time of assay, typically a date-time stamped folder generated by the acquisition instrument. YYYY-MM-DD hh:mm, where YYYY is the year, MM is the month with leading 0s, and DD is the day with leading 0s, hh is the hour with leading zeros, mm are the minutes with leading zeros.  
protocols_io_doi DOI for protocols.io referring to the protocol for this assay.  
operator Name of the person responsible for executing the assay.  
operator_email Email address for the operator.  
pi Name of the principal investigator responsible for the data.  
pi_email Email address for the principal investigator.  
assay_category Each assay is placed into one of the following 4 general categories: generation of images of microscopic entities, identification & quantitation of molecules by mass spectrometry, imaging mass spectrometry, and determination of nucleotide sequence. 'sequence'
assay_type The specific type of assay being executed. 'WGS'
analyte_class Analytes are the target molecules being measured with the assay. 'DNA'
is_targeted Specifies whether or not a specific molecule(s) is/are targeted for detection/measurement by the assay. 'Yes' 'No'
acquisition_instrument_vendor An acquisition instrument is the device that contains the signal detection hardware and signal processing software. Assays generate signals such as light of various intensities or color or signals representing the molecular mass.  
acquisition_instrument_model Manufacturers of an acquisition instrument may offer various versions (models) of that instrument with different features or sensitivities. Differences in features or sensitivities may be relevant to processing or interpretation of the data.  
gdna_fragmentation_quality_assurance Is the gDNA integrity good enough for WGS? This is usually checked through running a gel. 'Pass' 'Fail'
dna_assay_input_value Amount of DNA input into library preparation  
dna_assay_input_unit Units of DNA input into library preparation 'ug'
library_construction_method Describes DNA library preparation kit. Modality of isolating gDNA, Fragmentation and generating sequencing libraries.  
library_construction_protocols_io_doi A link to the protocol document containing the library construction method (including version) that was used.  
library_layout State whether the library was generated for single-end or paired end sequencing. 'single-end' 'paired-end'
library_adapter_sequence The adapter sequence to be used for adapter trimming starting with the 5’ end. (eg. 5-ATCCTGAGAA)  
library_final_yield Total amount of library after final pcr amplification step  
library_final_yield_unit Total units of library after final pcr amplification step 'ng'
library_average_fragment_size Average size in basepairs (bp) of sequencing library fragments estimated via gel electrophoresis or bioanalyzer/tapestation.  
sequencing_reagent_kit Reagent kit used for sequencing  
sequencing_read_format Slash-delimited list of the number of sequencing cycles for, for example, Read1, i7 index, i5 index, and Read2.  
sequencing_read_percent_q30 Q30 is the weighted average of all the reads (e.g. # bases UMI * q30 UMI + # bases R2 * q30 R2 + …)  
sequencing_phix_percent Percent PhiX loaded to the run  
contributors_path Relative path to file with ORCID IDs for contributors for this dataset.  
data_path Relative path to file or directory with instrument data. Downstream processing will depend on filename extension conventions.