RNAseq Metadata Attributes

Fields that are collected for RNAseq data, available at dataset.metadata.<attribute>  

* indicates a required field

Attribute Type Description Allowable Values
parent_sample_id Unique HuBMAP or SenNet identifier of the sample (i.e., block, section or suspension) used to perform this assay. For example, for a RNAseq assay, the parent would be the suspension, whereas, for one of the imaging assays, the parent would be the tissue section. If an assay comes from multiple parent samples then this should be a comma separated list. Example: HBM386.ZGKG.235, HBM672.MKPK.442 or SNT232.UBHJ.322, SNT329.ALSK.102  
lab_id A locally assigned identifier provided by the data provider for the dataset. It is used to reference an external metadata record that may be maintained independently, enabling traceability and supporting provenance tracking. Example: Visium_9OLC_A4_S1  
preparation_protocol_doi DOI for the protocols.io page that describes the assay or sample procurment and preparation. For example for an imaging assay, the protocol might include staining of a section through the creation of an OME-TIFF file. In this case the protocol would include any image processing steps required to create the OME-TIFF file. Example: https://dx.doi.org/10.17504/protocols.io.eq2lyno9qvx9/v1 https://dx.doi.org/10.17504/protocols.io.eq2lyno9qvx9/v1
dataset_type The specific type of dataset being produced.  
analyte_class Analytes are the target molecules being measured with the assay.  
is_targeted Specifies whether or not a specific molecule(s) is/are targeted for detection/measurement by the assay.  
acquisition_instrument_vendor An acquisition instrument is the device that contains the signal detection hardware and signal processing software. Assays generate signals such as light of various intensities or color or signals representing the molecular mass.  
acquisition_instrument_model Manufacturers of an acquisition instrument may offer various versions (models) of that instrument with different features or sensitivities. Differences in features or sensitivities may be relevant to processing or interpretation of the data.  
source_storage_duration_value How long was the source material stored, prior to this sample being processed? For assays applied to tissue sections, this would be how long the tissue section (e.g., slide) was stored, prior to the assay beginning (e.g., imaging). For assays applied to suspensions such as sequencing, this would be how long the suspension was stored before library construction began.  
source_storage_duration_unit The time duration unit of measurement  
time_since_acquisition_instrument_calibration_value The amount of time since the acqusition instrument was last serviced by the vendor. This provides a metric for assessing drift in data capture.  
time_since_acquisition_instrument_calibration_unit The time unit of measurement  
contributors_path Relative path to file with ORCID IDs for contributors for this dataset.  
data_path Relative path to file or directory with instrument data. Downstream processing will depend on filename extension conventions.  
barcode_offset Positions in the read at which the cell or capture spot barcodes start. Cell and capture spot barcodes are, for example, 3 x 8 bp sequences that are spaced by constant sequences (the offsets). First barcode at position 0, then 38, then 76. This should be included when constructing sequencing libraries with a non-commercial kit.  
barcode_read Which read file contains the cell or capture spot barcode. This should be included when constructing sequencing libraries with a non-commercial kit. This field is required if the source material is barcoded. This field is used to determine which analysis pipeline to run.  
barcode_size Length of the cell or capture spot barcode in base pairs. Cell and capture spot barcodes are, for example, 3 x 8 bp sequences that are spaced by constant sequences, the offsets. This should be included when constructing sequencing libraries with a non-commercial kit. This field is required if the source material is barcoded. This field is used to determine which analysis pipeline to run.  
umi_offset Position in the read at which the umi barcode starts.  
umi_read Which read file(s) contains the UMI (unique molecular identifier) barcode.  
umi_size Length of the umi barcode in base pairs.  
assay_input_entity This is the entity from which the analyte is being captured. For example, for bulk sequencing this would be “tissue”, while it would be “single cell” for single cell sequencing. This field is used to determine which analysis pipeline to run.  
number_of_input_cells_or_nuclei How many cells or nuclei were input to the assay? This is typically not available for preparations working with bulk tissue.  
amount_of_input_analyte_value The amount of RNA or DNA input to the assay, typically measured by a Qubit, BioAnalyzer, or TapeStation. In most single cell/nuclei assays, this value isn’t available.  
amount_of_input_analyte_unit Units of amount of entity input to assay value  
library_adapter_sequence Adapter sequence to be used for adapter trimming  
library_average_fragment_size Average size in basepairs (bp) of sequencing library fragments estimated via gel electrophoresis or bioanalyzer/tapestation.  
library_input_amount_value The amount of cDNA, after amplification, that was used for library construction.  
library_input_amount_unit unit of library input amount value  
library_output_amount_value Total amount (eg. nanograms) of library after the clean-up step of final pcr amplification step. Answer the question: What is the Qubit measured concentration (ng/ul) times the elution volume (ul) after the final clean-up step?  
library_output_amount_unit Units of library final yield.  
library_concentration_value The concentration value of the pooled library samples submitted for sequencing.  
library_concentration_unit Unit of library concentration value.  
library_layout State whether the library was generated for single-end or paired end sequencing.  
number_of_iterations_of_cdna_amplification This is the amplification of the cDNA prior to library construction. This is typically a PCR amplification, while for linear amplification methods like aRNA this would be the number of rounds of aRNA.  
number_of_pcr_cycles_for_indexing Number of PCR cycles performed in order to add adapters and amplify the library. This does not include the cDNA amplification which is captured in the “number of iterations of cDNA amplification” field.  
library_preparation_kit Reagent kit used for library preparation  
sample_indexing_kit Indexes are needed for multiplexing sequencing libraries for simultaneous sequencing (pooling) and proper attachment to the Illumina flowcell. Each indexing kit would have a number of compatible sequences (“sample indexing sets”) that are used to label some number of samples (the number of sets depend on the kit).  
sample_indexing_set The specific sequencing barcode index set used, selected from the sample indexing kit. Example: For 10X this might be “SI-GA-A1”, for Nextera “N505 - CTCCTTAC”  
is_technical_replicate Is the sequencing reaction run in replicate, TRUE or FALSE  
expected_entity_capture_count Number of cells, nuclei or capture spots expected to be captured by the assay. For Visium this is the total number of spots covered by tissue, within the capture area.  
sequencing_reagent_kit Reagent kit used for sequencing  
sequencing_read_format Slash-delimited list of the number of sequencing cycles for, for example, Read1, i7 index, i5 index, and Read2.  
sequencing_batch_id The ID for the sequencing run. This could, for example, be the chip ID and should allow users the ability to determine which samples were processed together in a sequencing run. It is recommended that data providers prefix the ID with the center name, to prevent values overlapping across centers.  
capture_batch_id A lab-generated ID to identify which cells were captured at the same time. This would, for example, be an ID to denote which datasets were derived from a single 10X Genomics Chromium Controller run. In the case of the 10X Controller this could be the chip ID and would allow users the ability to determine which samples were processed together in a Chromium controller. It is recommended that data providers prefix the ID with the center name, to prevent values overlapping across centers.  
preparation_instrument_vendor The manufacturer of the instrument used to prepare (staining/processing) the sample for the assay. If an automatic slide staining method was indicated this field should list the manufacturer of the instrument.  
preparation_instrument_model Manufacturers of a staining system instrument may offer various versions (models) of that instrument with different features. Differences in features or sensitivities may be relevant to processing or interpretation of the data.  
preparation_instrument_kit The reagent kit used with the preparation instrument.  
metadata_schema_id The string that serves as the definitive identifier for the metadata schema version and is readily interpretable by computers for data validation and processing. Example: 22bc762a-5020-419d-b170-24253ed9e8d9  

 

Deprecated Attributes

 

indicates a field that was previously required

Attribute Type Description Allowable Values
assay_category Each assay is placed into one of the following 4 general categories: generation of images of microscopic entities, identification & quantitation of molecules by mass spectrometry, imaging mass spectrometry, and determination of nucleotide sequence.  
bulk_rna_isolation_protocols_io_doi Link to a protocols document answering the question: How was tissue stored and processed for RNA isolation RNA_isolation_protocols_io_doi  
bulk_rna_isolation_quality_metric_value RIN value  
bulk_rna_yield_units_per_tissue_unit RNA amount per Tissue input amount. Valid values should be weight/weight (ng/mg).  
bulk_rna_yield_value RNA (ng) per Weight of Tissue (mg). Answer the question: How much RNA in ng was isolated? How much tissue in mg was initially used for isolating RNA? Calculate the yield by dividing total RNA isolated by amount of tissue used to isolate RNA from (ng/mg).  
donor_id HuBMAP Display ID of the donor of the assayed tissue.  
execution_datetime Start date and time of assay, typically a date-time stamped folder generated by the acquisition instrument. YYYY-MM-DD hh:mm, where YYYY is the year, MM is the month with leading 0s, and DD is the day with leading 0s, hh is the hour with leading zeros, mm are the minutes with leading zeros.  
library_construction_protocols_io_doi A link to the protocol document containing the library construction method (including version) that was used, e.g. “Smart-Seq2”, “Drop-Seq”, “10X v3”.  
library_id A library ID, unique within a TMC, which allows corresponding RNA and chromatin accessibility datasets to be linked.  
operator Name of the person responsible for executing the assay.  
operator_email Email address for the operator.  
pi Name of the principal investigator responsible for the data.  
pi_email Email address for the principal investigator.  
rnaseq_assay_method The kit used for the RNA sequencing assay  
sc_isolation_enrichment The method by which specific cell populations are sorted or enriched.  
sc_isolation_protocols_io_doi Link to a protocols document answering the question: How were single cells separated into a single-cell suspension?  
sc_isolation_quality_metric A quality metric by visual inspection prior to cell lysis or defined by known parameters such as wells with several cells or no cells. This can be captured at a high level.  
sc_isolation_tissue_dissociation The method by which tissues are dissociated into single cells in suspension.  
sc_isolation_cell_number Total number of cell/nuclei yielded post dissociation and enrichment  
sequencing_phix_percent Percent PhiX loaded to the run  
sequencing_read_percent_q30 Q30 is the weighted average of all the reads (e.g. # bases UMI * q30 UMI + # bases R2 * q30 R2 + …)  
version Version of the schema to use when validating this metadata.  
description Free-text description of this assay.