Occurrence download formats
Data downloads are available from GBIF in the following primary formats:
-
Simple. This format contains a selection of commonly used terms, after the data has been aligned to GBIF’s taxonomic and geographic indices and structured vocabularies
-
Downloads created on www.gbif.org or through the API using the format
SIMPLE_CSVare produced in a tab-separated text format, suitable for use with spreadsheets and programming/scripting languages -
Occurrence data accessed through cloud services, or with the API format
SIMPLE_PARQUET, are produced in Apache Parquet format. The fields are the same as for tab-separated text format.
-
-
Darwin Core Archive (API:
DWCA). This is a compressed Zip file, containing data in tab-separated text format, and metadata in XML format.-
occurrence.txtcontains occurrence data after interpretation by GBIF’s systems. -
multimedia.txtcontains information on multimedia (images, audio, video) relating to the occurrences. -
verbatim.txtcontains the original, uninterpreted data, without modifications by GBIF’s systems. -
optionally, additional verbatim Darwin Core Archive extensions. The data are as-received from the publisher. See GBIF Registered Extensions for documentation of these — note not all of them are maintained by GBIF.
-
-
FASTA Archive (API:
FASTA_ARCHIVE). This is a Darwin Core Archive download of occurrences with DNA or RNA sequences, with the sanitized sequences added in FASTA format, together with their quality metrics. -
Species List (API:
SPECIES_LIST). This is a summary format containing the distinct list of species names returned by the filter. -
Cube. This format allows you to aggregate occurrences by their taxonomic, temporal and/or spatial properties.
The header row (first row) of all these files contain the short name of the terms they contain. Most of the terms are defined by the Darwin Core standard. For example, the column catalogNumber contains data of the Darwin Core term http://rs.tdwg.org/dwc/terms/catalogNumber.
Simple download – Term definitions
The definitions marked with are from the Darwin Core standard.
The definitions marked with are from GBIF, and may reflect the result of interpretation and data quality procedures applied by GBIF, or they may not be part of Darwin Core.
| Column name | Data type | Nullable | Definition |
|---|---|---|---|
String |
No |
We aim to keep these keys stable, but this is not possible in every case. |
|
String |
No |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
This value is unaltered by GBIF’s processing; see also the GADM fields. |
|
String |
Yes |
For definitions, see the GBIF occurrence status vocabulary. |
|
Integer |
Yes |
|
|
String |
Yes |
|
|
Double |
Yes |
|
|
Double |
Yes |
|
|
Double |
Yes |
|
|
Double |
Yes |
|
|
Double |
Yes |
|
|
Double |
Yes |
|
|
Double |
Yes |
|
|
Double |
Yes |
|
|
String |
Yes |
|
|
Integer |
Yes |
||
Integer |
Yes |
||
Integer |
Yes |
|
|
Integer |
Yes |
|
|
Integer |
Yes |
|
|
String |
Yes |
See GBIF’s Darwin Core Type Vocabulary for definitions. |
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String array, delimited with |
Yes |
|
|
ISO 8601 Date |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String array, delimited with |
Yes |
|
|
String array, delimited with |
Yes |
|
|
String structure |
Yes |
Values are aligned to the GBIF EstablishmentMeans vocabulary, which is derived from the Darwin Core EstablishmentMeans vocabulary. |
|
ISO 8601 Date |
Yes |
This is the time the record was last changed in GBIF, not the time the record was last changed by the publisher. Data is also reprocessed when we changed the taxonomic backbone, geographic data sources or other interpretation procedures. An earlier interpretation system distinguished between “parsing” and “interpretation”, but in the current system there is only one process — the two dates will always be the same. |
|
String array, delimited with |
Yes |
|
|
String array, delimited with |
Yes |
See the list of occurrence issues and the OccurrenceIssue enumeration for possible values and definitions. |
DWCA downloads
Darwin Core Archive downloads from gbif.org contain the following files:
occurrence.txt-
Occurrence data after interpretation by GBIF. Described in detail below.
multimedia.txt-
Occurrence multimedia data after interpretation by GBIF. Described in detail below.
verbatim.txt-
Occurrence data without interpretation by GBIF. Described in detail below.
verbatim/*.txt-
Occurrence extension data without interpretation by GBIF. See GBIF Registered Extensions for documentation of these — note not all of them are maintained by GBIF.
dnaderiveddata.txt-
Interpreted DNA sequence data, included when the DNA Derived Data extension is requested with
interpretedExtensions. It containsgbifID,target_geneand the sanitizeddna_sequence; see DNA sequence sanitization. meta.xml-
The Darwin Core Archive metafile, describing the structure of the archive — the file formats, column names and their terms.
metadata.xml-
Metadata about the download in Ecological Metadata Language (EML).
rights.txt-
Licence information for all the datasets with occurrences in the download.
citations.txt-
Citations for all the datasets with occurrences in the download.
dataset/*.xml-
EML metadata for every dataset with occurrences in the download.
The data may be read without any special tools, including by spreadsheets such as Microsoft Excel and LibreOffice Calc (see the FAQ). The .txt files are tab-delimited, and all files are in UTF-8 encoding with Unix-style (\n) line endings.
There are libraries to read Darwin Core Archives in these programming languages:
-
Java — GBIF dwca-io
-
.NET — DwC-A_dotnet
-
Python — Python DWCA Reader
-
R — finch (NB: abandoned library)
-
Ruby — dwc-archive
Interpreted term definitions (occurrence.txt)
This is the Darwin Core Archive core entity, with row type Occurrence. Values are tab-delimited and in UTF-8 encoding.
| Column name | Data type | Nullable | Definition |
|---|---|---|---|
String |
No |
We aim to keep these keys stable, but this is not possible in every case. |
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
||
String |
Yes |
|
|
ISO 8601 Date |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String array, delimited with |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String array, delimited with |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
See GBIF’s Darwin Core Type Vocabulary for definitions. |
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String array, delimited with |
Yes |
|
|
String array, delimited with |
Yes |
|
|
Integer |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String structure |
Yes |
|
|
String structure |
Yes |
Values are aligned to the GBIF LifeStage vocabulary |
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String structure |
Yes |
Values are aligned to the GBIF EstablishmentMeans vocabulary, which is derived from the Darwin Core EstablishmentMeans vocabulary. |
|
String structure |
Yes |
Values are aligned to the GBIF DegreeOfEstablishment vocabulary, which is derived from the Darwin Core DegreeOfEstablishment vocabulary. |
|
String structure |
Yes |
Values are aligned to the GBIF Pathway vocabulary, which is derived from the Darwin Core Pathway vocabulary. |
|
String |
Yes |
|
|
String |
Yes |
For definitions, see the GBIF occurrence status vocabulary. |
|
String array, delimited with |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String array, delimited with |
Yes |
|
|
String |
Yes |
|
|
String array, delimited with |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String structure |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
Integer |
Yes |
|
|
Integer |
Yes |
||
Integer |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String array, delimited with |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String array, delimited with |
Yes |
|
|
String |
Yes |
In particular this splits the Americas into North and South America with North America including the Caribbean (except Trinidad and Tobago) and reaching down and including Panama. See the GBIF Continents for the exact divisions. This is a geographical division. See |
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
This value is unaltered by GBIF’s processing; see also the GADM fields. |
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
Double |
Yes |
|
|
Double |
Yes |
|
|
Double |
Yes |
|
|
Double |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String array, delimited with |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String structure |
Yes |
|
|
String structure |
Yes |
|
|
String structure |
Yes |
|
|
String structure |
Yes |
|
|
String structure |
Yes |
|
|
String structure |
Yes |
|
|
String structure |
Yes |
|
|
String structure |
Yes |
|
|
String structure |
Yes |
|
|
String structure |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String array, delimited with |
Yes |
|
|
String array, delimited with |
Yes |
|
|
String array, delimited with |
Yes |
|
|
ISO 8601 Date |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
No |
||
String |
Yes |
|
|
ISO 8601 Date |
Yes |
This is the time the record was last changed in GBIF, not the time the record was last changed by the publisher. Data is also reprocessed when we changed the taxonomic backbone, geographic data sources or other interpretation procedures. An earlier interpretation system distinguished between “parsing” and “interpretation”, but in the current system there is only one process — the two dates will always be the same. |
|
Double |
Yes |
|
|
Double |
Yes |
|
|
Double |
Yes |
|
|
Double |
Yes |
|
|
Double |
Yes |
|
|
String array, delimited with |
Yes |
See the list of occurrence issues and the OccurrenceIssue enumeration for possible values and definitions. |
|
String array, delimited with |
Yes |
– |
|
String array, delimited with |
Yes |
– |
|
String array, delimited with |
Yes |
|
|
Boolean |
Yes |
|
|
Boolean |
Yes |
|
|
Integer |
Yes |
|
|
Integer |
Yes |
|
|
Integer |
Yes |
|
|
Integer |
Yes |
|
|
Integer |
Yes |
|
|
Integer |
Yes |
|
|
String |
Yes |
– |
|
Integer |
Yes |
|
|
String |
Yes |
– |
|
String |
Yes |
– |
|
String |
Yes |
– |
|
Integer |
Yes |
|
|
Integer |
Yes |
|
|
Integer |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
This is not yet a Darwin Core term, see the proposal to add it. |
|
String |
Yes |
|
|
ISO 8601 Date |
Yes |
This is the time the record was last changed in GBIF, not the time the record was last changed by the publisher. Data is also reprocessed when we changed the taxonomic backbone, geographic data sources or other interpretation procedures. An earlier interpretation system distinguished between “parsing” and “interpretation”, but in the current system there is only one process — the two dates will always be the same. |
|
ISO 8601 Date |
Yes |
|
|
String |
Yes |
– |
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
This is a political division, part of GBIF’s governance structure. |
|
String |
Yes |
This is a political division, part of GBIF’s governance structure. |
|
String array, delimited with |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
See the GBIF vocabulary for the values and their definitions, and the IUCN Red List of Threatened Species dataset in GBIF for the version of the Red List GBIF’s interpretation procedures are using. |
Multimedia term definitions (multimedia.txt)
| Column name | Data type | Nullable | Definition |
|---|---|---|---|
String |
No |
We aim to keep these keys stable, but this is not possible in every case. |
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
Verbatim term definitions (verbatim.txt)
Data in this table is not modified by GBIF interpretation processes, except for conversion to Unicode and possible changes to whitespace (spaces, tabs, newlines etc).
Verbatim extensions (verbatim/*.txt)
Data in these tables is not modified by GBIF interpretation processes, except for conversion to Unicode and possible changes to whitespace (spaces, tabs, newlines etc).
See the GBIF Registered Extensions for documentation of the extensions.
Species list downloads – Term definitions
Species list downloads are a summary format containing the distinct list of species names returned by the filter.
The definitions marked with are from GBIF, and may reflect the result of interpretation and data quality procedures applied by GBIF, or they may not be part of Darwin Core.
| Column name | Data type | Nullable | Definition |
|---|---|---|---|
Integer |
No |
|
|
String |
Yes |
|
|
Integer |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
||
String |
Yes |
|
|
String |
Yes |
|
|
String |
Yes |
|
|
Integer |
Yes |
|
|
String |
Yes |
|
|
Integer |
Yes |
|
|
String |
Yes |
|
|
Integer |
Yes |
|
|
String |
Yes |
|
|
Integer |
Yes |
|
|
String |
Yes |
|
|
Integer |
Yes |
|
|
String |
Yes |
|
|
Integer |
Yes |
|
|
String |
Yes |
|
|
Integer |
Yes |
|
|
String |
Yes |
See the GBIF vocabulary for the values and their definitions, and the IUCN Red List of Threatened Species dataset in GBIF for the version of the Red List GBIF’s interpretation procedures are using. |
FASTA Archive downloads
The (API: FASTA_ARCHIVE) format is a Darwin Core Archive download of occurrences carrying DNA sequences, with two extra files added: sequences.fasta, holding the sanitized sequences in FASTA format, and sequences.txt, holding the quality metrics calculated for each of those sequences.
The format is intended for occurrence searches restricted to records with sequences, typically together with a filter on target gene (the nucleotideSequence.targetGene search parameter). Whatever filter you request, the download itself only ever includes occurrences with at least one sanitized sequence — records whose sequences are all missing or invalid are excluded from the whole archive, occurrence.txt and verbatim.txt included.
|
| This is version 1 of the format. The layout of the archive and the contents of the two sequence files may change in future versions. |
A FASTA Archive download contains the files of a Darwin Core Archive download, plus the two sequence files:
├── occurrence.txt (1)
├── verbatim.txt (1)
├── multimedia.txt (1)
├── verbatim/dnaderiveddata.txt (2)
├── sequences.fasta (3)
├── sequences.txt (3)
├── dataset/*.xml (1)
├── meta.xml (1)
├── metadata.xml (1)
├── citations.txt (4)
└── rights.txt (1)
| 1 | As in any Darwin Core Archive download, see above. |
| 2 | The verbatim DNA Derived Data extension, present when it was requested with verbatimExtensions. This holds the sequences and MIxS terms as published, before sanitization. |
| 3 | Specific to this format, described below. |
| 4 | Contains a citation for the download as a whole, described below. |
The two sequence files are not described in meta.xml: standard Darwin Core Archive readers will see the archive as an ordinary occurrence download and ignore them.
|
Sequences (sequences.fasta)
One FASTA record per sanitized sequence, with the sequence on a single line — it is not wrapped. The header line has three parts, separated by |:
>nucleotideSequenceID|gbifID|target_gene
For example:
>e4346dfb13ca73687e6ca421e1608d54|5866098319|COI
CTTATCAAGCATTACGGCCCATTCTGGGCCTGCAGTAGATTTGGCAATTTTTAGTCTACATATAGCAGGTGCG…
>d3970a214c7ee0c97542bf907d4a148a|5866098312|SSU_rRNA_12S_mitochondrial
CCATCGTGATAAATTCTTAGGTCATTACCAGTGCCAAATCTTGCTTCTACACCATCTACAAAATCTATGTTAC…
nucleotideSequenceID-
The MD5 hash of the sanitized sequence. Identical sequences share the same identifier, so this can be used to group or deduplicate records across datasets.
gbifID-
The GBIF identifier of the occurrence the sequence belongs to — the key to
occurrence.txt,verbatim.txtandsequences.txt. target_gene-
The interpreted target gene, normalised against the GBIF target gene vocabulary. Empty when the publisher supplied no value, or a value that could not be interpreted.
The sequence written to the file is the sanitized sequence, not the sequence as published: whitespace and gaps are removed, RNA is converted to DNA, long runs of N are capped, and the ends are trimmed. See DNA sequence sanitization for the full pipeline. Sequences flagged as invalid by the sanitization — those containing characters that are not IUPAC codes, or natural-language text — are not included in this file. The sequences as published are available in verbatim/dnaderiveddata.txt.
An occurrence may have more than one sequence, for instance for different target genes. Each of those sequences is a separate record in this file, and all sequences of an included occurrence are written, including sequences for genes other than the one filtered on.
Sequence metrics (sequences.txt)
A tab-delimited file with one row per record in sequences.fasta, in the same order. It repeats the identifiers from the FASTA header and adds the quality metrics produced by the sanitization pipeline, so that you can apply your own thresholds — on sequence length, ambiguity or GC content — without re-processing the sequences.
| Column | Type | Description |
|---|---|---|
|
long |
The GBIF identifier of the occurrence; the key to |
|
string |
MD5 hash of the sanitized sequence; the key to |
|
string |
Interpreted target gene, from the GBIF target gene vocabulary. |
|
integer |
Length of the sanitized sequence. |
|
float |
GC content, calculated over A/C/G/T bases only. |
|
float |
Fraction of characters that are not valid IUPAC DNA codes. |
|
float |
Fraction of characters that are ambiguous IUPAC codes (anything other than A, C, G, T or N). |
|
float |
Fraction of characters that are |
|
integer |
Number of runs of |
|
boolean |
Whether natural-language marker words were found in the published sequence. |
|
boolean |
Whether either end of the sequence was trimmed. |
|
boolean |
Whether any whitespace or gap characters were removed. |
|
boolean |
Whether the sequence was flagged as invalid. Invalid sequences are not part of this download, so this is always |
For example:
gbifID nucleotideSequenceID target_gene sequenceLength gcContent … invalid
5866098319 e4346dfb13ca73687e6ca421e1608d54 COI 313 0.31309904153354634 … false
5866098312 d3970a214c7ee0c97542bf907d4a148a SSU_rRNA_12S_mitochondrial 80 0.3875 … false
Citation (citations.txt)
As with other downloads, citations.txt lists a citation for every dataset contributing occurrences to the download. For FASTA Archive downloads it starts with an additional citation for the download as a whole, describing what the archive contains:
FASTA Archive Download of COI sequences from Insecta in Norway. GBIF.org (16 March 2026). https://doi.org/10.15468/dl.xxxxx
The target gene, taxonomic and geographic parts of the sentence are taken from the filters of the download request, and each is left out when the request had no such filter. Use this citation, together with the DOI of the download, when citing the archive; the dataset citations remain relevant if you use records from individual datasets.
Cube downloads
This download format allows you to aggregate occurrences by their taxonomic, temporal and/or spatial properties. For example, a data cube can be configured to aggregate occurrences by family, month and grid cell of the European Environment Agency reference grid (three dimensions) and count the number of occurrences (a measure) per combination. The result is a CSV file.
Once configured, a SQL query will be created to generate the data cube. For more advanced use, it is possible to further customize the requested download by editing the SQL query.
Read more about occurrence cubes.