Class: FileCollection
A collection of files with shared characteristics (format, purpose, structure). Represents a logical grouping of related files within a dataset, such as all training data files, all image files, or all raw data files. Maps to RO-Crate Dataset entities via schema:hasPart relationships.
URI: dcat:Dataset
classDiagram
class FileCollection
click FileCollection href "../FileCollection/"
Information <|-- FileCollection
click Information href "../Information/"
FileCollection : collection_type
FileCollection --> "0..1" FileCollectionTypeEnum : collection_type
click FileCollectionTypeEnum href "../FileCollectionTypeEnum/"
FileCollection : compression
FileCollection --> "0..1" CompressionEnum : compression
click CompressionEnum href "../CompressionEnum/"
FileCollection : conforms_to
FileCollection : conforms_to_class
FileCollection : conforms_to_schema
FileCollection : conforms_to_standard
FileCollection --> "*" DataStandardEnum : conforms_to_standard
click DataStandardEnum href "../DataStandardEnum/"
FileCollection : created_by
FileCollection : created_on
FileCollection : description
FileCollection : doi
FileCollection : download_url
FileCollection : external_resources
FileCollection --> "*" ExternalResource : external_resources
click ExternalResource href "../ExternalResource/"
FileCollection : file_count
FileCollection : id
FileCollection : issued
FileCollection : keywords
FileCollection : language
FileCollection : last_updated_on
FileCollection : license
FileCollection : modified_by
FileCollection : name
FileCollection : notes
FileCollection : page
FileCollection : path
FileCollection : publisher
FileCollection : resources
FileCollection --> "*" File : resources
click File href "../File/"
FileCollection : source_caveats
FileCollection : status
FileCollection : title
FileCollection : total_bytes
FileCollection : version
FileCollection : was_derived_from
Inheritance
- NamedThing
- Information
- FileCollection
- Information
Slots
| Name | Cardinality and Range | Description | Inheritance |
|---|---|---|---|
| path | 0..1 String |
Path or URL to the FileCollection | direct |
| compression | 0..1 CompressionEnum |
Compression format if the collection is packaged as a compressed archive (e | direct |
| external_resources | * ExternalResource |
External files or URLs referenced by this file collection | direct |
| resources | * File |
The individual files in this collection | direct |
| collection_type | 0..1 FileCollectionTypeEnum |
Type(s) of content in this file collection | direct |
| file_count | 0..1 Integer |
Number of files in this collection | direct |
| total_bytes | 0..1 Integer |
Total size of all files in this collection, in bytes (integer) | direct |
| conforms_to | 0..1 String |
A standard or specification the dataset's own content follows — an imagin... | Information |
| conforms_to_class | 0..1 String |
The class within conforms_to_schema that this record instantiates — `Datase... |
Information |
| conforms_to_schema | 0..1 String |
The metadata schema this record itself is written in — normally `https://... | Information |
| conforms_to_standard | * DataStandardEnum |
Which registered data standard conforms_to names, where one applies — a ter... |
Information |
| created_by | 0..1 String |
The person or organization primarily responsible for creating the resource | Information |
| created_on | 0..1 Datetime |
The date and time when the resource was created | Information |
| doi | 0..1 String |
Digital Object Identifier (DOI) in format 10 | Information |
| download_url | 0..1 Uri |
URL from which the data can be downloaded | Information |
| issued | 0..1 Datetime |
Date of formal issuance or publication of the resource | Information |
| keywords | * String |
Keywords or tags describing the resource for discovery and classification | Information |
| language | 0..1 String |
Language in which the information is expressed | Information |
| last_updated_on | 0..1 Datetime |
The date and time when the resource was most recently modified or updated | Information |
| license | 0..1 String |
The legal license under which the resource is made available (e | Information |
| modified_by | 0..1 String |
A person or organization that contributed to modifying or updating the resour... | Information |
| page | 0..1 String |
A landing page or web page providing access to or information about the resou... | Information |
| publisher | 0..1 Uriorcurie |
The organization or entity responsible for making the resource available | Information |
| status | 0..1 String |
The status of the resource (e | Information |
| title | 0..1 String |
The official title of the element | Information |
| version | 0..1 String |
The version identifier of the resource (e | Information |
| was_derived_from | 0..1 String |
A resource from which this resource was derived, in whole or in part | Information |
| id | 1 Uriorcurie |
A unique identifier for a thing | NamedThing |
| name | 0..1 String |
A human-readable name for a thing | NamedThing |
| description | 0..1 String |
A human-readable description for a thing | NamedThing |
| notes | 0..1 String |
Residual content only, after every fitting slot is used — structured slots fi... | NamedThing |
| source_caveats | 0..1 String |
Commentary on the evidence behind this object's values — source conflicts, wh... | NamedThing |
Usages
| used by | used in | type | used |
|---|---|---|---|
| Dataset | file_collections | range | FileCollection |
| DataSubset | file_collections | range | FileCollection |
Aliases
- file collection
- data files
- file group
Identifier and Mapping Information
Schema Source
- from schema: https://w3id.org/bridge2ai/data-sheets-schema
Mappings
| Mapping Type | Mapped Value |
|---|---|
| self | dcat:Dataset |
| native | data_sheets_schema:FileCollection |
| exact | schema:Dataset |
| close | dcat:Distribution |
LinkML Source
Direct
name: FileCollection
description: A collection of files with shared characteristics (format, purpose, structure).
Represents a logical grouping of related files within a dataset, such as all training
data files, all image files, or all raw data files. Maps to RO-Crate Dataset entities
via schema:hasPart relationships.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
aliases:
- file collection
- data files
- file group
exact_mappings:
- schema:Dataset
close_mappings:
- dcat:Distribution
is_a: Information
slots:
- path
- compression
- external_resources
- resources
slot_usage:
path:
name: path
description: Path or URL to the FileCollection. May be a directory path, archive
file path, or download URL depending on how the collection is distributed.
compression:
name: compression
description: Compression format if the collection is packaged as a compressed
archive (e.g., gzip, zip, bzip2). Omit this field for uncompressed collections
or purely logical groupings.
external_resources:
name: external_resources
description: External files or URLs referenced by this file collection.
range: ExternalResource
multivalued: true
inlined_as_list: true
resources:
name: resources
description: The individual files in this collection. To describe a group of files
as a unit rather than listing them, add another entry to the dataset's `file_collections`
— a collection does not nest inside another collection.
range: File
multivalued: true
inlined_as_list: true
attributes:
collection_type:
name: collection_type
description: Type(s) of content in this file collection. A collection may have
multiple types, for example a collection containing both raw_data and documentation
files would have both types listed.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema/file-collection
rank: 1000
slot_uri: d4d:collectionType
domain_of:
- FileCollection
range: FileCollectionTypeEnum
file_count:
name: file_count
annotations:
d4d:docExample:
tag: d4d:docExample
value: '47'
description: Number of files in this collection.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema/file-collection
rank: 1000
slot_uri: d4d:fileCount
domain_of:
- FileCollection
range: integer
total_bytes:
name: total_bytes
annotations:
d4d:docExample:
tag: d4d:docExample
value: 1073741824 (1 GiB = 1024³ bytes)
description: Total size of all files in this collection, in bytes (integer). Maps
to dcat:byteSize.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema/file-collection
exact_mappings:
- dcat:byteSize
rank: 1000
slot_uri: d4d:total_bytes
domain_of:
- FileCollection
range: integer
class_uri: dcat:Dataset
Induced
name: FileCollection
description: A collection of files with shared characteristics (format, purpose, structure).
Represents a logical grouping of related files within a dataset, such as all training
data files, all image files, or all raw data files. Maps to RO-Crate Dataset entities
via schema:hasPart relationships.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
aliases:
- file collection
- data files
- file group
exact_mappings:
- schema:Dataset
close_mappings:
- dcat:Distribution
is_a: Information
slot_usage:
path:
name: path
description: Path or URL to the FileCollection. May be a directory path, archive
file path, or download URL depending on how the collection is distributed.
compression:
name: compression
description: Compression format if the collection is packaged as a compressed
archive (e.g., gzip, zip, bzip2). Omit this field for uncompressed collections
or purely logical groupings.
external_resources:
name: external_resources
description: External files or URLs referenced by this file collection.
range: ExternalResource
multivalued: true
inlined_as_list: true
resources:
name: resources
description: The individual files in this collection. To describe a group of files
as a unit rather than listing them, add another entry to the dataset's `file_collections`
— a collection does not nest inside another collection.
range: File
multivalued: true
inlined_as_list: true
attributes:
collection_type:
name: collection_type
description: Type(s) of content in this file collection. A collection may have
multiple types, for example a collection containing both raw_data and documentation
files would have both types listed.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema/file-collection
rank: 1000
slot_uri: d4d:collectionType
alias: collection_type
owner: FileCollection
domain_of:
- FileCollection
range: FileCollectionTypeEnum
file_count:
name: file_count
annotations:
d4d:docExample:
tag: d4d:docExample
value: '47'
description: Number of files in this collection.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema/file-collection
rank: 1000
slot_uri: d4d:fileCount
alias: file_count
owner: FileCollection
domain_of:
- FileCollection
range: integer
total_bytes:
name: total_bytes
annotations:
d4d:docExample:
tag: d4d:docExample
value: 1073741824 (1 GiB = 1024³ bytes)
description: Total size of all files in this collection, in bytes (integer). Maps
to dcat:byteSize.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema/file-collection
exact_mappings:
- dcat:byteSize
rank: 1000
slot_uri: d4d:total_bytes
alias: total_bytes
owner: FileCollection
domain_of:
- FileCollection
range: integer
path:
name: path
annotations:
d4d:docExample:
tag: d4d:docExample
value: data/example_cohort/participants.csv
description: Path or URL to the FileCollection. May be a directory path, archive
file path, or download URL depending on how the collection is distributed.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: schema:contentUrl
alias: path
owner: FileCollection
domain_of:
- File
- FileCollection
range: string
compression:
name: compression
annotations:
d4d:docExample:
tag: d4d:docExample
value: zip
description: Compression format if the collection is packaged as a compressed
archive (e.g., gzip, zip, bzip2). Omit this field for uncompressed collections
or purely logical groupings.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcat:compressFormat
alias: compression
owner: FileCollection
domain_of:
- Information
- File
- FileCollection
range: CompressionEnum
external_resources:
name: external_resources
description: External files or URLs referenced by this file collection.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcterms:references
alias: external_resources
owner: FileCollection
domain_of:
- Dataset
- ExternalResource
- FileCollection
range: ExternalResource
multivalued: true
inlined_as_list: true
resources:
name: resources
description: The individual files in this collection. To describe a group of files
as a unit rather than listing them, add another entry to the dataset's `file_collections`
— a collection does not nest inside another collection.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: schema:hasPart
alias: resources
owner: FileCollection
domain_of:
- DatasetCollection
- Dataset
- FileCollection
range: File
multivalued: true
inlined_as_list: true
conforms_to:
name: conforms_to
annotations:
d4d:docExample:
tag: d4d:docExample
value: Brain Imaging Data Structure (BIDS) v1.9.0
description: 'A standard or specification the **dataset''s own content** follows
— an imaging, waveform, file-layout or common-data-model standard such as DICOM,
WFDB, BIDS, OMOP CDM, Open mHealth or ESDS. Not the schema this datasheet is
written in: that is `conforms_to_schema`.'
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcterms:conformsTo
alias: conforms_to
owner: FileCollection
domain_of:
- Information
range: string
conforms_to_class:
name: conforms_to_class
annotations:
d4d:docExample:
tag: d4d:docExample
value: Dataset
d4d:perRecord:
tag: d4d:perRecord
value: true
description: The class within `conforms_to_schema` that this record instantiates
— `Dataset` for a full datasheet, `CoreDataset` for a core one. Like `conforms_to_schema`,
a statement about the record rather than the data.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
broad_mappings:
- dcterms:conformsTo
rank: 1000
slot_uri: d4d:conformsToClass
alias: conforms_to_class
owner: FileCollection
domain_of:
- Information
range: string
conforms_to_schema:
name: conforms_to_schema
annotations:
d4d:docExample:
tag: d4d:docExample
value: https://w3id.org/bridge2ai/data-sheets-schema
d4d:perRecord:
tag: d4d:perRecord
value: true
description: The metadata schema **this record itself** is written in — normally
`https://w3id.org/bridge2ai/data-sheets-schema`. A statement about the datasheet,
not about the data it describes; a standard the *data* follows belongs in `conforms_to`.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
broad_mappings:
- dcterms:conformsTo
rank: 1000
slot_uri: d4d:conformsToSchema
alias: conforms_to_schema
owner: FileCollection
domain_of:
- Information
range: string
conforms_to_standard:
name: conforms_to_standard
annotations:
d4d:docExample:
tag: d4d:docExample
value: BIDS
description: 'Which registered data standard `conforms_to` names, where one applies
— a term rather than prose, so the corpus can be queried for "what standards
does this follow?" (#403).
`conforms_to` records what the sources say, in their words. This records which
standard that is. Populate both for the same standard: one is verbatim evidence,
the other is the normalized term, and they answer different questions.
Use `OTHER` when the sources name a standard this vocabulary lacks, and keep
the name in `conforms_to` so nothing is lost. Never pick a near-neighbour —
reporting BIDS because a dataset is neuroimaging, when the sources say otherwise,
is an invention.'
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcterms:conformsTo
alias: conforms_to_standard
owner: FileCollection
domain_of:
- Information
range: DataStandardEnum
multivalued: true
created_by:
name: created_by
annotations:
d4d:docExample:
tag: d4d:docExample
value: orcid:0000-0002-1234-5678
description: The person or organization primarily responsible for creating the
resource.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcterms:creator
alias: created_by
owner: FileCollection
domain_of:
- Information
range: string
created_on:
name: created_on
annotations:
d4d:docExample:
tag: d4d:docExample
value: '2021-03-01T00:00:00Z'
description: The date and time when the resource was created. Requires a UTC offset
(`…T00:00:00Z`); RFC 3339 `date-time` rejects a naive timestamp.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcterms:created
alias: created_on
owner: FileCollection
domain_of:
- Information
range: datetime
doi:
name: doi
annotations:
d4d:docExample:
tag: d4d:docExample
value: 10.xxxxx/example.1234
description: 'Digital Object Identifier (DOI) in format 10.xxxx/xxxxx — a registrant
prefix, a slash, a suffix — providing persistent identification. The bare DOI
only — not the doi: CURIE and not the doi.org resolver URL. The pattern is anchored
(#646): unanchored, JSON Schema''s search semantics accepted all three surface
forms, and one v5 record carried the same DOI three ways. The 2026-08-20b arm
already writes the bare form; records written before the anchor keep the validation
verdicts they were pinned with (#426) and are not rewritten (#520).'
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
exact_mappings:
- schema:identifier
broad_mappings:
- dcterms:identifier
rank: 1000
slot_uri: d4d:doiIdentifier
alias: doi
owner: FileCollection
domain_of:
- Information
range: string
pattern: ^10\.\d{4,}\/.+$
download_url:
name: download_url
annotations:
d4d:docExample:
tag: d4d:docExample
value: https://example.org/datasets/example-cohort/download
description: URL from which the data can be downloaded. This is not the same as
the landing page, which is a page that describes the dataset. Rather, this URL
points directly to the data itself.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
exact_mappings:
- schema:url
rank: 1000
slot_uri: dcat:downloadURL
alias: download_url
owner: FileCollection
domain_of:
- Information
- DistributionFormat
range: uri
issued:
name: issued
annotations:
d4d:docExample:
tag: d4d:docExample
value: '2022-06-30T00:00:00Z'
description: Date of formal issuance or publication of the resource. A `datetime`
renders as RFC 3339 `date-time`, which **requires a UTC offset** — write `2024-11-15T00:00:00Z`,
not `2024-11-15T00:00:00`. A date alone is also rejected.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcterms:issued
alias: issued
owner: FileCollection
domain_of:
- Information
range: datetime
keywords:
name: keywords
annotations:
d4d:docExample:
tag: d4d:docExample
value: chronic disease, medical imaging, multimodal, clinical data
description: Keywords or tags describing the resource for discovery and classification.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcat:keyword
alias: keywords
owner: FileCollection
domain_of:
- Information
range: string
multivalued: true
language:
name: language
annotations:
d4d:docExample:
tag: d4d:docExample
value: en
description: Language in which the information is expressed.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
exact_mappings:
- schema:inLanguage
rank: 1000
slot_uri: dcterms:language
alias: language
owner: FileCollection
domain_of:
- Information
range: string
last_updated_on:
name: last_updated_on
annotations:
d4d:docExample:
tag: d4d:docExample
value: '2022-06-30T00:00:00Z'
description: The date and time when the resource was most recently modified or
updated. Requires a UTC offset (`…T00:00:00Z`); RFC 3339 `date-time` rejects
a naive timestamp.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcterms:modified
alias: last_updated_on
owner: FileCollection
domain_of:
- Information
range: datetime
license:
name: license
annotations:
d4d:docExample:
tag: d4d:docExample
value: CC-BY-NC-4.0
description: The legal license under which the resource is made available (e.g.,
"MIT", "CC-BY-4.0").
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcterms:license
alias: license
owner: FileCollection
domain_of:
- Software
- Information
range: string
modified_by:
name: modified_by
annotations:
d4d:docExample:
tag: d4d:docExample
value: orcid:0000-0002-9876-5432
description: A person or organization that contributed to modifying or updating
the resource.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcterms:contributor
alias: modified_by
owner: FileCollection
domain_of:
- Information
range: string
page:
name: page
annotations:
d4d:docExample:
tag: d4d:docExample
value: https://example.org/datasets/example-cohort
description: A landing page or web page providing access to or information about
the resource.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcat:landingPage
alias: page
owner: FileCollection
domain_of:
- Information
range: string
publisher:
name: publisher
annotations:
d4d:docExample:
tag: d4d:docExample
value: 'ROR:0xxxxxxxx # a ROR CURIE (form only), a DOI, or a URL — not a
plain name'
description: The organization or entity responsible for making the resource available.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcterms:publisher
alias: publisher
owner: FileCollection
domain_of:
- Information
range: uriorcurie
status:
name: status
annotations:
d4d:docExample:
tag: d4d:docExample
value: published
description: The status of the resource (e.g., draft, published, deprecated).
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: d4d:publicationStatus
alias: status
owner: FileCollection
domain_of:
- Information
range: string
title:
name: title
annotations:
d4d:docExample:
tag: d4d:docExample
value: 'Example Cohort Study of Chronic Disease: Data Release 1.0'
description: The official title of the element.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: dcterms:title
alias: title
owner: FileCollection
domain_of:
- Information
range: string
version:
name: version
annotations:
d4d:docExample:
tag: d4d:docExample
value: 2.0.0
description: The version identifier of the resource (e.g., "1.0", "2.3.1").
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
rank: 1000
slot_uri: schema:version
alias: version
owner: FileCollection
domain_of:
- Software
- Information
range: string
was_derived_from:
name: was_derived_from
annotations:
d4d:docExample:
tag: d4d:docExample
value: https://example.org/datasets/example-cohort/versions/1
description: A resource from which this resource was derived, in whole or in part.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema
exact_mappings:
- dcterms:source
rank: 1000
slot_uri: prov:wasDerivedFrom
alias: was_derived_from
owner: FileCollection
domain_of:
- Information
range: string
id:
name: id
annotations:
d4d:docExample:
tag: d4d:docExample
value: https://example.org/dataset/my-dataset-001
description: A unique identifier for a thing.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema/base
rank: 1000
slot_uri: schema:identifier
identifier: true
alias: id
owner: FileCollection
domain_of:
- NamedThing
- Organization
- DatasetProperty
- Grant
range: uriorcurie
required: true
name:
name: name
annotations:
d4d:docExample:
tag: d4d:docExample
value: Example Cohort Dataset
description: A human-readable name for a thing.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema/base
rank: 1000
slot_uri: schema:name
alias: name
owner: FileCollection
domain_of:
- NamedThing
- Organization
- DatasetProperty
- Grant
range: string
description:
name: description
annotations:
d4d:docExample:
tag: d4d:docExample
value: A multimodal dataset of 2,500 participants in a longitudinal cohort
study.
description: A human-readable description for a thing.
from_schema: https://w3id.org/bridge2ai/data-sheets-schema/base
rank: 1000
slot_uri: schema:description
alias: description
owner: FileCollection
domain_of:
- NamedThing
- Organization
- DatasetProperty
- Grant
- DatasetRelationship
range: string
notes:
name: notes
description: Residual content only, after every fitting slot is used — structured
slots first, then description; notes solely for what description cannot hold
(#385). Never invent keys (#380).
from_schema: https://w3id.org/bridge2ai/data-sheets-schema/base
rank: 1000
slot_uri: schema:comment
alias: notes
owner: FileCollection
domain_of:
- NamedThing
- Organization
- DatasetProperty
- Grant
range: string
source_caveats:
name: source_caveats
description: 'Commentary on the evidence behind this object''s values — source
conflicts, what a value was transcribed from, questions the source material
leaves unanswered (#385). A trust annotation about the sibling slots, not dataset
content: content belongs in the structured slots, description or notes.'
from_schema: https://w3id.org/bridge2ai/data-sheets-schema/base
rank: 1000
slot_uri: dcterms:provenance
alias: source_caveats
owner: FileCollection
domain_of:
- NamedThing
- Organization
- DatasetProperty
- Grant
range: string
class_uri: dcat:Dataset