Skip to content

DatasetOperations

Access via client.datasets.

File operations are also available via client.files (by MFID) and ingestion via client.ingestions.

crucible.resources.datasets.DatasetOperations

Dataset-related API operations.

Access via: client.datasets.get(), client.datasets.list(), etc.

get(dataset_mfid, include_metadata=False, include_links=False, include_owner=True)

Get a dataset by its canonical MFID.

Parameters:

Name Type Description Default
dataset_mfid str

Dataset MFID

required
include_metadata bool

Whether to include scientific metadata

False
include_links bool

Whether to include immediate parent/child/associated links

False
include_owner bool

Resolve owner_orcid into a public-safe user object (default: True)

True

Returns:

Name Type Description
Dict Dict

Dataset object with optional metadata and links

list(sample_mfid=None, include_metadata=False, include_links=False, include_owner=False, limit=DEFAULT_LIMIT, offset=0, accessible_to_user=None, accessible_to_project=None, instrument_mfid=None, project_id=None, project_mfid=None, project_scope=None, **kwargs)

List datasets with optional filtering and automatic pagination.

Parameters:

Name Type Description Default
sample_mfid str

If provided, returns datasets for this sample

None
limit int

Maximum total results to return (default: 100). Larger requests are handled transparently by following the server's keyset cursor. Pass None to fetch all matches.

DEFAULT_LIMIT
offset int

Deprecated. The /datasets collection uses keyset pagination and ignores offset, including when filtered by sample_mfid.

0
include_metadata bool

Include scientific metadata in results

False
include_links bool

Include linked resources (parents, children, associated) per dataset

False
include_owner bool

Resolve owner_orcid into a public-safe user object per dataset

False
accessible_to_user Optional[Union[str, Sequence[str]]]

User reference or references whose effective access must include every result

None
accessible_to_project Optional[Union[str, Sequence[str]]]

Project reference or references whose direct access must include every result

None
instrument_mfid Optional[str]

Canonical instrument MFID assigned to every result

None
project_id Optional[str]

Project slug used to scope results by project relationship

None
project_mfid Optional[str]

Canonical project MFID used to scope results by project relationship

None
project_scope Optional[str]

Project relationship to include: assigned, shared, or all. Requires project_id or project_mfid and defaults to assigned.

None
**kwargs Any

Query parameters for filtering. Supported fields include: keyword, unique_id, public, dataset_name, owner_orcid, project_id, instrument_name, timestamp, size, data_format, data_type, measurement, session_name. Filters expect exact matches (case sensitive) except for keywords, which are case insensitive and match substrings.

{}

Returns:

Type Description
List[Dict]

List[Dict]: Dataset objects matching filter criteria

count(project_id=None, project_mfid=None, project_scope=None, **kwargs)

Return the total number of datasets matching the given filters without fetching items.

search(q, project_id=None, limit=20)

Fuzzy name search across datasets. Available to all authenticated users.

Matches against dataset_name. Returns datasets the caller can read. For scientific metadata search use search_metadata().

Parameters:

Name Type Description Default
q str

Search term (min 3 chars). Typo-tolerant — "pero" finds "perovskite".

required
project_id Optional[str]

Optional project to scope results to.

None
limit int

Max results (default 20, max 50).

20

Returns:

Type Description
List[Dict]

List[Dict]: Matching DatasetResponse records, ranked by relevance.

create(dataset, scientific_metadata=None, keywords=None, files=None, upload_files=True, files_to_upload=None, ingestor=None, verbose=False, wait_for_ingestion_response=False, skip_ingestion=False)

Create a new dataset record with scientific metadata and keywords.

Parameters:

Name Type Description Default
dataset Dataset

Dataset object with dataset details. Use owner with an ORCID, MFID, username, or email to create for a specific owner. owner_orcid is deprecated for creation. Providing both owner fields is invalid. project_id and instrument_id accept human-readable IDs; project_mfid and instrument_mfid accept canonical MFIDs. Matching ID and MFID selectors may be supplied together.

required
scientific_metadata dict

Scientific metadata

None
keywords list

Keywords to associate with dataset

None
files list

Files to attach. Each item is either a local path (str) or an AssociatedFile describing a file that lives elsewhere (Globus, NERSC, a shared filesystem, etc.). - A str path is uploaded to GCS when upload_files=True (default), or cataloged by its resolved absolute path (storage_backend='local', no upload) when upload_files=False. - An AssociatedFile is always cataloged via add_remote_file() - it must set storage_backend to something other than 'gcs', since there's no local file to upload from a model description.

None
upload_files bool

Whether str paths in files are uploaded to GCS (default: True) or just cataloged by their local path (False). Only affects str items - AssociatedFile items are routed by their own storage_backend regardless of this flag.

True
files_to_upload list

Deprecated alias for files (str paths only, always uploaded). Use files instead.

None

Returns: Dict: created_record, scientific_metadata_record, dataset_mfid, dsid, files (files is the per-item result of adding each entry in files, in the same order). dsid is retained for compatibility.

update(dataset_mfid, **updates)

Update an existing dataset with new field values.

'owner_orcid' and 'project_id' are no longer accepted here (422) - use transfer_ownership() / reassign_project() instead. The deprecated 'public' field delegates to set_public() or set_private(). Instrument reassignment is not available through generic PATCH. Omit 'instrument_id' and 'instrument_name' unless resubmitting their current values for compatibility.

Parameters:

Name Type Description Default
dataset_mfid str

Dataset MFID

required
**updates Any

Fields to update (e.g., dataset_name="New Name")

{}

Returns:

Name Type Description
Dict Dict

Updated dataset object

Example

client.datasets.update("my-dataset-id", dataset_name="Updated Name")

add_file(dataset_mfid, file_path, ingestion_class=None, wait_for_ingestion_response=False, multipart=True, chunk_size_mb=None, max_workers=None, skip_ingestion=False)

Upload a file to a dataset and request ingestion.

Parameters:

Name Type Description Default
dataset_mfid str

Dataset MFID

required
file_path str

Local path to the file

required
ingestion_class Optional[str]

Ingestion class for the worker (e.g. 'lammps', 'nexus'). Defaults to the server-side default if omitted.

None
wait_for_ingestion_response bool

Block until ingestion completes. Ignored if skip_ingestion is True.

False
multipart bool

Use parallel multipart upload (default: True). Set to False to use the sequential resumable upload (slower but simpler).

True
chunk_size_mb Optional[int]

Override chunk size in MiB (uses config/default if None).

None
max_workers Optional[int]

Override number of upload threads (uses config/default if None).

None
skip_ingestion bool

Upload the file without requesting ingestion. Request it later with client.files.request_ingestion(mfid).

False

Returns:

Name Type Description
Dict Dict

{'associated_file': AssociatedFileRead, 'ingestion_request': IngestionRequest}

add_remote_file(dataset_mfid, file)

Register a file that lives outside GCS (Globus, NERSC, a shared filesystem path, etc.) without uploading it.

Crucible only catalogs the pointer here — it never verifies the file exists or fetches bytes on your behalf. Internally this is a two-step API call (create, then set storage_path) hidden behind one method.

Parameters:

Name Type Description Default
dataset_mfid str

Dataset MFID

required
file AssociatedFile

Must have storage_backend set to something other than 'gcs' (e.g. 'globus', 'local'). storage_path is optional — omit it to catalog the file now and set its location later via client.files.update(mfid, storage_path=...).

required

Returns:

Name Type Description
Dict Dict

The created (and, if storage_path was given, updated) file record.

list_files(dataset_mfid)

List files attached to a dataset.

Parameters:

Name Type Description Default
dataset_mfid str

Dataset MFID

required

Returns:

Type Description
List[Dict]

List[Dict]: File records (mfid, filename, storage_path, storage_backend, access_note, size, sha256_hash, dataset_mfid). For a 'gcs' file, storage_path is null until it has been ingested.

Get signed download URLs for all ingested files in a dataset.

Returns:

Name Type Description
Dict Dict

Mapping of file MFID → signed URL. Empty dict if no ingested files.

download(dataset_mfid, file_name=None, output_dir='crucible-downloads', no_files=False, no_record=False, overwrite_existing=True, include=None, exclude=None)

Download a dataset's files and optionally save its record as JSON.

Parameters:

Name Type Description Default
dataset_mfid str

Dataset MFID

required
file_name Optional[str]

Deprecated. Use include=['pattern'] with glob syntax.

None
output_dir str

Directory to save files (default: 'crucible-downloads/')

'crucible-downloads'
no_files bool

Skip file download, save record.json only.

False
no_record bool

Skip saving record.json, just download files.

False
overwrite_existing bool

Overwrite existing files (default: True)

True
include Optional[List[str]]

Glob patterns - only download matching files

None
exclude Optional[List[str]]

Glob patterns - skip matching files

None

Returns:

Type Description
List[str]

List[str]: Paths of all downloaded items (record.json + data files)

add_thumbnail(dataset_mfid, image, thumbnail_name=None)

Encode and add a thumbnail, returning the API thumbnail record.

get_thumbnails(dataset_mfid, limit=DEFAULT_LIMIT)

Get all thumbnails for a dataset.

The limit parameter is retained for compatibility. The API route is not paginated and returns every thumbnail.

update_thumbnail(dataset_mfid, thumbnail_id, image=None, thumbnail_name=None)

Rename or replace a thumbnail and return the updated API record.

delete_thumbnail(dataset_mfid, thumbnail_id)

Delete a thumbnail from a dataset.

add_keyword(dataset_mfid, keyword)

Add a keyword to a dataset.

Parameters:

Name Type Description Default
dataset_mfid str

Dataset MFID

required
keyword str

Keyword/tag to associate with dataset

required

Returns:

Name Type Description
Dict Dict

Keyword object with updated usage count

get_keywords(dataset_mfid, limit=DEFAULT_LIMIT)

List keywords, optionally filtered by dataset.

Parameters:

Name Type Description Default
dataset_mfid str

Dataset MFID

required
limit int

Maximum number of results to return

DEFAULT_LIMIT

Returns:

Type Description
List[Dict]

List[Dict]: Keyword objects with keyword text and num_datasets counts

Link a sample to a dataset.

Parameters:

Name Type Description Default
dataset_mfid str

Dataset MFID

required
sample_mfid str

Sample MFID

required

Returns:

Name Type Description
Dict Dict

Information about the created link

Remove the link between a dataset and a sample.

Requires admin permissions.

Parameters:

Name Type Description Default
dataset_mfid str

Dataset MFID

required
sample_mfid str

Sample MFID

required

Returns:

Name Type Description
Dict Dict

Deletion confirmation

Link two datasets with a parent-child relationship.

Parameters:

Name Type Description Default
parent_mfid str

Parent dataset MFID

required
child_mfid str

Child dataset MFID

required
relationship_type str

Kind of link, one of crucible.constants.RELATIONSHIP_TYPES. Describes the child relative to the parent.

None

Returns:

Name Type Description
Dict Dict

Information about the created link

Remove the parent-child link between two datasets.

Parameters:

Name Type Description Default
parent_mfid str

Parent dataset MFID

required
child_mfid str

Child dataset MFID

required

Returns:

Name Type Description
Dict Dict

Deletion confirmation

list_parents(child_mfid, limit=DEFAULT_LIMIT, offset=0, relationship_type=None, **kwargs)

List the parents of a given dataset with optional filtering.

Parameters:

Name Type Description Default
child_mfid str

Child dataset MFID

required
limit int

Maximum number of results to return

DEFAULT_LIMIT
offset int

Starting position in the full result set (default: 0)

0
relationship_type str

Only return parents linked with this kind of link, one of crucible.constants.RELATIONSHIP_TYPES.

None
**kwargs Any

Query parameters for filtering datasets

{}

Returns:

Type Description
List[Dict]

List[Dict]: Parent datasets

list_children(parent_mfid, limit=DEFAULT_LIMIT, offset=0, relationship_type=None, **kwargs)

List the children of a given dataset with optional filtering.

Parameters:

Name Type Description Default
parent_mfid str

Parent dataset MFID

required
limit int

Maximum number of results to return

DEFAULT_LIMIT
offset int

Starting position in the full result set (default: 0)

0
relationship_type str

Only return children linked with this kind of link, one of crucible.constants.RELATIONSHIP_TYPES.

None
**kwargs Any

Query parameters for filtering datasets

{}

Returns:

Type Description
List[Dict]

List[Dict]: Children datasets

graph(dataset_mfid, recursive=False, as_networkx=False)