DatasetOperations¶
Access via client.datasets.
File operations are also available via client.files (by MFID) and ingestion via client.ingestions.
crucible.resources.datasets.DatasetOperations
¶
Dataset-related API operations.
Access via: client.datasets.get(), client.datasets.list(), etc.
get(dataset_mfid, include_metadata=False, include_links=False, include_owner=True)
¶
Get a dataset by its canonical MFID.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset_mfid
|
str
|
Dataset MFID |
required |
include_metadata
|
bool
|
Whether to include scientific metadata |
False
|
include_links
|
bool
|
Whether to include immediate parent/child/associated links |
False
|
include_owner
|
bool
|
Resolve owner_orcid into a public-safe user object (default: True) |
True
|
Returns:
| Name | Type | Description |
|---|---|---|
Dict |
Dict
|
Dataset object with optional metadata and links |
list(sample_mfid=None, include_metadata=False, include_links=False, include_owner=False, limit=DEFAULT_LIMIT, offset=0, accessible_to_user=None, accessible_to_project=None, instrument_mfid=None, project_id=None, project_mfid=None, project_scope=None, **kwargs)
¶
List datasets with optional filtering and automatic pagination.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sample_mfid
|
str
|
If provided, returns datasets for this sample |
None
|
limit
|
int
|
Maximum total results to return (default: 100). Larger requests are handled transparently by following the server's keyset cursor. Pass None to fetch all matches. |
DEFAULT_LIMIT
|
offset
|
int
|
Deprecated. The /datasets collection uses keyset pagination and ignores offset, including when filtered by sample_mfid. |
0
|
include_metadata
|
bool
|
Include scientific metadata in results |
False
|
include_links
|
bool
|
Include linked resources (parents, children, associated) per dataset |
False
|
include_owner
|
bool
|
Resolve owner_orcid into a public-safe user object per dataset |
False
|
accessible_to_user
|
Optional[Union[str, Sequence[str]]]
|
User reference or references whose effective access must include every result |
None
|
accessible_to_project
|
Optional[Union[str, Sequence[str]]]
|
Project reference or references whose direct access must include every result |
None
|
instrument_mfid
|
Optional[str]
|
Canonical instrument MFID assigned to every result |
None
|
project_id
|
Optional[str]
|
Project slug used to scope results by project relationship |
None
|
project_mfid
|
Optional[str]
|
Canonical project MFID used to scope results by project relationship |
None
|
project_scope
|
Optional[str]
|
Project relationship to include: assigned, shared, or all. Requires project_id or project_mfid and defaults to assigned. |
None
|
**kwargs
|
Any
|
Query parameters for filtering. Supported fields include: keyword, unique_id, public, dataset_name, owner_orcid, project_id, instrument_name, timestamp, size, data_format, data_type, measurement, session_name. Filters expect exact matches (case sensitive) except for keywords, which are case insensitive and match substrings. |
{}
|
Returns:
| Type | Description |
|---|---|
List[Dict]
|
List[Dict]: Dataset objects matching filter criteria |
count(project_id=None, project_mfid=None, project_scope=None, **kwargs)
¶
Return the total number of datasets matching the given filters without fetching items.
search(q, project_id=None, limit=20)
¶
Fuzzy name search across datasets. Available to all authenticated users.
Matches against dataset_name. Returns datasets the caller can read. For scientific metadata search use search_metadata().
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
q
|
str
|
Search term (min 3 chars). Typo-tolerant — "pero" finds "perovskite". |
required |
project_id
|
Optional[str]
|
Optional project to scope results to. |
None
|
limit
|
int
|
Max results (default 20, max 50). |
20
|
Returns:
| Type | Description |
|---|---|
List[Dict]
|
List[Dict]: Matching DatasetResponse records, ranked by relevance. |
create(dataset, scientific_metadata=None, keywords=None, files=None, upload_files=True, files_to_upload=None, ingestor=None, verbose=False, wait_for_ingestion_response=False, skip_ingestion=False)
¶
Create a new dataset record with scientific metadata and keywords.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset
|
Dataset
|
Dataset object with dataset details. Use |
required |
scientific_metadata
|
dict
|
Scientific metadata |
None
|
keywords
|
list
|
Keywords to associate with dataset |
None
|
files
|
list
|
Files to attach. Each item is either a local path (str) or an AssociatedFile describing a file that lives elsewhere (Globus, NERSC, a shared filesystem, etc.). - A str path is uploaded to GCS when upload_files=True (default), or cataloged by its resolved absolute path (storage_backend='local', no upload) when upload_files=False. - An AssociatedFile is always cataloged via add_remote_file() - it must set storage_backend to something other than 'gcs', since there's no local file to upload from a model description. |
None
|
upload_files
|
bool
|
Whether str paths in |
True
|
files_to_upload
|
list
|
Deprecated alias for |
None
|
Returns:
Dict: created_record, scientific_metadata_record, dataset_mfid, dsid, files
(files is the per-item result of adding each entry in files,
in the same order). dsid is retained for compatibility.
update(dataset_mfid, **updates)
¶
Update an existing dataset with new field values.
'owner_orcid' and 'project_id' are no longer accepted here (422) - use transfer_ownership() / reassign_project() instead. The deprecated 'public' field delegates to set_public() or set_private(). Instrument reassignment is not available through generic PATCH. Omit 'instrument_id' and 'instrument_name' unless resubmitting their current values for compatibility.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset_mfid
|
str
|
Dataset MFID |
required |
**updates
|
Any
|
Fields to update (e.g., dataset_name="New Name") |
{}
|
Returns:
| Name | Type | Description |
|---|---|---|
Dict |
Dict
|
Updated dataset object |
Example
client.datasets.update("my-dataset-id", dataset_name="Updated Name")
add_file(dataset_mfid, file_path, ingestion_class=None, wait_for_ingestion_response=False, multipart=True, chunk_size_mb=None, max_workers=None, skip_ingestion=False)
¶
Upload a file to a dataset and request ingestion.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset_mfid
|
str
|
Dataset MFID |
required |
file_path
|
str
|
Local path to the file |
required |
ingestion_class
|
Optional[str]
|
Ingestion class for the worker (e.g. 'lammps', 'nexus'). Defaults to the server-side default if omitted. |
None
|
wait_for_ingestion_response
|
bool
|
Block until ingestion completes. Ignored if skip_ingestion is True. |
False
|
multipart
|
bool
|
Use parallel multipart upload (default: True). Set to False to use the sequential resumable upload (slower but simpler). |
True
|
chunk_size_mb
|
Optional[int]
|
Override chunk size in MiB (uses config/default if None). |
None
|
max_workers
|
Optional[int]
|
Override number of upload threads (uses config/default if None). |
None
|
skip_ingestion
|
bool
|
Upload the file without requesting ingestion. Request it later with client.files.request_ingestion(mfid). |
False
|
Returns:
| Name | Type | Description |
|---|---|---|
Dict |
Dict
|
{'associated_file': AssociatedFileRead, 'ingestion_request': IngestionRequest} |
add_remote_file(dataset_mfid, file)
¶
Register a file that lives outside GCS (Globus, NERSC, a shared filesystem path, etc.) without uploading it.
Crucible only catalogs the pointer here — it never verifies the file exists or fetches bytes on your behalf. Internally this is a two-step API call (create, then set storage_path) hidden behind one method.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset_mfid
|
str
|
Dataset MFID |
required |
file
|
AssociatedFile
|
Must have |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Dict |
Dict
|
The created (and, if storage_path was given, updated) file record. |
list_files(dataset_mfid)
¶
List files attached to a dataset.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset_mfid
|
str
|
Dataset MFID |
required |
Returns:
| Type | Description |
|---|---|
List[Dict]
|
List[Dict]: File records (mfid, filename, storage_path, storage_backend, access_note, size, sha256_hash, dataset_mfid). For a 'gcs' file, storage_path is null until it has been ingested. |
get_download_links(dataset_mfid)
¶
Get signed download URLs for all ingested files in a dataset.
Returns:
| Name | Type | Description |
|---|---|---|
Dict |
Dict
|
Mapping of file MFID → signed URL. Empty dict if no ingested files. |
download(dataset_mfid, file_name=None, output_dir='crucible-downloads', no_files=False, no_record=False, overwrite_existing=True, include=None, exclude=None)
¶
Download a dataset's files and optionally save its record as JSON.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset_mfid
|
str
|
Dataset MFID |
required |
file_name
|
Optional[str]
|
Deprecated. Use include=['pattern'] with glob syntax. |
None
|
output_dir
|
str
|
Directory to save files (default: 'crucible-downloads/') |
'crucible-downloads'
|
no_files
|
bool
|
Skip file download, save record.json only. |
False
|
no_record
|
bool
|
Skip saving record.json, just download files. |
False
|
overwrite_existing
|
bool
|
Overwrite existing files (default: True) |
True
|
include
|
Optional[List[str]]
|
Glob patterns - only download matching files |
None
|
exclude
|
Optional[List[str]]
|
Glob patterns - skip matching files |
None
|
Returns:
| Type | Description |
|---|---|
List[str]
|
List[str]: Paths of all downloaded items (record.json + data files) |
add_thumbnail(dataset_mfid, image, thumbnail_name=None)
¶
Encode and add a thumbnail, returning the API thumbnail record.
get_thumbnails(dataset_mfid, limit=DEFAULT_LIMIT)
¶
Get all thumbnails for a dataset.
The limit parameter is retained for compatibility. The API route
is not paginated and returns every thumbnail.
update_thumbnail(dataset_mfid, thumbnail_id, image=None, thumbnail_name=None)
¶
Rename or replace a thumbnail and return the updated API record.
delete_thumbnail(dataset_mfid, thumbnail_id)
¶
Delete a thumbnail from a dataset.
add_keyword(dataset_mfid, keyword)
¶
Add a keyword to a dataset.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset_mfid
|
str
|
Dataset MFID |
required |
keyword
|
str
|
Keyword/tag to associate with dataset |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Dict |
Dict
|
Keyword object with updated usage count |
get_keywords(dataset_mfid, limit=DEFAULT_LIMIT)
¶
List keywords, optionally filtered by dataset.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset_mfid
|
str
|
Dataset MFID |
required |
limit
|
int
|
Maximum number of results to return |
DEFAULT_LIMIT
|
Returns:
| Type | Description |
|---|---|
List[Dict]
|
List[Dict]: Keyword objects with keyword text and num_datasets counts |
link_sample(dataset_mfid, sample_mfid)
¶
Link a sample to a dataset.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset_mfid
|
str
|
Dataset MFID |
required |
sample_mfid
|
str
|
Sample MFID |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Dict |
Dict
|
Information about the created link |
unlink_sample(dataset_mfid, sample_mfid)
¶
Remove the link between a dataset and a sample.
Requires admin permissions.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset_mfid
|
str
|
Dataset MFID |
required |
sample_mfid
|
str
|
Sample MFID |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Dict |
Dict
|
Deletion confirmation |
link(parent_mfid, child_mfid, relationship_type=None)
¶
Link two datasets with a parent-child relationship.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
parent_mfid
|
str
|
Parent dataset MFID |
required |
child_mfid
|
str
|
Child dataset MFID |
required |
relationship_type
|
str
|
Kind of link, one of crucible.constants.RELATIONSHIP_TYPES. Describes the child relative to the parent. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
Dict |
Dict
|
Information about the created link |
unlink(parent_mfid, child_mfid)
¶
Remove the parent-child link between two datasets.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
parent_mfid
|
str
|
Parent dataset MFID |
required |
child_mfid
|
str
|
Child dataset MFID |
required |
Returns:
| Name | Type | Description |
|---|---|---|
Dict |
Dict
|
Deletion confirmation |
list_parents(child_mfid, limit=DEFAULT_LIMIT, offset=0, relationship_type=None, **kwargs)
¶
List the parents of a given dataset with optional filtering.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
child_mfid
|
str
|
Child dataset MFID |
required |
limit
|
int
|
Maximum number of results to return |
DEFAULT_LIMIT
|
offset
|
int
|
Starting position in the full result set (default: 0) |
0
|
relationship_type
|
str
|
Only return parents linked with this kind of link, one of crucible.constants.RELATIONSHIP_TYPES. |
None
|
**kwargs
|
Any
|
Query parameters for filtering datasets |
{}
|
Returns:
| Type | Description |
|---|---|
List[Dict]
|
List[Dict]: Parent datasets |
list_children(parent_mfid, limit=DEFAULT_LIMIT, offset=0, relationship_type=None, **kwargs)
¶
List the children of a given dataset with optional filtering.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
parent_mfid
|
str
|
Parent dataset MFID |
required |
limit
|
int
|
Maximum number of results to return |
DEFAULT_LIMIT
|
offset
|
int
|
Starting position in the full result set (default: 0) |
0
|
relationship_type
|
str
|
Only return children linked with this kind of link, one of crucible.constants.RELATIONSHIP_TYPES. |
None
|
**kwargs
|
Any
|
Query parameters for filtering datasets |
{}
|
Returns:
| Type | Description |
|---|---|
List[Dict]
|
List[Dict]: Children datasets |