mgnipy.V2.datasets package

Contents

mgnipy.V2.datasets package#

class MGazine(downloads, config=None, *, client=None, mgnify_studies=None, mgnify_analyses=None, mgnify_runs=None, mgnify_samples=None, mgnify_assemblies=None, biosamples_metadata=None, obs=None)[source]#

Bases: StreamMixin, ClientManagerMixin, MetadataSettersMixin

Reads or downloads datasets from MGnify.

MGazine is a class for managing and downloading datasets from MGnify. - Accepts a list of download-like dictionaries (for example the objects returned by the MGnify API for downloads) and provides simple streaming and download helpers. - Supports grouping datasets by pipeline version and short description, and provides methods for downloading individual files or all files in the MGazine.

Parameters:
  • downloads (list of dict ) – A list of download-like dictionaries, each containing keys such as alias, url, file_type, download_group, short_description, and pipeline_version.

  • config (MGnipyConfig, optional) – An optional configuration object for MGnipy. If not provided, a default configuration is used.

  • client (Client or AuthenticatedClient, optional) – An optional client object for making HTTP requests. If not provided, a default client is used.

  • mgnify_[studies|analyses|runs|samples|assemblies] (list of dict , optional) – Lists of dictionaries containing metadata for each respective MGnify dataset.

  • biosamples_metadata (list of dict , optional) – A list of dictionaries containing metadata for BioSamples.

  • mgnify_studies (list [dict [str , Any]] | None)

  • mgnify_analyses (list [dict [str , Any]] | None)

  • mgnify_runs (list [dict [str , Any]] | None)

  • mgnify_samples (list [dict [str , Any]] | None)

  • mgnify_assemblies (list [dict [str , Any]] | None)

  • obs (list [dict [str , Any]] | None)

downloads#

The list of download-like dictionaries provided during initialization.

Type:

list of dict

downloads_df[source]#

A DataFrame representation of the downloads, with columns such as alias, url, and file_type.

Type:

pandas.DataFrame

Return type:

DataFrame

aliases#

A list of all download aliases extracted from the downloads.

Type:

list of str

urls#

An alias for url_list, providing a list of all download URLs.

Type:

list of str

url_list#

A list of URLs extracted from the downloads.

Type:

list of str

url_dict#

A dictionary mapping each download alias to its corresponding URL.

Type:

dict

lazy_merged#

A lazy frame containing the merged datasets, if initialized.

Type:

polars.LazyFrame or None

short_desc#

The short description of the MGazine, derived from the downloads. If multiple short descriptions are present, a warning is issued.

Type:

str

Example

>>> downloads = [
...    {"alias": "a", "url": "/tmp/a.txt", "file_type": "txt", "short_description": "desc1", "pipeline_version": "v5"},
...    {"alias": "boop", "url": "/tmp/b.fasta", "file_type": "fasta", "short_description": "desc2", "pipeline_version": "v5"},
... ]
>>> mg = MGazine(downloads)
>>> print(mg)
MGazine containing:
- MGnify pipeline versions: ['v5']
- Number of downloads: 2
- Short descriptions: ['desc1', 'desc2']
- Nonempty metadata sets:
async aclose()#
async adownload(to_dir, alias=None, *, url=None, filename=None, overwrite=False, hide_progress=False)[source]#

Asynchronously download a file from an alias or URL.

Parameters:
  • to_dir (DirectoryPath) – Directory where the file will be saved.

  • alias (str or None, optional) – Download alias known to this MGazine instance.

  • url (str or None, optional) – Direct URL to fetch. Either alias or url must be provided.

  • filename (str or None, optional) – Filename to use for the saved file. When omitted the alias is used.

  • httpx_aclient (httpx.AsyncClient, optional) – Optional httpx.AsyncClient to use for the HTTP request.

  • overwrite (bool , optional) – If False and the destination file already exists the download is skipped. When True the existing file will be overwritten.

  • hide_progress (bool , optional) – Disable the progress bar when True.

Raises:

ValueError – If neither alias nor url is provided.

Examples

downloads = [ … { … “alias”: “example.txt”, … “url”: “http://ex/x ”, … “file_type”: “txt”, … }] mg = MGazine(downloads) await mg.adownload(“download_to_here”, alias=”example.txt”) # doctest: +SKIP

async adownload_all(to_dir, overwrite=False, hide_progress=False)[source]#

Asynchronously download all files known to this MGazine.

Parameters:
  • to_dir (DirectoryPath) – Directory where the files will be saved.

  • overwrite (bool , optional) – Passed to adownload to control overwriting behavior.

  • hide_progress (bool , optional) – Disable progress bars when True.

Notes

This helper creates a single async HTTP client and schedules concurrent adownload calls for all aliases.

Examples

>>> downloads = [
...     {"alias": "example.txt", "url": "http://ex/x", "file_type": "txt"},
...     {"alias": "example2.fasta.gz", "url": "http://ex/x2", "file_type": "fasta"},
... ]
>>> mg = MGazine(downloads)
>>> await mg.adownload_all("download_to_here")
property aliases: list [str ]#

Return a list of all download aliases.

Example

>>> downloads = [{"alias": "example.txt", "url": "http://ex/x"}]
>>> MGazine(downloads).aliases
['example.txt']
append_biosamples_metadata(value)#
Parameters:

value (dict [str , Any ])

append_mgnify_analyses(value)#
Parameters:

value (dict [str , Any ])

append_mgnify_assemblies(value)#
Parameters:

value (dict [str , Any ])

append_mgnify_runs(value)#
Parameters:

value (dict [str , Any ])

append_mgnify_samples(value)#
Parameters:

value (dict [str , Any ])

append_mgnify_studies(value)#
Parameters:

value (dict [str , Any ])

append_obs(value)#
Parameters:

value (dict [str , Any ])

property async_httpx_client: AsyncClient#

Get the asynchronous httpx client instance from the AuthenticatedClient.

Returns:

The asynchronous httpx client instance.

Return type:

httpx.AsyncClient

property available_metadata_sets: list [str ]#

Return a list of available metadata sets in the MGazine.

This property checks which metadata sets (e.g., studies, analyses, runs, samples, assemblies, biosamples) are non-empty and returns their names as a list.

Returns:

A list of names of non-empty metadata sets available in the MGazine.

Return type:

list of str

Examples

>>> mg = MGazine(downloads)
>>> mg.available_metadata_sets
['mgnify_studies', 'mgnify_analyses', 'mgnify_runs']
property biosamples_metadata: ResultsHandler#
by_downloads_col(col)[source]#

Group downloads by a specified column in the downloads dataframe.

Parameters:

col (str ) – The column name to group by.

Returns:

A dictionary where keys are unique values from the specified column and values are lists of download dictionaries.

Return type:

dict

Raises:

ValueError – If the specified column is not present in the downloads dataframe.

close()#
download(to_dir, alias=None, *, url=None, filename=None, overwrite=False, hide_progress=False)[source]#

Download a file by its alias or URL.

Download a file from an alias or URL to a local directory.

Parameters:
  • to_dir (DirectoryPath) – Directory where the file will be saved.

  • alias (str or None, optional) – Download alias known to this MGazine instance. When provided the corresponding URL from the instance’s downloads list is used.

  • url (str or None, optional) – Direct URL to fetch. Either alias or url must be provided.

  • filename (str or None, optional) – Filename to use for the saved file. When omitted the alias is used.

  • overwrite (bool , optional) – If False and the destination file already exists the download is skipped. When True the existing file will be overwritten.

  • hide_progress (bool , optional) – Disable the progress bar when True.

Raises:

ValueError – If neither alias nor url is provided.

Examples

mg = MGazine(downloads) # doctest: +SKIP mg.download(“download_to_here”, alias=”example.txt”) # doctest: +SKIP

download_all(to_dir, hide_progress=False, overwrite=False)[source]#

Download all files known to this MGazine instance.

Parameters:
  • to_dir (DirectoryPath) – Directory where the files will be saved.

  • hide_progress (bool , optional) – Disable per-file and overall progress bars when True.

  • overwrite (bool , optional) – Passed to download to control overwriting behavior.

Notes

This helper calls download for each alias present in the instance’s downloads list.

Examples

>>> downloads = [
...     {"alias": "example.txt", "url": "http://ex/x", "file_type": "txt"},
...     {"alias": "example2.fasta.gz", "url": "http://ex/x2", "file_type": "fasta"},
... ]
>>> mg = MGazine(downloads)
>>> mg.download_all("download_to_here")
downloads_df(**pd_kwargs)[source]#

The downloads as a DataFrame.

This returns a pandas.DataFrame of all downloads. The dataframe should contain columns such as alias, url and file_type (TODO pandera).

Parameters:

pd_kwargs (dict ) – Additional keyword arguments to pass to the pandas.DataFrame constructor.

Returns:

A DataFrame containing the downloads information

Return type:

pd.DataFrame

Examples

>>> downloads = [{"alias": "example.txt", "url": "http://ex/x", "file_type": "txt"}]
>>> mag = MGazine(downloads)
>>> df = mag.downloads_df(index=["boop"])
property httpx_client: Client#

Get the synchronous httpx client instance from the AuthenticatedClient.

Returns:

The synchronous httpx client instance.

Return type:

httpx.Client

lazy_concat(aliases=None, urls=None, how='vertical_relaxed', **pl_kwargs)[source]#

Return a concatenated Polars LazyFrame of the datasets corresponding to the provided aliases or URLs.

Parameters:
  • aliases (list [str ] or None, optional) – List of download aliases to stream and concatenate. If provided, this takes precedence over urls.

  • urls (list [str ] or None, optional) – List of download URLs to stream and concatenate. Used only if aliases is not provided.

  • how (str , optional) – Concatenation method. Options include ‘vertical’, ‘horizontal’, ‘vertical_relaxed’, etc. See Polars documentation for details.

  • **pl_kwargs – Additional keyword arguments to pass to the Polars concatenation function.

Returns:

A Polars LazyFrame representing the concatenated datasets.

Return type:

pl.LazyFrame

property lazy_merged: LazyFrame | None #

Return the current lazy merged Polars LazyFrame if available.

Returns:

The current lazy merged Polars LazyFrame, or None if not set.

Return type:

pl.LazyFrame or None

list_pipeline_version()[source]#

A list of unique pipeline versions in the MGazine.

Returns:

A list of unique pipeline versions extracted from the downloads.

Return type:

list of str

Examples

>>> downloads = [
...     {"alias": "example.txt", "url": "http://ex/x", "pipeline_version": 'v4_1'},
...     {"alias": "example2.txt", "url": "http://ex/x2", "pipeline_version": 'v5'},
... ]
>>> MGazine(downloads).list_pipeline_version()
['v4_1', 'v5']
list_short_descriptions()[source]#

A list of unique short descriptions of the downloads.

The unique short descriptions in the given column

Returns:

A list of unique short descriptions extracted from the downloads.

Return type:

list of str

Examples

>>> downloads = [
...     {"alias": "example.txt", "short_description": "shortdesc1"},
...     {"alias": "boo.txt", "short_description": "shortdesc1"},
...     {"alias": "example2.txt", "short_description": "shortdesc2"},
... ]
>>> MGazine(downloads).list_short_descriptions()
['shortdesc1', 'shortdesc2']
property mgnify_analyses: MGnifyMetadata#
property mgnify_assemblies: MGnifyMetadata#
property mgnify_runs: MGnifyMetadata#
property mgnify_samples: MGnifyMetadata#
property mgnify_studies: MGnifyMetadata#
property obs: ResultsHandler#
obs_metadata(df_engine='pandas', expand_nested_dicts=True, drop_duplicates=False, how='left', coalesce=True, for_runs=None, index_col_name='_mgnipy_runs_accs')#
Parameters:
Return type:

DataFrame | DataFrame

renew_client()#

Init a new client instance and replace the existing one. This is useful if the current client has been closed or is no longer valid, allowing for a fresh start with a new HTTP client session.

property short_desc: str #

The short description of the MGazine.

This property returns the FIRST short description of the MGazine, which is derived from the downloads. If multiple short descriptions are present, a warning is issued.

status()#

Print the status of the MGnipy client, including the type of client and whether the synchronous and asynchronous httpx client sessions are open.

Return type:

None

stream(*, alias=None, url=None, chunksize=None, max_skip=5, **kwargs)#

Streams a single download based on its alias or url.

If chunksize is specified then iterators of dataframes or strings will be returned; otherwise the full data will be returned as a single object.

Supported formats and their handlers#

param alias:

The alias of the download to stream.

type alias:

Optional[str]

param url:

The url of the download to stream.

type url:

Optional[HttpUrl]

param chunksize:

The size of the chunks to read from the stream.

type chunksize:

Optional[int]

param max_skip:

The maximum number of rows to skip before raising an error. Default is 5.

type max_skip:

int, optional

param **kwargs:

Additional keyword arguments to pass to the streamer function.

returns:

The streamer result for the resolved alias or url.

rtype:

Any

Parameters:
  • alias (str | None)

  • url (HttpUrl | None)

  • chunksize (int | None)

  • max_skip (int )

Return type:

Any

stream_biom(url, **skbio_kwargs)#

Stream a biom file from a URL using scikit-bio’s read function. Refer there for more info.

Parameters:
  • url (str ) – The URL to the biom file to stream.

  • **skbio_kwargs – Additional keyword arguments passed to skbio.io.read(), such as into and verify.

Returns:

A generator yielding scikit-bio Sequence objects parsed from the biom file.

Return type:

Generator

stream_fasta(url, **skbio_kwargs)#

Stream a FASTA file from a URL using scikit-bio’s read function. Refer there for more info.

Parameters:
  • url (str ) – The URL to the FASTA file to stream.

  • **skbio_kwargs – Additional keyword arguments passed to skbio.io.read(), such as into and verify.

Returns:

A generator yielding scikit-bio Sequence objects parsed from the FASTA file.

Return type:

Generator

stream_gff(url, **skbio_kwargs)#

Stream a GFF file from a URL using scikit-bio’s read function. Refer there for more info.

Parameters:
  • url (str ) – The URL to the GFF file to stream.

  • **skbio_kwargs – Additional keyword arguments passed to skbio.io.read(), such as into and verify.

Returns:

A generator yielding scikit-bio Sequence objects parsed from the GFF file.

Return type:

Generator

stream_gzipped(url, chunksize=None, decode=False, encoding='utf-8', errors='replace', **httpx_kwargs)#

Stream a gzipped HTTP resource and present a file-like interface.

When chunksize is None the entire compressed payload is fetched and decompressed into memory. When chunksize is provided a streaming file-like object is returned.

Parameters:
Return type:

bytes | str | BufferedReader | TextIOWrapper

stream_html(url, **web_kwargs)#

Open an HTML URL in the default web browser.

Parameters:
  • url (str ) – The URL to open in the web browser.

  • **web_kwargs – Additional keyword arguments passed to webbrowser.open(), such as new and autoraise.

Returns:

True if the URL was opened successfully, False otherwise.

Return type:

bool

stream_json(url, chunksize=None, **httpx_kwargs)#
Parameters:
Return type:

dict | Generator

stream_jsonl(url, orient=None, chunksize=None, df_engine='pandas', **df_kwargs)#
Parameters:
  • url (str )

  • orient (Literal ['records', 'split', 'index', 'columns', 'values', 'table'] | None)

  • chunksize (int | None)

  • df_engine (Literal ['pandas', 'polars'] | None)

Return type:

dict

stream_pandas(url, sep='\t', chunksize=None, max_skip=5, low_memory=False, **pd_kwargs)#

Read a TSV from a URL or local file with resilient header handling.

The helper will retry with increasing skiprows when pandas raises a ParserError (useful for files with extra header lines). When chunksize is provided an iterator is returned.

Parameters:
  • url (str ) – The URL or local file path to read the TSV from.

  • sep (str ) – The delimiter to use (default is tab).

  • chunksize (int or None) – If an integer is provided, returns an iterator that yields DataFrames of that many rows. If None, returns a single DataFrame.

  • max_skip (int ) – The maximum number of lines to skip when trying to parse the TSV.

  • **pd_kwargs – Additional keyword arguments passed to pd.read_csv.

  • low_memory (bool )

Returns:

A DataFrame containing the TSV data, or an iterator yielding DataFrames if chunksize is specified.

Return type:

pd.DataFrame or TextFileReader

Raises:
  • ValueError – If chunksize is not a positive integer or None.

  • RuntimeError – If the TSV cannot be parsed after skipping up to max_skip lines.

  • Pandas ParserError – If the TSV cannot be parsed due to a format error (after retries).

stream_polars(url, sep='\t', chunksize=None, max_skip=5, low_memory=False, **pl_kwargs)#

Read a TSV from a URL or local file into a Polars DataFrame with resilient header handling.

The helper will retry with increasing skip_rows when Polars raises an error (useful for files with extra header lines). When chunksize is provided an iterator is returned.

Parameters:
  • url (str ) – The URL or local file path to read the TSV from.

  • sep (str ) – The delimiter to use (default is tab).

  • chunksize (int or None) – If an integer is provided, returns an iterator that yields DataFrames of that many rows. If None, returns a single DataFrame.

  • max_skip (int ) – The maximum number of lines to skip when trying to parse the TSV.

  • **pl_kwargs – Additional keyword arguments passed to pl.read_csv.

  • low_memory (bool )

Returns:

A Polars DataFrame containing the TSV data, or an iterator yielding DataFrames if chunksize is specified.

Return type:

pl.DataFrame or Iterator[pl.DataFrame]

Raises:
  • ValueError – If chunksize is not a positive integer or None.

  • RuntimeError – If the TSV cannot be parsed after skipping up to max_skip lines.

  • Polars Error – If the TSV cannot be parsed due to a format error (after retries).

stream_tree(url, **skbio_kwargs)#
Parameters:

url (str )

Return type:

Generator

stream_txt(url, chunksize=None, **httpx_kwargs)#

Stream a plain-text resource. When chunksize is None the full text is returned as a string. When chunksize is an integer the function yields lists of lines.

Parameters:
  • url (str ) – The URL to stream the text from.

  • chunksize (int or None) – If an integer is provided, yields lists of lines of that size. If None, yields the entire text as a single string.

  • httpx_client (httpx.Client, optional) – An optional httpx.Client to use for the request. If None, a new client will be created for the request.

  • **httpx_kwargs – Additional keyword arguments passed to the httpx.Client.request() method

Returns:

The full text as a string if chunksize is None, or a generator yielding lists of lines if chunksize is an integer.

Return type:

str or Generator

to_pandas(**pd_kwargs)[source]#
Return type:

DataFrame

to_polars()[source]#
Return type:

DataFrame

property url_dict: dict [str , dict ]#

Return mapping of alias to URL for all downloads.

Returns:

Dictionary mapping alias -> url (or None when no url is available).

Return type:

dict

Examples

>>> downloads = [{"alias": "example.txt", "url": "http://ex/x"}]
>>> MGazine(downloads).url_dict
{'example.txt': 'http://ex/x'}
property url_list#

Return a list of all download URLs.

Examples

>>> downloads = [{"alias": "example.txt", "url": "http://ex/x"}]
>>> MGazine(downloads).urls
['http://ex/x']
property urls: list [str | None ]#

Return a list of all download URLs. Same as url_list().

Examples

>>> downloads = [{"alias": "example.txt", "url": "http://ex/x"}]
>>> MGazine(downloads).urls
['http://ex/x']
class MTG(dataset, *, var_cols=None, var_index=None, obs_index='_mgnipy_runs_accs', mgnify_studies=None, mgnify_analyses=None, mgnify_runs=None, mgnify_samples=None, mgnify_assemblies=None, biosamples_metadata=None, obs=None)[source]#

Bases: MetadataSettersMixin

MGic the Gatherer combines a MGnify dataset with its metadata.

The MGic gatherer (MTG) takes a dataset as pandas or polars dataframe and MGnify or BioSamples metadata and combines them into a single object. MTG can be used to enrich the dataset with metadata, and to convert the dataset into different formats such as pandas, polars, or anndata.

Parameters:
  • dataset (pandas.DataFrame or polars.DataFrame) – The dataset to be combined with metadata. This can be a pandas or polars dataframe.

  • var_cols (list of str , optional) – A list of column names in the dataset that are considered variable columns. These columns will be in var_metadata() and excluded from obs_metadata()

  • mgnify_[studies|analyses|runs|samples|assemblies] (list of dict , optional) – Lists of dictionaries containing metadata for each respective MGnify dataset.

  • biosamples_metadata (list of dict , optional) – A list of dictionaries containing metadata for BioSamples.

  • var_index (str | None)

  • obs_index (str )

  • mgnify_studies (list [dict [str , Any]] | None)

  • mgnify_analyses (list [dict [str , Any]] | None)

  • mgnify_runs (list [dict [str , Any]] | None)

  • mgnify_samples (list [dict [str , Any]] | None)

  • mgnify_assemblies (list [dict [str , Any]] | None)

  • obs (list [dict [str , Any]] | None)

runs_accessions#

A list of all run accessions in the dataset. This is derived from the columns of the dataset that are not in var_cols.

Type:

list

X(df_engine='pandas')[source]#

Gets the feature matrix (X) from the dataset.

Basically transposes.

Parameters:

df_engine (Literal["polars", "pandas"], optional) – The DataFrame engine to use for the output. If “polars” is specified, a polars.DataFrame is returned; if “pandas” is specified, a pandas.DataFrame is returned.

Returns:

The feature matrix (X) containing the non-var columns from the dataset

Return type:

pl.DataFrame or pd.DataFrame

append_biosamples_metadata(value)#
Parameters:

value (dict [str , Any ])

append_mgnify_analyses(value)#
Parameters:

value (dict [str , Any ])

append_mgnify_assemblies(value)#
Parameters:

value (dict [str , Any ])

append_mgnify_runs(value)#
Parameters:

value (dict [str , Any ])

append_mgnify_samples(value)#
Parameters:

value (dict [str , Any ])

append_mgnify_studies(value)#
Parameters:

value (dict [str , Any ])

append_obs(value)#
Parameters:

value (dict [str , Any ])

property available_metadata_sets: list [str ]#

Return a list of available metadata sets in the MTG.

This property checks which metadata sets (e.g., studies, analyses, runs, samples, assemblies, biosamples) are non-empty and returns their names as a list.

Returns:

A list of names of non-empty metadata sets available in the MTG.

Return type:

list of str

property biosamples_metadata: ResultsHandler#
property mgnify_analyses: MGnifyMetadata#
property mgnify_assemblies: MGnifyMetadata#
property mgnify_runs: MGnifyMetadata#
property mgnify_samples: MGnifyMetadata#
property mgnify_studies: MGnifyMetadata#
property obs: ResultsHandler#
obs_metadata(*args, **kwargs)[source]#
Return type:

DataFrame | DataFrame

property runs_accessions: list #
to_anndata(drop_duplicates=True, **anndata_kwargs)[source]#
Parameters:

drop_duplicates (bool )

Return type:

AnnData

to_pandas()[source]#
Return type:

DataFrame

to_polars()[source]#
Return type:

DataFrame

var_metadata(df_engine='pandas')[source]#

Return the variable metadata as a dataframe.

Parameters:

df_engine (str , optional) – The dataframe engine to use. Can be “polars” or “pandas”. Default is “pandas”.

Returns:

A dataframe containing the variable metadata.

Return type:

pd.DataFrame or pl.DataFrame

Submodules#