Connecting to the MGnify API#
The MGnify API has many endpoints providing access to multiple types of resources such as studies, samples, analyses, genomes, and more. This notebook shows you how to
Discover what resources are available
Inspect what query parameters each resource accepts
Build a set of queries
# uncomment below if colab
#!pip install mgnipy
Starting up a mgnipy.MGnipy client#
For more details on configuring mgnipy and the default configuration go to the config info page
from mgnipy import MGnipy
# init
MG = MGnipy(
# add a configuration
cache_dir=None,
)
# print the MGnipy instance to see its configuration (credentials are not printed)
print(MG)
MGnipy(config=api_version=<SupportedApiVersions.V2: 'v2'> base_url=HttpUrl('https://www.ebi.ac.uk/') cache_dir=None)
Exploring the available resources#
We can learn more about the MGnify API and its available resources via the MGnipy client.
# to list all avail resources
print(MG.list_resources())
['analyses', 'analysis', 'assemblies', 'assembly', 'genomes', 'genome', 'publications', 'publication', 'samples', 'sample', 'studies', 'study', 'runs', 'run', 'biomes', 'biome', 'miscellaneous', 'catalogues', 'catalogue', 'private_studies']
the plural resources (e.g.
analyses,studies) represent collection/list endpoints from the APIe.g.
Studies: Lists of MGnify studies
Analyses: Lists of MGnify pipeline analyses
Usually we use
MGnifyListendpoints to search or filter for a list of the resourcethe singular (e.g.
analysis,study) represent a detail endpoint (i.e., getting the details of a single study, analysis, etc)e.g.
Study: Detailed metadata for a study given its study accession id
Analysis: Detailed metadata for a MGnify Analysis given its MGnify analysis accession id
MGnifyDetailendpoints are used to get the metadata for a given item.
A description of the resource and corresponding API endpoint can be viewed using the helper methods .describe_resource() for a given one or .describe_resources() to see all.
print("studies list endpoint:")
MG.describe_resource("studies")
print("\n----------\n")
print("analysis detail endpoint:")
MG.describe_resource("analysis")
studies list endpoint:
List all studies analysed by MGnify
MGnify studies inherit directly from studies (or projects) in ENA.
Supported parameters:
- order: ListMgnifyStudiesOrderType0 | None | Unset
- biome_lineage: None | str | Unset The lineage to match, including all descendant biomes
- has_analyses_from_pipeline: None | PipelineVersions | Unset If set, will only show studies with analyses from the specified MGnify pipeline version
- search: None | str | Unset Search within study titles and accessions
- page: int | Unset Default: 1.
- page_size: int | None | Unset
----------
analysis detail endpoint:
Get MGnify analysis by accession
MGnify analyses are accessioned with an MYGA-prefixed identifier and correspond to an individual Run
or Assembly analysed by a Pipeline.
Supported parameters:
- accession: str
Accessing a resource#
To use a given endpoint you can access it as an attribute of your mgnipy.MGnipy instance (e.g. MG.<chosen_resource>) which returns a resource proxy (aka endpoint-specific MGnifier)
# accessing Studies proxy as an attribute of MGnipy instance
studies = MG.studies
again to help there are helper functions for each resource proxy such as .list_supported_params() .describe_endpoint()
# print for more info
print(studies, "\n----------\n")
# or helper to list supported query params for the endpoint
print(studies.list_supported_params(), "\n----------\n")
# or a helper to describe corresponding API endpoint
studies.describe_endpoint()
<class 'mgnipy.V2.proxies.studies.Studies'> for 'studies' resource
- Endpoint: 'mgnipy.emgapi_v2_client.api.studies.list_mgnify_studies'
- Params: {}
- Child resource: 'study'
----------
['order', 'biome_lineage', 'has_analyses_from_pipeline', 'search', 'page', 'page_size']
----------
List all studies analysed by MGnify
MGnify studies inherit directly from studies (or projects) in ENA.
Supported parameters:
- order: ListMgnifyStudiesOrderType0 | None | Unset
- biome_lineage: None | str | Unset The lineage to match, including all descendant biomes
- has_analyses_from_pipeline: None | PipelineVersions | Unset If set, will only show studies with analyses from the specified MGnify pipeline version
- search: None | str | Unset Search within study titles and accessions
- page: int | Unset Default: 1.
- page_size: int | None | Unset
Notice how the configuration (e.g. cache_dir=None) was automatically passed to the proxy instance 🙌
Building a query set#
Using the supported params we can refine our query of the resource. For example, for Studies list we can .filter by search and has_analyses_from_pipeline
We can pass our search params:
when calling the resource e.g.
MG.studies(<param>=<value>)or byusing
.filter(<param>=<value>)after
# MGnifyList example with studies endpoint
# 1. at init of resource
filtered_studies = MG.studies(search="chicken")
filtered_studies.explain()
# or
print("\n----------\n")
# 2. filter method
filtered_studies = studies.filter(search="chicken")
filtered_studies.explain()
https://www.ebi.ac.uk/metagenomics/api/v2/studies?search=chicken&page=1
https://www.ebi.ac.uk/metagenomics/api/v2/studies?search=chicken&page=2
----------
https://www.ebi.ac.uk/metagenomics/api/v2/studies?search=chicken&page=1
https://www.ebi.ac.uk/metagenomics/api/v2/studies?search=chicken&page=2
Tip
explain provides a preview of the set of query urls to be called to fulfil our search and populate the Studies list
MGnifyDetails can also be “filtered” but basically only by accession/id. For example:
# MGnifyDetail example with .study
# 1. at init of resource
study = MG.study(accession="MGYS00000653")
# 2. filter method
study = MG.study
study = study.filter(accession="MGYS00000653")
study.explain()
https://www.ebi.ac.uk/metagenomics/api/v2/studies/MGYS00000653
Wrap Up:#
This page was a quick start demonstration of:
✅ Start up a
mgnipy.MGnipyclient with your desired configuration✅ Search in MGnify resources using a
MGnifierglass⬜ Receive a MGazine of MGnify datasets (go to next page)