kg_microbe package

Subpackages

Submodules

kg_microbe.bactotraits_to_mongo module

BactoTraits CSV to MongoDB JSON converter.

Converts BactoTraits semicolon-separated CSV files with hierarchical headers into MongoDB-compatible JSON format with proper field nesting and one-hot encoding.

kg_microbe.bactotraits_to_mongo.build_nested_dict(path_parts, value)

Build nested dictionary from path parts.

Return type:

Dict

kg_microbe.bactotraits_to_mongo.clean_field_name(name)

Convert field name to MongoDB-compatible format.

Return type:

str

kg_microbe.bactotraits_to_mongo.create_field_path(category, field_name)

Create hierarchical field path from category and field name.

Return type:

str

kg_microbe.bactotraits_to_mongo.forward_fill_headers(header_row)

Forward fill empty cells in header row.

Return type:

List[str]

kg_microbe.bactotraits_to_mongo.main()

Convert BactoTraits CSV to MongoDB JSON format.

kg_microbe.bactotraits_to_mongo.merge_dicts(dict1, dict2)

Deep merge two dictionaries.

Return type:

Dict

kg_microbe.bactotraits_to_mongo.parse_bactotraits_to_mongo(input_file)

Parse BactoTraits CSV into MongoDB-compatible JSON. :rtype: List[Dict[str, Any]]

  • Uses hierarchical header structure (category + field name)

  • Creates nested paths using forward-filled categories

  • Only includes non-zero, non-NA, non-empty values

  • Handles one-hot encoding by preserving all non-zero values

  • Splits comma-separated values into arrays

kg_microbe.bactotraits_to_mongo.parse_value(value)

Parse value, converting numbers and filtering out NA/empty.

Return type:

Any

kg_microbe.bactotraits_to_mongo.split_list_values(value)

Split comma-separated values into arrays where appropriate.

Return type:

Any

kg_microbe.download module

Download resources from YAML file.

exception kg_microbe.download.UnknownDownloadTagError

Bases: ValueError

Raised when a requested -t tag matches no entry in the download config.

A dedicated type so the CLI can report bad tags as usage errors without also swallowing unrelated ValueErrors raised deeper in the download (JSON decode failures and pydantic ValidationError are both ValueError subclasses, and were being presented as “Invalid value” for -t).

kg_microbe.download.download(yaml_file, output_dir, snippet_only, ignore_cache=False, tags=None)

Download data files from list of URLs.

DL based on config (default: download.yaml) into data directory (default: data/).

Parameters:

yaml_file (str) – A string pointing to the yaml file

:param utilized to facilitate the downloading of data. :type output_dir: str :param output_dir: A string pointing to the location to download data to. :type snippet_only: bool :param snippet_only: Downloads only the first 5 kB of the source,for testing and file checks. :type ignore_cache: bool :param ignore_cache: Ignore cache and download files even if they exist [false] :type tags: Optional[Sequence[str]] :param tags: Only download entries carrying one of these tags. None or empty

means download everything.

Return type:

None

Returns:

None.

kg_microbe.query module

Query module.

kg_microbe.query.parse_query_yaml(yaml_file)

Parse a YAML file and return the results as a dictionary.

Parameters:

yaml_file – YAML file to parse.

Return type:

dict

Returns:

A dictionary of results from the YAML file.

kg_microbe.query.result_dict_to_tsv(result_dict, outfile)

Write a dictionary to a TSV file.

Parameters:
  • result_dict (dict) – Dictionary to write to TSV file.

  • outfile (str) – TSV file to write to.

Return type:

None

kg_microbe.query.run_query(query, endpoint, return_format='json')

Run a SPARQL query and return the results as a dictionary.

Parameters:
  • query (str) – SPARQL query to run.

  • endpoint (str) – SPARQL endpoint to query.

  • return_format – Format of the returned data.

Return type:

dict

Returns:

A dictionary of results from the SPARQL query.

kg_microbe.run module

Drive KG download, transform, merge steps.

kg_microbe.transform module

Transform module.

class kg_microbe.transform.LazyTransform(dotted_path)

Bases: object

Resolve one transform class only when it is used.

property transform_class

Import and return the registered transform class.

exception kg_microbe.transform.TransformBatchError(failed, skipped)

Bases: RuntimeError

One or more sources in a batch failed; raised after the others ran.

Carries failed (source -> exception) and skipped (source -> the upstream sources it declares in TRANSFORM_INPUTS that failed first), and renders both, so the exit is non-zero and the summary is unmissable.

kg_microbe.transform.transform(input_dir, output_dir, sources=None, show_status=True)

Transform based on resource and class declared in DATA_SOURCES.

Call scripts in kg_microbe/transform/[source name]/ to transform each source into a graph format that KGX can ingest directly, in either TSV or JSON format: https://github.com/biolink/kgx/blob/master/data-preparation.md

Parameters:
  • input_dir (Optional[Path]) – A string pointing to the directory to import data from.

  • output_dir (Optional[Path]) – A string pointing to the directory to output data to.

  • sources (Optional[List[str]]) – A list of sources to transform. A registered source name, or one ontology name from ONTOLOGIES_MAP (ec, chebi, …) to refresh that ontology alone (#690).

Raises:
  • ValueError – If a requested source is not registered in DATA_SOURCES.

  • FileNotFoundError – If a selected source declares a curation input that is not on disk; nothing runs (#685).

  • TransformBatchError – After the batch, if any source failed. A failure is isolated to its source: later sources still run unless they declare the failed one in TRANSFORM_INPUTS, in which case they are skipped rather than built on stale upstream output (#685). BaseException (FatalOntologyError, Ctrl-C) still aborts at once.

Return type:

None

Module contents

kg-microbe package.

kg_microbe.download(yaml_file, output_dir, snippet_only, ignore_cache=False, tags=None)

Download data files from list of URLs.

DL based on config (default: download.yaml) into data directory (default: data/).

Parameters:

yaml_file (str) – A string pointing to the yaml file

:param utilized to facilitate the downloading of data. :type output_dir: str :param output_dir: A string pointing to the location to download data to. :type snippet_only: bool :param snippet_only: Downloads only the first 5 kB of the source,for testing and file checks. :type ignore_cache: bool :param ignore_cache: Ignore cache and download files even if they exist [false] :type tags: Optional[Sequence[str]] :param tags: Only download entries carrying one of these tags. None or empty

means download everything.

Return type:

None

Returns:

None.