kg_microbe package
Subpackages
- kg_microbe.merge_utils package
- Submodules
- kg_microbe.merge_utils.artifact_manifest module
- kg_microbe.merge_utils.external_node_closure module
- kg_microbe.merge_utils.invariants module
- kg_microbe.merge_utils.kgx_source module
- kg_microbe.merge_utils.local_context module
- kg_microbe.merge_utils.merge_kg module
- kg_microbe.merge_utils.progress module
- kg_microbe.merge_utils.source_admission module
- kg_microbe.merge_utils.source_freshness module
- kg_microbe.merge_utils.stats_provenance module
- Module contents
- kg_microbe.query_utils package
- kg_microbe.transform_utils package
- Subpackages
- kg_microbe.transform_utils.bacdive package
- kg_microbe.transform_utils.bactotraits package
- kg_microbe.transform_utils.bakta package
- kg_microbe.transform_utils.cog package
- kg_microbe.transform_utils.ctd package
- kg_microbe.transform_utils.disbiome package
- kg_microbe.transform_utils.example_transform package
- kg_microbe.transform_utils.gold package
- kg_microbe.transform_utils.gtdb package
- kg_microbe.transform_utils.kegg package
- kg_microbe.transform_utils.lpsn package
- kg_microbe.transform_utils.lpsn_api package
- kg_microbe.transform_utils.madin_etal package
- kg_microbe.transform_utils.mediadive package
- kg_microbe.transform_utils.metatraits package
- kg_microbe.transform_utils.metatraits_gtdb package
- kg_microbe.transform_utils.microbedecoder package
- kg_microbe.transform_utils.mim_ingredients package
- kg_microbe.transform_utils.ontologies package
- kg_microbe.transform_utils.ontologies_stubs package
- kg_microbe.transform_utils.prego package
- kg_microbe.transform_utils.rhea_mappings package
- kg_microbe.transform_utils.uniprot_functional_microbes package
- kg_microbe.transform_utils.uniprot_human package
- kg_microbe.transform_utils.uniprot_trembl package
- kg_microbe.transform_utils.wallen_etal package
- Submodules
- kg_microbe.transform_utils.constants module
- kg_microbe.transform_utils.transform module
TransformTransform.CODE_INPUTSTransform.DATA_DIRTransform.DATA_INPUTSTransform.DEFAULT_INPUT_DIRTransform.DEFAULT_OUTPUT_DIRTransform.OPTIONAL_CONSUMED_INPUTSTransform.OPTIONAL_RAW_CONSUMED_INPUTSTransform.REQUIRED_CONSUMED_INPUTSTransform.TRANSFORM_INPUTSTransform.begin_consumed_inputs()Transform.begin_dependency_admission()Transform.consume_input()Transform.consume_optional_input()Transform.consumed_input_snapshotsTransform.finalize()Transform.optional_consumed_inputsTransform.pass_through()Transform.run()Transform.verify_consumed_inputs()Transform.verify_declared_dependencies()
- Module contents
- Subpackages
- kg_microbe.utils package
- Submodules
- kg_microbe.utils.atomic_io module
- kg_microbe.utils.biolink_hierarchy module
- kg_microbe.utils.biolink_model module
- kg_microbe.utils.cas module
- kg_microbe.utils.chemical_mapping_utils module
ChemicalMappingLoaderChemicalMappingLoader.find_chebi_by_formula()ChemicalMappingLoader.find_chebi_by_name()ChemicalMappingLoader.find_chebi_by_xref()ChemicalMappingLoader.from_ingredient_lookup_bundle()ChemicalMappingLoader.get_canonical_name()ChemicalMappingLoader.get_category()ChemicalMappingLoader.get_formula()ChemicalMappingLoader.get_identifier_annotation_owners()ChemicalMappingLoader.get_node_enrichment()ChemicalMappingLoader.get_parents()ChemicalMappingLoader.get_synonyms()ChemicalMappingLoader.get_xrefs()ChemicalMappingLoader.resolve_ingredient_source()
PREDICATE_SEMANTICS_KEYSKOS_SEMANTICSfind_chebi_by_formula()find_chebi_by_name()find_chebi_by_xref()get_canonical_name()get_category()get_formula()get_hydrate_equivalents()get_mapping_load_audit()get_node_enrichment()get_parents()get_synonyms()get_xrefs()load_unified_mappings()normalize_chemical_primes()normalize_name()read_predicate_semantics()
- kg_microbe.utils.consolidate_categories module
- kg_microbe.utils.download_bacdive module
- kg_microbe.utils.download_manifest module
- kg_microbe.utils.download_utils module
- kg_microbe.utils.dummy_tqdm module
- kg_microbe.utils.external_identifiers module
- kg_microbe.utils.fix_list_representations module
- kg_microbe.utils.foodon_classification module
- kg_microbe.utils.go_authority module
- kg_microbe.utils.graph_canonicalization module
- kg_microbe.utils.graph_schema module
- kg_microbe.utils.ingredient_bundle module
ReviewedIngredientBundleReviewedIngredientBundle.active_xrefs()ReviewedIngredientBundle.annotation_document()ReviewedIngredientBundle.audit()ReviewedIngredientBundle.bind_to_transform()ReviewedIngredientBundle.canonical_owner()ReviewedIngredientBundle.enrich_node()ReviewedIngredientBundle.identifier_claims()ReviewedIngredientBundle.identifier_owners()ReviewedIngredientBundle.mappings()ReviewedIngredientBundle.occurrences()ReviewedIngredientBundle.policy_parity()ReviewedIngredientBundle.products()ReviewedIngredientBundle.resolve_source()ReviewedIngredientBundle.verify_current()ReviewedIngredientBundle.write_annotations()
build_ingredient_lookup_bundle()load_ingredient_lookup_bundle()
- kg_microbe.utils.ingredient_bundle_contract module
- kg_microbe.utils.ingredient_identity module
- kg_microbe.utils.ingredient_kgx module
- kg_microbe.utils.ingredient_scope module
- kg_microbe.utils.isolation_source_mapping_utils module
- kg_microbe.utils.lpsn_utils module
- kg_microbe.utils.mapping_file_utils module
MetpoTreeNodegenerate_assay_entity_edges()generate_assay_entity_nodes()generate_assay_nodes()load_assay_kit_mappings()load_metpo_enzyme_mappings()load_metpo_mappings()load_metpo_metabolite_production_mappings()load_metpo_metabolite_utilization_mappings()normalize_biolink_category()uri_to_curie()
- kg_microbe.utils.mediadive_bulk_download module
- kg_microbe.utils.metpo_liveness module
- kg_microbe.utils.metpo_predicates module
- kg_microbe.utils.microbial_trait_mappings module
- kg_microbe.utils.ner_utils module
- kg_microbe.utils.oak_utils module
- kg_microbe.utils.ontology_resolution module
- kg_microbe.utils.ontology_utils module
ChebiDbUnavailableErrorDbEnsureResultFatalOntologyErrorKeptTargetOntologyDbUnavailableErrorOntologyVersionMismatchErrorassert_chebi_version_alignment()assert_go_version_alignment()assert_ncbitaxon_version_alignment()get_chebi_adapter()get_chebi_category()get_ec_adapter()get_foodon_category()get_go_adapter()get_go_aspect()get_go_category_by_aspect()get_ncbitaxon_adapter()get_ncbitaxon_category()get_ontology_adapter()get_pato_category()get_uberon_category()ontology_db_path()replace_category_ontology()replace_deprecated_categories()resolve_adapter()
- kg_microbe.utils.optional_consumed_inputs module
- kg_microbe.utils.pandas_utils module
- kg_microbe.utils.parse_taxon_rank module
- kg_microbe.utils.postprocess_artifacts module
- kg_microbe.utils.provenance module
- kg_microbe.utils.review_claims module
- kg_microbe.utils.robot_utils module
- kg_microbe.utils.sanitize_curies module
- kg_microbe.utils.source_finalization module
- kg_microbe.utils.sssom_identity_policy module
- kg_microbe.utils.string_coding module
- kg_microbe.utils.stub_curie_collection module
- kg_microbe.utils.transform_fingerprint module
FINGERPRINT_FILEFINGERPRINT_VERSIONSCHEMA_FILESSHARED_CODEbounded_file_fingerprint()code_fingerprint()data_fingerprint()declared_data_inputs()finalization_inputs_current()migrate_markers()read_fingerprint()resolve_data_input()schema_fingerprint()shared_code_fingerprint()upstream_fingerprint()write_fingerprint()
- kg_microbe.utils.trembl_utils module
- kg_microbe.utils.tsv_io module
- kg_microbe.utils.unipathways_utils module
- kg_microbe.utils.uniprot_utils module
check_string_in_tar()convert_omim_diseases()create_pool()get_go_category_trees()get_go_relation_and_obsolete_terms()get_nodes_and_edges()go_category_trees_is_complete()is_float()parse_binding_site()parse_disease()parse_ec()parse_gene()parse_go_entry()parse_rhea_entry()prepare_go_dictionary()prepare_mondo_dictionary()process_lines()write_obsolete_file_header()
- kg_microbe.utils.validation_utils module
- Module contents
Submodules
kg_microbe.bactotraits_to_mongo module
BactoTraits CSV to MongoDB JSON converter.
Converts BactoTraits semicolon-separated CSV files with hierarchical headers into MongoDB-compatible JSON format with proper field nesting and one-hot encoding.
- kg_microbe.bactotraits_to_mongo.build_nested_dict(path_parts, value)
Build nested dictionary from path parts.
- Return type:
Dict
- kg_microbe.bactotraits_to_mongo.clean_field_name(name)
Convert field name to MongoDB-compatible format.
- Return type:
str
- kg_microbe.bactotraits_to_mongo.create_field_path(category, field_name)
Create hierarchical field path from category and field name.
- Return type:
str
- kg_microbe.bactotraits_to_mongo.forward_fill_headers(header_row)
Forward fill empty cells in header row.
- Return type:
List[str]
- kg_microbe.bactotraits_to_mongo.main()
Convert BactoTraits CSV to MongoDB JSON format.
- kg_microbe.bactotraits_to_mongo.merge_dicts(dict1, dict2)
Deep merge two dictionaries.
- Return type:
Dict
- kg_microbe.bactotraits_to_mongo.parse_bactotraits_to_mongo(input_file)
Parse BactoTraits CSV into MongoDB-compatible JSON. :rtype:
List[Dict[str,Any]]Uses hierarchical header structure (category + field name)
Creates nested paths using forward-filled categories
Only includes non-zero, non-NA, non-empty values
Handles one-hot encoding by preserving all non-zero values
Splits comma-separated values into arrays
- kg_microbe.bactotraits_to_mongo.parse_value(value)
Parse value, converting numbers and filtering out NA/empty.
- Return type:
Any
- kg_microbe.bactotraits_to_mongo.split_list_values(value)
Split comma-separated values into arrays where appropriate.
- Return type:
Any
kg_microbe.download module
Download resources from YAML file.
- exception kg_microbe.download.UnknownDownloadTagError
Bases:
ValueErrorRaised when a requested -t tag matches no entry in the download config.
A dedicated type so the CLI can report bad tags as usage errors without also swallowing unrelated ValueErrors raised deeper in the download (JSON decode failures and pydantic ValidationError are both ValueError subclasses, and were being presented as “Invalid value” for -t).
- kg_microbe.download.download(yaml_file, output_dir, snippet_only, ignore_cache=False, tags=None)
Download data files from list of URLs.
DL based on config (default: download.yaml) into data directory (default: data/).
- Parameters:
yaml_file (
str) – A string pointing to the yaml file
:param utilized to facilitate the downloading of data. :type output_dir:
str:param output_dir: A string pointing to the location to download data to. :type snippet_only:bool:param snippet_only: Downloads only the first 5 kB of the source,for testing and file checks. :type ignore_cache:bool:param ignore_cache: Ignore cache and download files even if they exist [false] :type tags:Optional[Sequence[str]] :param tags: Only download entries carrying one of these tags. None or emptymeans download everything.
- Return type:
None- Returns:
None.
kg_microbe.query module
Query module.
- kg_microbe.query.parse_query_yaml(yaml_file)
Parse a YAML file and return the results as a dictionary.
- Parameters:
yaml_file – YAML file to parse.
- Return type:
dict- Returns:
A dictionary of results from the YAML file.
- kg_microbe.query.result_dict_to_tsv(result_dict, outfile)
Write a dictionary to a TSV file.
- Parameters:
result_dict (
dict) – Dictionary to write to TSV file.outfile (
str) – TSV file to write to.
- Return type:
None
- kg_microbe.query.run_query(query, endpoint, return_format='json')
Run a SPARQL query and return the results as a dictionary.
- Parameters:
query (
str) – SPARQL query to run.endpoint (
str) – SPARQL endpoint to query.return_format – Format of the returned data.
- Return type:
dict- Returns:
A dictionary of results from the SPARQL query.
kg_microbe.run module
Drive KG download, transform, merge steps.
kg_microbe.transform module
Transform module.
- class kg_microbe.transform.LazyTransform(dotted_path)
Bases:
objectResolve one transform class only when it is used.
- property transform_class
Import and return the registered transform class.
- exception kg_microbe.transform.TransformBatchError(failed, skipped)
Bases:
RuntimeErrorOne or more sources in a batch failed; raised after the others ran.
Carries
failed(source -> exception) andskipped(source -> the upstream sources it declares inTRANSFORM_INPUTSthat failed first), and renders both, so the exit is non-zero and the summary is unmissable.
- kg_microbe.transform.transform(input_dir, output_dir, sources=None, show_status=True)
Transform based on resource and class declared in DATA_SOURCES.
Call scripts in kg_microbe/transform/[source name]/ to transform each source into a graph format that KGX can ingest directly, in either TSV or JSON format: https://github.com/biolink/kgx/blob/master/data-preparation.md
- Parameters:
input_dir (
Optional[Path]) – A string pointing to the directory to import data from.output_dir (
Optional[Path]) – A string pointing to the directory to output data to.sources (
Optional[List[str]]) – A list of sources to transform. A registered source name, or one ontology name fromONTOLOGIES_MAP(ec,chebi, …) to refresh that ontology alone (#690).
- Raises:
ValueError – If a requested source is not registered in DATA_SOURCES.
FileNotFoundError – If a selected source declares a curation input that is not on disk; nothing runs (#685).
TransformBatchError – After the batch, if any source failed. A failure is isolated to its source: later sources still run unless they declare the failed one in
TRANSFORM_INPUTS, in which case they are skipped rather than built on stale upstream output (#685).BaseException(FatalOntologyError, Ctrl-C) still aborts at once.
- Return type:
None
Module contents
kg-microbe package.
- kg_microbe.download(yaml_file, output_dir, snippet_only, ignore_cache=False, tags=None)
Download data files from list of URLs.
DL based on config (default: download.yaml) into data directory (default: data/).
- Parameters:
yaml_file (
str) – A string pointing to the yaml file
:param utilized to facilitate the downloading of data. :type output_dir:
str:param output_dir: A string pointing to the location to download data to. :type snippet_only:bool:param snippet_only: Downloads only the first 5 kB of the source,for testing and file checks. :type ignore_cache:bool:param ignore_cache: Ignore cache and download files even if they exist [false] :type tags:Optional[Sequence[str]] :param tags: Only download entries carrying one of these tags. None or emptymeans download everything.
- Return type:
None- Returns:
None.