kg_microbe.utils package
Submodules
kg_microbe.utils.atomic_io module
Atomic writes for derived caches, so a failed run leaves no half-written file.
- kg_microbe.utils.atomic_io.atomic_write(path, mode='w', mark_complete=False, **open_kwargs)
Write to
pathvia a temp file that is renamed into place only on success.Several derived caches in this repo are generated once and then guarded by a bare
path.exists()on every later run. Writing them in place makes any mid-write failure permanent: the header lands, the generator raises, the context manager closes a truncated file, and every subsequent run sees a file that exists and skips regeneration.go_category_trees.tsvis the case that bit us — a header-only file makesprepare_go_dictionaryreturn{}, which drops every protein→GO edge and logs every GO term as obsolete, with a zero exit code, forever.os.replaceis atomic on POSIX and Windows, so a reader either sees the old file or the complete new one, never a partial. The temp file is removed on any unwind — includingBaseException, which is what the fatal ontology errors are — because the cleanup lives infinally.Callers that need the old file left intact on failure get that for free: it is never touched until the rename.
- Parameters:
path (
Union[str,Path]) – Final destination path.mode (
str) – Mode for the underlyingopen; must be a writing mode.open_kwargs – Passed through to
open(encoding, newline, …).
- Yield:
The open file handle to write to.
- kg_microbe.utils.atomic_io.cache_is_complete(path, delimiter='\\t')
Report whether a derived cache was completely produced.
Complete means either an explicit completion marker written by
atomic_write()(mark_complete=True), or — for caches written before markers existed — at least one data row. The marker is what lets a legitimately empty result count as finished: judging on row count alone, a run that correctly produced zero annotations looked identical to a truncated one and was regenerated on every subsequent run forever.- Parameters:
path (
Union[str,Path]) – Cache path.delimiter (
str) – Field delimiter of the cache.
- Return type:
bool- Returns:
True if the cache exists and is known to be complete.
- kg_microbe.utils.atomic_io.has_data_rows(path, delimiter='\\t')
Report whether a delimited cache exists and holds at least one data row.
The companion to
atomic_write(). Atomicity stops a cache from being poisoned again, but it cannot repair one poisoned before the fix landed — and every consumer of these caches guards regeneration with a bare.exists(), so a header-only file is accepted forever. Checking for content is what lets an already-poisoned cache heal on the next run instead of requiring the user to know which file to delete.Conservative on error: an unreadable file reports False, so the caller regenerates rather than trusting something it could not inspect.
- Parameters:
path (
Union[str,Path]) – Cache path.delimiter (
str) – Field delimiter of the cache.
- Return type:
bool- Returns:
True if the file exists and has a row beyond the header.
kg_microbe.utils.biolink_hierarchy module
Biolink Model hierarchy utility for category specificity comparison.
- class kg_microbe.utils.biolink_hierarchy.BiolinkHierarchy(biolink_yaml_path=None)
Bases:
objectUtility for Biolink Model category hierarchy operations.
Loads biolink-model.yaml and provides methods to: - Determine most specific category from a list - Check if category A is more specific than category B - Traverse hierarchy to find ancestors
- Example usage:
>>> hierarchy = BiolinkHierarchy() >>> categories = ["biolink:ChemicalEntity", "biolink:SmallMolecule"] >>> result = hierarchy.get_most_specific_category(categories) >>> print(result) biolink:SmallMolecule
- get_ancestors(category)
Get all ancestor categories up to NamedThing root.
- Return type:
List[str]
- Args:
category: Category string (can include or omit “biolink:” prefix)
- Returns:
List of ancestor categories with “biolink:” prefix, ordered from immediate parent to root (NamedThing)
- Example:
>>> hierarchy = BiolinkHierarchy() >>> hierarchy.get_ancestors("biolink:SmallMolecule") ['biolink:ChemicalEntity', 'biolink:NamedThing']
- get_depth(category)
Get the depth of a category in the hierarchy.
- Return type:
Optional[int]
- Args:
category: Category string (can include or omit “biolink:” prefix)
- Returns:
Depth as integer (0 = NamedThing root), or None if category not found
- Example:
>>> hierarchy = BiolinkHierarchy() >>> hierarchy.get_depth("biolink:NamedThing") 0 >>> hierarchy.get_depth("biolink:ChemicalEntity") 2 >>> hierarchy.get_depth("biolink:SmallMolecule") 3
- get_most_specific_category(categories)
Return the most specific (deepest in hierarchy) category.
- Return type:
str
- Args:
- categories: List of category strings (e.g., [“ChemicalEntity”, “SmallMolecule”])
Can include or omit “biolink:” prefix
- Returns:
Most specific category string with “biolink:” prefix
- Example:
>>> hierarchy = BiolinkHierarchy() >>> hierarchy.get_most_specific_category(["biolink:ChemicalEntity", "biolink:SmallMolecule"]) 'biolink:SmallMolecule'
- is_more_specific(category_a, category_b)
Check if category_a is more specific than category_b.
Returns True if category_a is a descendant of category_b (i.e., category_a has greater depth in the hierarchy).
- Return type:
bool
- Args:
category_a: First category (can include or omit “biolink:” prefix) category_b: Second category (can include or omit “biolink:” prefix)
- Returns:
True if category_a is more specific (deeper) than category_b
- Example:
>>> hierarchy = BiolinkHierarchy() >>> hierarchy.is_more_specific("biolink:SmallMolecule", "biolink:ChemicalEntity") True
kg_microbe.utils.biolink_model module
Configure KGX/BMT to use KG-Microbe’s pinned local Biolink model.
- kg_microbe.utils.biolink_model.prepare_kgx()
Install a local-default BMT Toolkit before importing KGX.
KGX 2.x creates
Toolkit()at module import time. BMT’s defaults are remote URLs for both its schema and predicate map, so importing KGX can otherwise perform network I/O. Replacing BMT’s exported Toolkit class before KGX imports preserves explicit caller arguments while redirecting its no-argument construction to pinned repository files.- Return type:
None
- kg_microbe.utils.biolink_model.sibling_imports(schema_path)
Return the files
schema_pathimports from its own directory.biolink-model.yamldeclaresimports: [linkml:types, attributes]. A prefixed name such aslinkml:typesis resolved from the installed linkml runtime, but a bare one is a sibling file that must be downloaded alongside the model – andattributes.yamlnever was, so the pinned model raisedFileNotFoundErrorfrom deep inside linkml on every load (#939). Reporting it here names the missing file and the fix instead.The list is read with a real YAML parser rather than a regex over the
imports:block. A regex handles today’s block style and silently returns nothing for the equally valid flow style, which would restore the #939 failure with no signal that the guard had stopped working – a silent degradation in the guard against silent degradation (#942). The parse costs about 50 ms withCSafeLoaderagainst the 264 ms BMT already spends building aSchemaViewfrom the same file, and runs once per pipeline invocation, not once per record.- Parameters:
schema_path (
Path) – Path to the local Biolink model YAML.- Return type:
list[Path]- Returns:
Paths the model expects beside itself, whether or not they exist.
kg_microbe.utils.cas module
Structural CAS registry-number admission, without inferring chemical identity.
- kg_microbe.utils.cas.invalid_cas_identifier(value, *, allow_bare=False)
Reject recognizable invalid CAS inputs; leave other namespaces/names alone.
Case and the legacy CAS-RN prefix are representation aliases only. Bare registry-shaped names are checked only at lexical lookup boundaries. No replacement digit, substance, or equivalence is ever supplied.
- Return type:
bool
- kg_microbe.utils.cas.valid_cas(identifier)
Validate a canonical CAS CURIE’s syntax and checksum, never repair it.
- Return type:
bool
kg_microbe.utils.chemical_mapping_utils module
Utilities for chemical mapping lookups.
Reads the unified ingredient SSSOM mapping set
(mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz) and reconstructs an
entity-centric in-memory index grouped on object_id. The SSSOM carries
the per-entity attributes (canonical_name via object_label,
formula/category via extension columns) as well as the mappings
themselves (xrefs, canonical-name rows, synonym rows).
- class kg_microbe.utils.chemical_mapping_utils.ChemicalMappingLoader(mappings_path=None, *, ingredient_bundle=None, mappings_sha256=None)
Bases:
objectLoader class for unified chemical mappings.
Provides convenient API for chemical entity lookups. Uses module-level caching internally.
- find_chebi_by_formula(formula)
Lookup ChEBI IDs by molecular formula.
- Parameters:
formula (
str) – Molecular formula- Return type:
List[str]- Returns:
List of ChEBI IDs
- find_chebi_by_name(name, synonyms=True, fuzzy_stereochemistry=False, fuzzy_hydrate=False)
Lookup ChEBI ID by chemical name.
- Parameters:
name (
str) – Chemical name to search forsynonyms (
bool) – If True, search both canonical names and synonymsfuzzy_stereochemistry (
bool) – If True, retry with stereochemistry prefixes removedfuzzy_hydrate (
bool) – Compatibility argument; never strips hydration to infer identity
- Return type:
Optional[str]- Returns:
ChEBI ID or None if not found
- find_chebi_by_xref(xref)
Lookup ChEBI ID by cross-reference.
- Parameters:
xref (
str) – Cross-reference identifier- Return type:
Optional[str]- Returns:
ChEBI ID or None if not found
- classmethod from_ingredient_lookup_bundle(directory, *, manifest_sha256)
Load a pinned candidate with separate legacy and reviewed scoped inputs.
- get_canonical_name(chebi_id)
Get canonical ChEBI name.
- Parameters:
chebi_id (
str) – ChEBI ID- Return type:
Optional[str]- Returns:
Canonical name or None if not found
- get_category(curie)
Get the biolink category recorded for a CURIE.
- Parameters:
curie (
str) – Primary CURIE.- Return type:
Optional[str]- Returns:
Biolink category string or None.
- get_formula(chebi_id)
Get molecular formula for ChEBI ID.
- Parameters:
chebi_id (
str) – ChEBI ID- Return type:
Optional[str]- Returns:
Molecular formula or None if not found
- get_identifier_annotation_owners(identifier, *, include_history=False)
Query all owners of registry annotations without adding identity aliases.
- Return type:
List[str]
- get_node_enrichment(curie)
Get KGX node enrichment fields (xref, synonym, name) for a CURIE.
- Parameters:
curie (
str) – Primary CURIE.- Return type:
Dict[str,str]- Returns:
Dict with
xref,synonym,namekeys.
- get_parents(curie)
Get parent CURIEs (asymmetric narrowMatch) for an entity.
- Parameters:
curie (
str) – child CURIE- Return type:
List[str]- Returns:
List of parent CURIEs (empty if none recorded)
- get_synonyms(chebi_id)
Get all synonyms for ChEBI ID.
- Parameters:
chebi_id (
str) – ChEBI ID- Return type:
List[str]- Returns:
List of synonyms
- get_xrefs(chebi_id)
Get all cross-references for ChEBI ID.
- Parameters:
chebi_id (
str) – ChEBI ID- Return type:
List[str]- Returns:
List of xrefs
- resolve_ingredient_source(source_id, *, occurrence_id=None)
Apply the explicit scoped source concept before a context-free name lookup.
- Return type:
Optional[str]
- kg_microbe.utils.chemical_mapping_utils.PREDICATE_SEMANTICS_KEY = 'predicate_semantics'
Header key by which a mapping set declares which semantics its asymmetric predicates follow. Proposed on #822 / CultureBotAI/MediaIngredientMech#390. The name is MIM’s to settle; it is a constant here so a rename is one line.
- kg_microbe.utils.chemical_mapping_utils.SKOS_SEMANTICS = 'skos'
Value meaning “read skos:narrowMatch / skos:broadMatch per the SKOS spec”.
- kg_microbe.utils.chemical_mapping_utils.find_chebi_by_formula(formula)
Lookup ChEBI IDs by molecular formula.
Note: May return multiple ChEBI IDs if formula matches multiple compounds.
- Parameters:
formula (
str) – Molecular formula (e.g., “H2O”, “C6H12O6”)- Return type:
List[str]- Returns:
List of ChEBI IDs (may be empty if not found)
- kg_microbe.utils.chemical_mapping_utils.find_chebi_by_name(name, synonyms=True, fuzzy_stereochemistry=False, fuzzy_hydrate=False)
Lookup entry CURIE by ingredient name.
The unified mapping is keyed on a generic CURIE (id), so this function returns any supported ontology CURIE — CHEBI for chemicals, or FOODON/UBERON/ENVO for food/anatomy/environment ingredients — when the name matches. The legacy function name is retained for API stability.
- Parameters:
name (
str) – Ingredient name to search forsynonyms (
bool) – If True, search both canonical names and synonyms If False, only search canonical namesfuzzy_stereochemistry (
bool) – If True and exact match fails, retry with stereochemistry prefixes removedfuzzy_hydrate (
bool) – Retained for call compatibility; hydration stripping no longer establishes identity. Use get_hydrate_equivalents for separately asserted weak recipe relations, not ingredient substitution.
- Return type:
Optional[str]- Returns:
CURIE (e.g., “CHEBI:12345”, “FOODON:00002441”) or None if not found
- kg_microbe.utils.chemical_mapping_utils.find_chebi_by_xref(xref)
Lookup ChEBI ID by cross-reference (KEGG, CAS, etc.).
- Parameters:
xref (
str) – Cross-reference identifier (e.g., “cas:50-00-0”, “kegg.compound:C00001”)- Return type:
Optional[str]- Returns:
ChEBI ID (e.g., “CHEBI:12345”) or None if not found
- kg_microbe.utils.chemical_mapping_utils.get_canonical_name(chebi_id)
Get canonical name for a given CURIE.
- Parameters:
chebi_id (
str) – Primary CURIE (e.g., “CHEBI:12345”, “FOODON:00002441”)- Return type:
Optional[str]- Returns:
Canonical name or None if not found
- kg_microbe.utils.chemical_mapping_utils.get_category(curie)
Get the biolink category stored for a CURIE in the unified mapping.
Category is a data column on each row of the unified mapping — there is no prefix-to-category table in code. Returns None if the CURIE is not present or has no category recorded.
- Parameters:
curie (
str) – Primary CURIE (e.g. “CHEBI:15377”, “FOODON:00002441”)- Return type:
Optional[str]- Returns:
Biolink category (e.g. “biolink:ChemicalEntity”, “biolink:Food”) or None if not found.
- kg_microbe.utils.chemical_mapping_utils.get_formula(chebi_id)
Get molecular formula for a given CURIE.
- Parameters:
chebi_id (
str) – Primary CURIE (e.g., “CHEBI:12345”)- Return type:
Optional[str]- Returns:
Molecular formula or None if not found
- kg_microbe.utils.chemical_mapping_utils.get_hydrate_equivalents(curie)
Return CHEBI CURIEs that are media-recipe interchangeable with
curie.Anhydrous and hydrated forms of a salt (e.g. CaCl2 ↔ CaCl2·2H2O) are DIFFERENT chemical entities – different formula, different molecular weight – but routinely substituted for each other in microbial growth media with a concentration adjustment. This accessor returns the other-form CURIE(s) for recipe-comparison purposes only; it must NOT be used to claim chemical identity (use
get_xrefsfor that, which assertsskos:exactMatch).The relationship is symmetric: looking up either form returns the other. Sourced from rows whose
predicate_id == 'skos:closeMatch'andcomment == 'recipe_equivalent_hydrate'in the unified SSSOM, emitted by the consolidator’s hydrate logic.- Parameters:
curie (
str) – CHEBI CURIE for either the anhydrous or hydrated form- Return type:
List[str]- Returns:
sorted list of recipe-equivalent CURIEs (empty if none)
- kg_microbe.utils.chemical_mapping_utils.get_mapping_load_audit()
Return counts and bounded row diagnostics from the last mapping load.
- Return type:
dict
- kg_microbe.utils.chemical_mapping_utils.get_node_enrichment(curie, *, ingredient_bundle=None)
Return KGX enrichment fields for a chemical/ingredient CURIE.
Returns a dict with keys
xref,synonym,namesuitable for populating the corresponding KGX node columns. Values are pipe-joined strings (KGX multivalued convention) or empty strings when absent.xref: unified cross-references plus independently reviewed CAS annotations. CAS annotations do not create SSSOM exactMatch or same_as.synonym: alternative free-text names from thesynonymscolumn.name: canonical name for the CURIE, or empty string when unknown.
- Parameters:
curie (
str) – Primary CURIE (e.g. “CHEBI:15377”, “FOODON:00002441”).- Return type:
Dict[str,str]- Returns:
Dict with
xref,synonym,namekeys.
- kg_microbe.utils.chemical_mapping_utils.get_parents(curie)
Get parent CURIEs for an entity (asymmetric narrowMatch / broadMatch).
Returns the list of broader / parent CURIEs that
curieis a kind-of. Populated from MIMskos:narrowMatchrows in the unified SSSOM (where e.g.MIM:Vermont_Soil → ENVO:00001998 narrowMatchtranslates tokgmicrobe.ingredient:vermont_soilhaving parentENVO:00001998).Transforms call this when emitting an ingredient / solution / sample edge to also write a
biolink:broad_matchedge to the broader OBO term, so consumers can navigate from kg-microbe-minted CURIEs back to the canonical hierarchy. It is notbiolink:subclass_of: MIM’s parent anchor means “closest broader term” and covers salts, hydrates and solutions of the parent, which ChEBI relates by has-part, not is_a (MIM MAPPING_SEMANTICS.md Section 1, #245).- Parameters:
curie (
str) – child CURIE- Return type:
List[str]- Returns:
sorted list of parent CURIEs (empty if no narrowMatch row exists)
- kg_microbe.utils.chemical_mapping_utils.get_synonyms(chebi_id)
Get all synonyms for a given CURIE.
- Parameters:
chebi_id (
str) – Primary CURIE (e.g., “CHEBI:12345”)- Return type:
List[str]- Returns:
List of synonyms (may be empty)
- kg_microbe.utils.chemical_mapping_utils.get_xrefs(chebi_id)
Get cross-reference annotations, without changing literal identity lookup keys.
The historical doubled KEGG compound prefix is presentation-only (#1259). Keep original SSSOM rows and the independently keyed xref lookup index intact: canonical and alias keys can have different historical mapping winners.
- Parameters:
chebi_id (
str) – Primary CURIE (e.g., “CHEBI:12345”)- Return type:
List[str]- Returns:
List of xrefs (e.g., [“cas:50-00-0”, “kegg.compound:C00001”])
- kg_microbe.utils.chemical_mapping_utils.load_unified_mappings(mappings_path=None, *, expected_sha256=None)
Load the unified ingredient SSSOM mapping set.
Reads
mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz(or the explicit path given) and builds the in-memory per-entity indices used by the lookup API (find_chebi_by_name,get_canonical_name, …). Per-entity attributes that SSSOM cannot express natively —object_formulaandobject_category— are read from the SSSOM extension columns emitted byscripts/consolidate_chemical_mappings.py.- Row-shape semantics (matches
export_unified_sssom): kgm.name:subject +comment == "canonical_name"→ contributes the canonical name viasubject_label/object_label.kgm.name:subject +comment == "synonym"→subject_labelis added as an entity synonym.other CURIE subject (not equal to object) with exactMatch → identity xref.
subject == object (attribute_carrier) → no-op mapping; used only to carry extension columns for entities with no other rows.
Uses module-level caching to avoid reloading on multiple calls.
- Parameters:
mappings_path (
Optional[Path]) – Path to the unified SSSOM. If None, uses the default path relative to this file.- Return type:
int- Returns:
Number of distinct entities loaded (zero before first load).
- Row-shape semantics (matches
- kg_microbe.utils.chemical_mapping_utils.normalize_chemical_primes(name)
Canonicalize typographic prime marks without losing chemical locants.
- Return type:
str
- kg_microbe.utils.chemical_mapping_utils.normalize_name(name, strip_stereochemistry=False, strip_hydrate=False)
Normalize chemical name for comparison.
- Parameters:
name (
str) – Chemical name to normalizestrip_stereochemistry (
bool) – If True, remove stereochemistry prefixes like (R)-, (S)-, D-, L-, (+)-, (-)-strip_hydrate (
bool) – If True, strip trailing hydrate suffixes like “ x n H2O”, “ · 6 H2O”, “ . 2H2O”
- Return type:
str- Returns:
Normalized name retaining hyphens and chemical prime locants
- kg_microbe.utils.chemical_mapping_utils.read_predicate_semantics(path)
Read a mapping set’s declared predicate semantics from its header.
Absence means legacy, and so does anything unrecognised. MIM writes
skos:narrowMatchto mean “the object is the parent”, which is the inverse of the SKOS spec; this repo has always read it that way, so the two agree and every asymmetric row yields a correctly directed parent edge (#822).Failing closed is the whole point. It makes the two halves of the fix order-independent: an old file, or a rebuild of old content, carries no declaration and is read as legacy however new the reader is. A date threshold could not do this — MIM’s
mapping_set_versionis build time, regenerated on every build, so rebuilding unfixed content would stamp a post-cutover date onto legacy rows and silently invert 141 edges.- Parameters:
path (
Path) – SSSOM file, plain or gzipped.- Return type:
str- Returns:
The declared value, or
""when absent or unreadable.
kg_microbe.utils.consolidate_categories module
Consolidate multi-category nodes in merged knowledge graph.
This script resolves nodes with pipe-delimited categories (e.g., “biolink:ChemicalEntity|biolink:SmallMolecule”) using explicit, verified consolidation rules that account for semantic meaning and Biolink Model hierarchy.
- Usage:
python kg_microbe/utils/consolidate_categories.py –input data/merged/merged-kg_nodes.tsv –output data/merged/merged-kg_nodes_consolidated.tsv –report data/merged/category_consolidation_report.txt
- kg_microbe.utils.consolidate_categories.consolidate_node_categories(input_file, output_file, report_file=None, rules_file=None)
Read merged nodes TSV and consolidate multi-category nodes.
For nodes with pipe-delimited categories, uses explicit verified rules that account for semantic meaning (e.g., ChemicalRole for functional classifications vs SmallMolecule for structural entities).
- Return type:
None
- Args:
input_file: Path to input nodes TSV file (with multi-category nodes) output_file: Path to output consolidated nodes TSV file (single categories) report_file: Optional path to consolidation report file rules_file: Optional path to rules YAML file (defaults to package location)
- Example:
>>> consolidate_node_categories( ... "data/merged/merged-kg_nodes.tsv", ... "data/merged/merged-kg_nodes_consolidated.tsv", ... "data/merged/category_consolidation_report.txt" ... ) ✓ Consolidated 1,117 multi-category nodes ✓ Output written to: data/merged/merged-kg_nodes_consolidated.tsv ✓ Report written to: data/merged/category_consolidation_report.txt
- kg_microbe.utils.consolidate_categories.main()
Parse command-line arguments and run consolidation.
kg_microbe.utils.download_bacdive module
kg_microbe.utils.download_manifest module
Record which URL produced each downloaded artifact, and re-fetch when it changes.
kghub_downloader decides whether to download by asking whether the file
exists, and stores nothing about where it came from. So editing a pinned version
in download.yaml changes nothing on disk: the run reports success and the
pipeline keeps consuming the old artifact, with the declared pin and the cached
bytes silently disagreeing. That is how a METPO pin moved three releases while
data/raw/metpo.json stayed on the version before it (#900, #911).
Existence is the wrong cache key. This module makes the URL part of it: a file
whose recorded URL no longer matches the declared one is removed before the
download runs, so the next ordinary kg download picks up the change.
A file with no record is left alone. Every artifact predates this manifest, so treating “unknown” as “stale” would re-fetch the entire corpus on the first run – including an hour of MediaDive crawling – to learn what is already known. Records accumulate as things are downloaded.
- kg_microbe.utils.download_manifest.MANIFEST_NAME = '.download_manifest.json'
Sits beside the artifacts it describes, under the gitignored data directory.
- kg_microbe.utils.download_manifest.drop_stale(yaml_file, output_dir, tags=None)
Remove artifacts whose declared URL has changed, so they are fetched again.
- Parameters:
yaml_file (
str) – Download config being applied.output_dir (
str) – Directory downloads are written to.tags (
Optional[Sequence[str]]) – Tags this run is restricted to, or None for all.
- Return type:
List[Path]- Returns:
Paths removed.
- kg_microbe.utils.download_manifest.manifest_path(output_dir)
Return the manifest location for a download directory.
- Parameters:
output_dir (
str) – Directory downloads are written to.- Return type:
Path- Returns:
Path to the manifest, which may not exist.
- kg_microbe.utils.download_manifest.read_manifest(output_dir)
Return the recorded
local_name -> urlmap.- Parameters:
output_dir (
str) – Directory downloads are written to.- Return type:
Dict[str,str]- Returns:
Mapping, empty when absent or unreadable. A corrupt manifest is treated as no manifest: it must never be able to block a download.
- kg_microbe.utils.download_manifest.record(yaml_file, output_dir, tags=None, skip_names=None)
Record the URL behind every artifact this run left on disk.
Only files that exist are recorded, so a skipped or failed entry does not gain a provenance claim it has not earned.
skip_namesexcludes entries the caller knows were never downloaded – pending-hosting sources are satisfied by hand, so their artifact exists without having come from the placeholder URL in the config (#929).- Parameters:
yaml_file (
str) – Download config that was applied.output_dir (
str) – Directory downloads were written to.tags (
Optional[Sequence[str]]) – Tags this run was restricted to, or None for all.skip_names (
Optional[Sequence[str]]) – Local names to leave unrecorded.
- Return type:
int- Returns:
Number of artifacts recorded.
- kg_microbe.utils.download_manifest.stale_by_url(yaml_file, output_dir, tags=None)
Return downloaded files whose recorded URL no longer matches the config.
- Parameters:
yaml_file (
str) – Download config being applied.output_dir (
str) – Directory downloads are written to.tags (
Optional[Sequence[str]]) – Tags this run is restricted to, or None for all.
- Return type:
List[Path]- Returns:
Paths that exist on disk and came from a different URL.
kg_microbe.utils.download_utils module
Download utilities for KG-Microbe data sources.
This module provides functions for downloading files from various sources including YAML-configured downloads, API-based downloads, and mirroring to cloud storage buckets.
- kg_microbe.utils.download_utils.check_for_file_existence_in_batch(batch, outyamls, empty_orgs)
Check which organisms in a batch already have downloaded files.
Args:
batch: List of organism IDs to check. outyamls: Output directory path where files are stored. empty_orgs: List of organism IDs known to have no data.
Returns:
Updated batch list with existing organisms removed.
- kg_microbe.utils.download_utils.download_from_api(yaml_item, outfile)
Download data from an API based on YAML configuration.
- Return type:
None
Args:
yaml_item: Item to be downloaded, parsed from YAML. outfile: Path where to write out the downloaded file.
- kg_microbe.utils.download_utils.download_from_yaml(yaml_file, output_dir, ignore_cache=False, snippet_only=False, tags=None, mirror=None)
Download files listed in a download.yaml file.
- Return type:
None
Args:
yaml_file: A string pointing to the download.yaml file, to be parsed for things to download. output_dir: A string pointing to where to write out downloaded files. ignore_cache: Ignore cache and download files even if they exist [false]. snippet_only: Downloads only the first 5 kB of each uncompressed source, for testing and file checks. tags: Limit to only downloads with this tag. mirror: Optional remote storage URL to mirror download to. Supported buckets: Google Cloud Storage.
Returns:
None.
- kg_microbe.utils.download_utils.elastic_search_query(es_connection, index, query, scroll='1m', request_timeout=60, preserve_order=True)
Fetch records from Elasticsearch using the given query parameters.
Args:
es_connection: Elasticsearch connection. index: The Elasticsearch index for query. query: Query to execute. scroll: Scroll parameter passed to Elasticsearch. request_timeout: Timeout parameter passed to Elasticsearch. preserve_order: Preserve order parameter passed to Elasticsearch.
Returns:
All records for query.
- kg_microbe.utils.download_utils.get_jobs(url, values)
Fetch all paginated results from a UniProt query URL.
Args:
url: The UniProt query URL. values: List to append results to.
- kg_microbe.utils.download_utils.get_uniprot_values_organism(organism_ids, base_url, fields, keywords, size, batch_size, outyamls)
Query UniProt for enzyme data per organism in batches.
Args:
organism_ids: List of NCBI Taxon organism IDs to query. base_url: Base URL for the UniProt API. fields: List of fields to retrieve from UniProt. keywords: List of keywords to filter results. size: Maximum number of results per query. batch_size: Number of organisms to query in each batch. outyamls: Output directory path for storing results.
- kg_microbe.utils.download_utils.mirror_to_bucket(local_file, bucket_url, remote_file)
Mirror a local file to a remote cloud storage bucket.
- Return type:
None
Args:
local_file: Path to the local file to upload. bucket_url: URL of the remote storage bucket (gs:// for Google Cloud Storage). remote_file: Name for the file in the remote bucket.
- kg_microbe.utils.download_utils.parse_ncbitaxon_json(input_file)
Parse NCBITaxon organism IDs from a JSON file.
Args:
input_file: Path to the JSON file containing NCBITaxon nodes.
Returns:
List of organism IDs extracted from the NCBITaxon nodes.
- kg_microbe.utils.download_utils.parse_response(res, values)
Parse a TSV response from UniProt and append to values list.
Args:
res: Response object from requests. values: List to append parsed records to.
Returns:
Updated values list.
- kg_microbe.utils.download_utils.parse_url(url)
Parse a URL for any environment variables enclosed in {curly braces}.
kg_microbe.utils.dummy_tqdm module
Define a dummy context manager to use when tqdm is disabled.
- class kg_microbe.utils.dummy_tqdm.DummyTqdm(*args, **kwargs)
Bases:
objectA dummy context manager that provides a no-operation replacement for tqdm.
This class is intended to be used as a drop-in replacement for tqdm progress bars when the display of progress is not needed. It implements the same methods as tqdm, but each method performs no action.
- Parameters:
*args –
Arbitrary positional arguments.
**kwargs –
Arbitrary keyword arguments.
- set_description(desc=None)
No-op implementation of the tqdm set_description method.
Sets the description of the progress bar, but here it does nothing.
- Parameters:
desc (str, optional) – The description text for the progress bar.
- update(n=1)
No-op implementation of the tqdm update method.
Intended to be called with an increment, but ignores it.
- Parameters:
n (int, optional) – The increment by which to update the progress (default is 1).
kg_microbe.utils.external_identifiers module
Exact external identifier retirement and assembly alias evidence.
- kg_microbe.utils.external_identifiers.load_assembly_aliases(nodes_file)
Return declared assemblies and uniquely proven, version-exact GTDB same_as aliases.
Never infer a GenBank/RefSeq pair from a shared numeric accession. Both accessions come from the GTDB row, and accession versions remain intact.
- Return type:
tuple
- kg_microbe.utils.external_identifiers.load_taxid_merges(raw_dir)
Read NCBI merged.dmp proof; absence means no locally available retirement evidence.
- Return type:
Dict[str,str]
- kg_microbe.utils.external_identifiers.resolve_replacement(identifier, replacements)
Follow only explicitly supplied replacements, rejecting cyclic evidence.
- Return type:
str
kg_microbe.utils.fix_list_representations module
Fix Python list representations in KGX merged output.
KGX sometimes writes list-valued fields as Python list representations (e.g., “[‘infores:chebi’, ‘infores:chebi’]”) instead of pipe-delimited strings (e.g., “infores:chebi|infores:chebi”). This utility converts them to proper format.
- Usage:
python kg_microbe/utils/fix_list_representations.py –input data/merged/merged-kg_edges.tsv –output data/merged/merged-kg_edges_fixed.tsv
- kg_microbe.utils.fix_list_representations.fix_list_field(value, delimiter='|', deduplicate=True)
Fix Python list representation to pipe-delimited string.
- Return type:
str
- Args:
value: String that may be a Python list representation delimiter: Output delimiter (default: pipe) deduplicate: Remove duplicates while preserving order (default: True)
- Returns:
Properly formatted pipe-delimited string
- Example:
>>> fix_list_field("['infores:chebi', 'infores:chebi']") 'infores:chebi' >>> fix_list_field("['infores:madin_etal', 'infores:bactotraits']") 'infores:madin_etal|infores:bactotraits' >>> fix_list_field("infores:chebi") 'infores:chebi'
- kg_microbe.utils.fix_list_representations.fix_tsv(input_file, output_file, delimiter='\\t', list_delimiter='|', deduplicate=True)
Fix Python list representations in TSV file.
- Return type:
None
- Args:
input_file: Path to input TSV file output_file: Path to output fixed TSV file delimiter: TSV field delimiter (default: tab) list_delimiter: List value delimiter (default: pipe) deduplicate: Remove duplicates while preserving order (default: True)
- kg_microbe.utils.fix_list_representations.main()
Parse command-line arguments and run fixes.
- kg_microbe.utils.fix_list_representations.parse_list_representation(value)
Parse Python list representation to actual list.
- Return type:
List[str]
- Args:
value: String that may be a Python list representation
- Returns:
List of strings if parseable, else [value]
- Example:
>>> parse_list_representation("['infores:chebi', 'infores:chebi']") ['infores:chebi', 'infores:chebi'] >>> parse_list_representation("infores:chebi") ['infores:chebi']
kg_microbe.utils.foodon_classification module
Asserted FOODON organism-class boundaries, distinct from foods and body parts.
- kg_microbe.utils.foodon_classification.authoritative_foodon_category(identifier, *, path=None)
Distinguish FOODON organism classes from the default food-material category.
- Return type:
str
- kg_microbe.utils.foodon_classification.foodon_dispositions()
Read evidence-bearing curation; pipeline data fingerprints separately hash its content.
- Return type:
dict
- kg_microbe.utils.foodon_classification.foodon_organism_classes(path=None)
Read local asserted ancestry when present; never fetch ontology data implicitly.
- Return type:
frozenset
- kg_microbe.utils.foodon_classification.organism_classes(graphs)
Follow asserted subclass edges downward from documented whole-organism roots.
- Return type:
frozenset
kg_microbe.utils.graph_canonicalization module
Source-independent KGX category and identifier normalization (#1052, #1054).
- kg_microbe.utils.graph_canonicalization.canonical_node_category(identifier, category, *, foodon_path=None)
Replace imported OntologyClass fallbacks without discarding substantive multi-typing.
- Return type:
str
- kg_microbe.utils.graph_canonicalization.compact_identifier(identifier)
Compact registered identifiers without inventing prefixes for arbitrary URLs.
- Return type:
str
kg_microbe.utils.graph_schema module
One published KGX TSV schema shared by source finalization and merge serialization.
- kg_microbe.utils.graph_schema.canonical_header(header, is_node)
Order required and present optional columns, retaining all extensions deterministically.
- kg_microbe.utils.graph_schema.validate_canonical_tsv(path, *, is_node)
Stream a literal LF TSV contract without changing headers, values, or provenance.
kg_microbe.utils.ingredient_bundle module
Consume a pinned reviewed ingredient cohort without deriving identity from xrefs.
- class kg_microbe.utils.ingredient_bundle.ReviewedIngredientBundle(directory, *, manifest_sha256)
Bases:
objectHold one immutable, verified snapshot with explicit source-context lookup.
- active_xrefs(owner_id)
Return reviewed current references for this owner, never equivalence candidates.
- Return type:
list[str]
- annotation_document()
Project owners alongside intact claims and resolvable bundled evidence references.
- Return type:
dict
- audit()
Account for every input claim without silently consuming deferred products/history.
- Return type:
dict
- bind_to_transform(transform)
Register exactly the bytes already validated for this producer’s finalization.
- Return type:
None
- canonical_owner(owner_id)
Replace only explicitly authorized scoped source concepts; products stay distinct.
- Return type:
str
- enrich_node(owner_id, existing)
Use this cohort’s reviewed CAS set on covered nodes and preserve other fields.
An omitted legacy CAS is outside this selected bundle’s active evidence, not a global rejection of its use by another independent source. Product catalog values, preparations and historical claims never become synonyms.
- Return type:
dict[str,str]
- identifier_claims(owner_id=None)
Return intact source/status/evidence tuples; original owners are never rewritten.
- Return type:
list[dict]
- identifier_owners(identifier, *, include_history=False)
Query all annotation owners without selecting a synonym-clique representative.
- Return type:
list[str]
- mappings()
Return every scoped mapping, including reviewed nonidentity relations.
- Return type:
list[dict]
- occurrences()
Return structured source occurrences with their unselected alternative groups.
- Return type:
list[dict]
- policy_parity()
Compare the existing case-specific CAS annotations before any policy migration.
- Return type:
dict
- products()
Return product specifications without attaching their qualifiers to generic ingredients.
- Return type:
list[dict]
- resolve_source(source_id, *, occurrence_id=None)
Resolve a source concept or verified occurrence before any label preference.
Native generic trait identifiers are not inferred to be MIM ingredients. Unknown source IDs return None; mismatched occurrence ownership is an error.
- Return type:
str|None
- verify_current()
Reject changed, deleted or undeclared bundle members before reusing a snapshot.
- Return type:
None
- write_annotations(path)
Write a deterministic graph-associated annotation artifact with its source pin.
- Return type:
dict
- kg_microbe.utils.ingredient_bundle.build_ingredient_lookup_bundle(consolidator, bundle, output)
Publish explicit legacy and scoped inputs together without flattening their semantics.
The existing consolidator handles its established native/legacy sources. The reviewed scoped cohort stays intact as a separate required input, resolved by source concept/occurrence before the legacy name index is consulted.
- Return type:
dict
- kg_microbe.utils.ingredient_bundle.load_ingredient_lookup_bundle(directory, *, manifest_sha256)
Verify both explicit lookup inputs and every companion artifact before index use.
kg_microbe.utils.ingredient_bundle_contract module
Portable, versioned validation for registry and occurrence companion claims.
- kg_microbe.utils.ingredient_bundle_contract.alternative_id(occurrence, group_key)
Identify an alternatives group independently of its current option payloads.
- Return type:
str
- kg_microbe.utils.ingredient_bundle_contract.annotation_id(payload)
Keep source/version/status claims distinct while binding revisions separately.
- Return type:
str
- kg_microbe.utils.ingredient_bundle_contract.annotation_table(rows)
Serialize one complete source claim per TSV row; evidence is one JSON cell.
- Return type:
bytes
- kg_microbe.utils.ingredient_bundle_contract.canonical_json(value)
Encode a deterministic JSON value, rejecting non-finite numbers.
- Return type:
bytes
- kg_microbe.utils.ingredient_bundle_contract.content_sha256(content)
Identify exact bytes, without trusting a caller-supplied digest.
- Return type:
str
- kg_microbe.utils.ingredient_bundle_contract.contract_resource(name)
Read the pinned packaged schema/profile, not a remote default.
- Return type:
bytes
- kg_microbe.utils.ingredient_bundle_contract.occurrence_id(source_id, source_occurrence_key)
Keep an explicitly named source occurrence stable across payload corrections.
- Return type:
str
- kg_microbe.utils.ingredient_bundle_contract.product_id(supplier, catalog_number)
Identify a supplier/catalog product specification, never an individual lot.
- Return type:
str
- kg_microbe.utils.ingredient_bundle_contract.read_annotation_table(content)
Restore exact identifier/source/status tuples with strict column and boolean types.
- Return type:
list[dict]
- kg_microbe.utils.ingredient_bundle_contract.read_json(content)
Read JSON without silently choosing between duplicate keys or NaN values.
- kg_microbe.utils.ingredient_bundle_contract.reviewed_products(review, sources)
Reconstruct artifacts from complete, independently owned reviewed claims.
The original mapping review format supplies each base disposition; companion decisions use the same verifier and additionally bind the explicit owner path. A consumer trusts a pinned bundle/review revision, never a self-issued checksum.
- Return type:
tuple[dict[str,bytes],dict]
- kg_microbe.utils.ingredient_bundle_contract.safe_member(root, name)
Confine manifest members to the bundle directory, including resolved symlinks.
- Return type:
Path
- kg_microbe.utils.ingredient_bundle_contract.stable_id(kind, key)
Identify a source claim or entity by its documented natural key.
- Return type:
str
- kg_microbe.utils.ingredient_bundle_contract.validate_bundle(root, *, expected_manifest_sha256=None, capabilities=frozenset({'catalog-products-v1', 'identifier-annotations-v1', 'ingredient-occurrences-v1', 'ingredient-scope-v1', 'owner-bound-review-v1'}))
Replay reviews and every projection before any consumer builds its indexes.
- Return type:
dict
- kg_microbe.utils.ingredient_bundle_contract.validate_member_hashes(root, manifest, capabilities=frozenset({'catalog-products-v1', 'identifier-annotations-v1', 'ingredient-occurrences-v1', 'ingredient-scope-v1', 'owner-bound-review-v1'}))
Reject unsupported versions, capabilities, changed bytes and undeclared files.
- Return type:
dict[str,bytes]
- kg_microbe.utils.ingredient_bundle_contract.validate_payload(kind, payload)
Validate claim shape and nonidentity rules before review or graph use.
- Return type:
None
kg_microbe.utils.ingredient_identity module
Structural admission and target-scoped exclusions for ingredient identities.
Curated policies reject specific lexical groundings and equivalence pairs, not ontology identifiers themselves. Separately, malformed CAS identifiers cannot establish a grounding. Neither check supplies a replacement identity.
- kg_microbe.utils.ingredient_identity.accepted_name_scope(subject, target, resolved, explicit_target)
Accept only the reviewed generic/specific Xanthine distinction with both routes intact.
- kg_microbe.utils.ingredient_identity.ingredient_authority_label(target)
Return the authority label recorded with a reviewed exclusion.
- Return type:
str
- kg_microbe.utils.ingredient_identity.ingredient_cas_annotations(target)
Return verified CAS node annotations; these are not same_as claims.
- kg_microbe.utils.ingredient_identity.ingredient_case_sensitive_name_scope(name)
Return (recognized family, exact-case target); unknown spelling in a family fails closed.
- kg_microbe.utils.ingredient_identity.ingredient_hydration_compatible(name, authority_label, authority_names=())
Reject hydration-scope changes without asserting a replacement chemical identity.
A separately supplied mapping still establishes the base identity. An explicit hydrate additionally needs a declared target label with the same water count. Independently native synonyms may refine an unspecified target hydrate only when their numeric scopes agree. Two unspecified hydrate scopes are compatible, but neither authorizes a specific water count. Missing evidence fails closed. Native labels remain queryable.
- kg_microbe.utils.ingredient_identity.ingredient_identity_policy()
Read and validate immutable-in-process curated identity rules.
- kg_microbe.utils.ingredient_identity.ingredient_mapping_allowed(name, target)
Reject invalid CAS inputs and reviewed ingredient-name/target pairs.
- Return type:
bool
- kg_microbe.utils.ingredient_identity.ingredient_name_scopes()
Return existing (ordinary query routes, CAS annotations), excluding exact-case aliases.
- kg_microbe.utils.ingredient_identity.ingredient_name_target(name)
Return the reviewed target for an explicitly scoped query, if any.
- kg_microbe.utils.ingredient_identity.ingredient_xref_allowed(subject, target)
Reject reviewed false equivalences in either serialization direction.
- Return type:
bool
kg_microbe.utils.ingredient_kgx module
Typed scalar transport for reviewed ingredient assertions in finalized KGX.
- kg_microbe.utils.ingredient_kgx.ingredient_graph_id(identifier)
Disambiguate a local ingredient from Biolink’s existing MIM namespace.
- kg_microbe.utils.ingredient_kgx.ingredient_kgx_profile()
Read the packaged application profile paired with the pinned Biolink model.
- kg_microbe.utils.ingredient_kgx.json_scalar(value)
Keep JSON arrays and source text inside one scalar, never a KGX pipe list.
- kg_microbe.utils.ingredient_kgx.validate_ingredient_fields(row, *, is_node)
Reject malformed, pooled or unversioned extension values at graph boundaries.
kg_microbe.utils.ingredient_scope module
Versioned ingredient scope contract, independent of identifier preference.
- kg_microbe.utils.ingredient_scope.profile_metadata(metadata, *, default_profile=True)
Declare extension meanings, rejecting conflicting prefix or slot definitions.
- Return type:
dict
- kg_microbe.utils.ingredient_scope.read_profile_table(content)
Read a required-profile TSV without losing extensions or accepting legacy fallback.
- Return type:
tuple[dict,list[str],list[dict]]
- kg_microbe.utils.ingredient_scope.read_scope_metadata(header)
Parse declared scope metadata without accepting duplicate YAML keys.
- Return type:
dict
- kg_microbe.utils.ingredient_scope.scope_profile()
Return the public, serializable profile without exposing cached mutable state.
- Return type:
dict
- kg_microbe.utils.ingredient_scope.validate_scope_row(row, metadata)
Validate scope declarations and return identity permission, or None for legacy.
SUPPORTED and authorization are asserted review outcomes. A release consumer must additionally verify the bundle’s content/owner-bound review evidence. Equal scope vocabulary values alone never establish chemical equivalence.
- Return type:
bool|None
- kg_microbe.utils.ingredient_scope.write_profile_table(metadata, fields, rows)
Serialize declared profile extensions deterministically as standard SSSOM/TSV.
- Return type:
bytes
kg_microbe.utils.isolation_source_mapping_utils module
Loader and validator for the BacDive isolation-source → ontology mapping table.
The mapping table at mappings/isolation_source_to_ontology.tsv records both
high-quality manual mappings and lower-quality automated lexical hits. The
loader applies a conservative confidence policy so that only trustworthy rows
are honored at runtime; everything else is dropped (and the BacDive transform
emits the raw isolation_source:* placeholder node instead).
Policy (applied in order; first match wins):
1. Drop the row if it carries no object_id (explicitly unmapped).
2. Drop the row if its object_source is in DISALLOWED_OBJECT_SOURCES
(e.g.
UO— units of measurement are never an isolation source).
Drop the row if its
object_labelmatches a banned-family substring for the subject’s apparent kind (anatomy / host / generic environment) — these indicate a family mismatch like an anatomical label being mapped to a facility, document, or clinical condition.Honor the row if
predicate_id == 'skos:exactMatch'andconfidence == 'high'.Honor the row if
mapping_justification == 'semapv:ManualMappingCuration'(any row that has been touched by a human curator).Otherwise drop. In practice this rejects
ols4_autolexicalskos:closeMatchrows that are most prone to false positives.
This loader is intentionally strict. Promoting an automated lexical hit to “trusted” should be a deliberate curator action (set predicate to exactMatch or change the mapping_justification to ManualMappingCuration) rather than an implicit upgrade.
- kg_microbe.utils.isolation_source_mapping_utils.load_isolation_source_mappings(mapping_path=None)
Load trusted isolation-source mappings keyed by normalized subject label.
The key is the lowercased
subject_label_normalized(or, as a fallback, the lowercasedsubject_label) — this matches the canonicalization the BacDive transform already applies before lookups (lowercasing the raw BacDive isolation-source string).- Parameters:
mapping_path (
Optional[Path]) – Optional override for the TSV path. Defaults to the committedmappings/isolation_source_to_ontology.tsvnext to this repo.- Return type:
Dict[str,Tuple[str,str,str,Optional[str]]]- Returns:
{normalized_label: (object_id, object_label, object_source, predicate_override)}for every row that passes both the family-compatibility check and the trust policy. The fourth element is the override CURIE (METPO:2000067 / METPO:2000068) if the row’snotescolumn references one, otherwiseNone. Rows that are explicitly unmapped, family-mismatched, or below the trust threshold are silently skipped (with INFO-level logging of the skip count).
- kg_microbe.utils.isolation_source_mapping_utils.normalize_isolation_source_label(label)
Normalize a raw BacDive isolation-source label to the loader key form.
The mapping table stores keys with hyphens, commas, underscores, and slashes collapsed to single spaces, then lowercased. This matches the
subject_label_normalizedcolumn inmappings/isolation_source_to_ontology.tsv(e.g.Bovinae-Cow,-Cattle→bovinae cow cattle).- Return type:
str
kg_microbe.utils.lpsn_utils module
Shared LPSN nomenclature helpers.
LPSN’s GSS export marks each record’s nomenclatural standing in a status
column. Only rows whose status contains correct name are the currently
accepted name; everything else — synonyms, illegitimate names, rejected names —
defers to another record through record_lnk.
Two transforms anchor edges on lpsn:<record_no> and so need the same
resolution:
bacdive matches a strain’s reported species against GSS names (#684);
microbedecoder takes
LPSN_IDstraight from its source column (#746).
Both were placing live strains under deprecated records. The resolver lived on
BacDiveTransform when only bacdive needed it; it holds no BacDive state, so
it moved here rather than being duplicated.
- kg_microbe.utils.lpsn_utils.LPSN_CORRECT_NAME = 'correct name'
LPSN GSS marks the currently accepted name with this phrase in
status.
- kg_microbe.utils.lpsn_utils.LPSN_MAX_SYNONYM_HOPS = 25
Deepest synonym chain observed in the shipped GSS is 3 hops. The cap is a backstop against a cyclic chain in a future release, not a real limit — and it matters because the failure mode of an unbounded walk is a hang rather than a wrong answer, which costs a CI job its whole budget (#742).
- kg_microbe.utils.lpsn_utils.resolve_accepted_records(gss_rows)
Map each non-current LPSN record to the accepted name it defers to.
Anchoring on a synonym puts a living strain under a deprecated class, which reasoners and any closure step then propagate. Dropping those edges would lose real information — the organism genuinely is that taxon, LPSN has simply renamed it — so
record_lnk, LPSN’s own crosswalk, is followed instead.Measured on the shipped GSS: every
record_lnkresolves to a real row, 7,018 synonyms reach a correct name in one hop, 250 in two, 2 in three, and no chain cycles. 192 dead-end with no link and are left alone, there being nothing better to point them at.Bounded as well as cycle-guarded: the
seenset is the real protection, but if it were lost the walk would hang rather than answer wrongly, so the iteration count is capped too.- Parameters:
gss_rows (
Dict[str,dict]) –record_noto its parsed GSS row.- Return type:
Dict[str,str]- Returns:
record_noto acceptedrecord_no, for those that move. A record already carryingcorrect name, or whose chain dead-ends, is absent from the mapping.
kg_microbe.utils.mapping_file_utils module
Utilities for handling mapping files from remote sources.
- class kg_microbe.utils.mapping_file_utils.MetpoTreeNode(iri, label, synonyms=None, biolink_equivalent=None, bacdive_json_paths=None)
Bases:
objectRepresents a node in the METPO class hierarchy tree.
For example, consider METPO:1000602 The node would have:
iri = “METPO:1000602” label = “aerobic” synonyms = [“aerobic”, “aerobe”, “Ox_aerobe”] bacdive_json_paths = [] biolink_equivalent = “” children = [] parent = {‘iri’: ‘METPO:1000601’, ‘label’: ‘oxygen preference’, …}
- add_child(child)
Add a child node and set its parent.
- find_biolink_equivalent_parent()
Find the closest parent (including self) that has a biolink equivalent.
(value populated in the biolink equivalent column).
- Return type:
Optional[str]
- find_synonym_node(synonym)
Find the node that contains the given synonym.
- Return type:
Optional[MetpoTreeNode]
- kg_microbe.utils.mapping_file_utils.generate_assay_entity_edges(assay_data, edge_header, *, go_authority=None)
Generate assay→entity edges from assay_kits_simple.json data.
Creates methodological reference edges showing what each assay tests: - Assays → GO molecular functions / EC activities: MICRO:0001206 - Assays → GO biological processes: MICRO:0001215 - Chemical assays → ChEBI reagents: MICRO:0000065
These edges are created once upfront, independent of organism data.
- Parameters:
assay_data (
dict) – Dictionary loaded from assay_kits_simple.jsonedge_header (
List[str]) – List of column names for edge rows (from Transform base class)go_authority – Validated, in-memory GO authority shared with target nodes
- Returns:
List of edge rows formatted for writerows(), each row matching edge_header length
- Return type:
List[List]
- Example edges:
kgmicrobe.assay:API_zym_alkaline_phosphatase → MICRO:0001206 → GO:0004035 kgmicrobe.assay:API_20NE_GLU__Ferm → MICRO:0001215 → GO:0019660 kgmicrobe.assay:API_50CHac_ERY → MICRO:0000065 → CHEBI:17113
- kg_microbe.utils.mapping_file_utils.generate_assay_entity_nodes(assay_data, node_header, *, go_authority=None)
Generate stub node rows for CHEBI / EC / GO entities referenced by assays.
The companion
generate_assay_entity_edgesemits assay→CHEBI/EC/GO edges whose targets are normally supplied by the ontologies transform. When an obsolete CHEBI ID (e.g. CHEBI:17004 ‘D-Tagatose’, missing from current chebi.json) is referenced, the merge step falls back to a KGX-generatedbiolink:NamedThingstub with empty name/category — which fails biolink domain/range checks and erases the label.GO declarations come from shared local authority, including historical labels/aspects/deprecation. Imported defaults must not pollute canonical categories;
considerhints never assert identity with another GO ID.- Parameters:
assay_data (
dict) – Dictionary loaded from assay_kits_simple.jsonnode_header (
List[str]) – Node header (from Transform base class)go_authority – Validated, in-memory GO authority shared with assay edges
- Return type:
List[List]- Returns:
List of node rows
- kg_microbe.utils.mapping_file_utils.generate_assay_nodes(assay_data, node_header)
Generate assay node rows from assay_kits_simple.json data.
Creates one node per well/test component across all API kits, with rich metadata combined into the description field (kit name, well name, test type, and original description).
- Parameters:
assay_data (
dict) – Dictionary loaded from assay_kits_simple.jsonnode_header (
List[str]) – List of column names for node rows (from Transform base class)
- Returns:
List of node rows formatted for writerows(), each row matching node_header length
- Return type:
List[List]
- Example node structure:
id: kgmicrobe.assay:API_zym_alkaline_phosphatase category: biolink:Procedure name: API zym - Alkaline phosphatase description: Tests for Alkaline phosphatase activity using chromogenic substrate.
Kit: API zym, Well: Alkaline phosphatase, Type: enzyme
- kg_microbe.utils.mapping_file_utils.load_assay_kit_mappings(assay_file=None)
Load assay kit mappings from the local assay_kits_simple.json file.
This function processes API kit data (e.g., “API zym”, “API coryne”) from BacDive and creates mappings for enzyme and chemical tests.
For example, from the “API zym” kit: - Well “Esterase” with type “enzyme” and ec_number “3.1.1.1” or go_terms “GO:0004806” - Well “GLY” with type “chemical” and chebi_id “CHEBI:17754”
The structure of the returned mapping is: {
- “API zym”: { # kit_name (exact match from JSON)
- “wells”: { # well-specific mappings
- “Esterase”: { # name from well (matches BacDive API keys)
“type”: “enzyme”, “go_terms”: [”GO:0004806”], “ec_number”: [“3.1.1.1”], “chebi_id”: []
}, “metpo_predicates”: { # kit-level predicates
- “positive”: {
“id”: “METPO:2000302”, “label”: “shows activity of”
}, “negative”: {
“id”: “METPO:2000303”, “label”: “does not show activity of”
}
}
}, “API 50CHac”: {
- “wells”: {
- “GLY”: { # name from well (matches BacDive API keys like “GLY”, “ERY”)
“type”: “chemical”, “chebi_id”: [“CHEBI:17754”], “go_terms”: [], “ec_number”: []
}, “metpo_predicates”: {
- “positive”: {
“id”: “METPO:2000011”, “label”: “ferments”
}, “negative”: {
“id”: “METPO:2000037”, “label”: “does not ferment”
}
}
}
}
- Parameters:
assay_file (
Optional[Path]) – Exact selected raw-root file; defaults to the repository download location- Returns:
Dictionary mapping kit names to well labels to their properties
- Return type:
Dict[str, Dict[str, Dict]]
- Raises:
FileNotFoundError – If the assay kits file is not found
ValueError – If the JSON content is invalid
- kg_microbe.utils.mapping_file_utils.load_metpo_enzyme_mappings()
Load METPO enzyme activity mappings from the METPO properties sheet.
This function looks for rows with the synonym ‘Physiology and metabolism.enzymes.[].activity’ and maps BacDive’s enzyme “activity” values (“+”/”-”) to METPO predicates based on the “assay outcome” column.
For example, from the rows: | ID | label | synonym TUPLES | assay outcome | |---------------|—————————-|-------------------------------|—————| | METPO:2000302 | shows activity of | hasRelatedSynonym ‘Phys…[]’ | + | | METPO:2000303 | does not show activity of | hasRelatedSynonym ‘Phys…[]’ | - |
This creates mappings: {
‘+’: {‘curie’: ‘METPO:2000302’, ‘label’: ‘shows activity of’}, ‘-’: {‘curie’: ‘METPO:2000303’, ‘label’: ‘does not show activity of’}
}
- Returns:
Dictionary mapping activity values (“+”/”-”) to METPO predicate info
- Return type:
Dict[str, Dict[str, str]]
- Raises:
requests.exceptions.HTTPError – If unable to fetch from remote URL
ValueError – If the response content is empty or invalid
- kg_microbe.utils.mapping_file_utils.load_metpo_mappings(synonym_column)
Load METPO mappings from METPO classes ROBOT template file for a given synonym column.
Implements the logic to find appropriate _predicates_ by traversing the parent hierarchy to find biolink equivalent and then mapping to properties.
For ambiguous synonyms (e.g., “yes” or “no” that appear under multiple parent concepts), compound keys are created using the parent label as context (e.g., “motility.yes”, “sporulation.yes”).
- Parameters:
synonym_column (
str) – The column name to use for synonyms (e.g., ‘bacdive keyword synonym’, ‘madin synonym or field’, etc.)- Returns:
Dictionary mapping synonyms to METPO curie, label, and predicate information. Format: {synonym: {‘curie’: metpo_curie, ‘label’: metpo_label, ‘predicate’: predicate_label}} For ambiguous values, also includes compound keys like “parent.synonym”
- Return type:
Dict[str, Dict[str, str]]
- Raises:
requests.exceptions.HTTPError – If unable to fetch from remote URL
ValueError – If the response content is empty or invalid
- kg_microbe.utils.mapping_file_utils.load_metpo_metabolite_production_mappings()
Load METPO metabolite production mappings from the METPO properties sheet.
This function specifically looks for rows with the synonym ‘produces’ and maps BacDive’s “production” values (“yes”/”no”) to METPO predicates based on the “assay outcome” column.
For example, from the rows: | ID | label | synonym property and value TUPLES | assay outcome | |---------------|——————|---------------------------------------------|—————| | METPO:2000202 | produces | oboInOwl:hasRelatedSynonym ‘produces’ | + | | METPO:2000222 | does not produce | oboInOwl:hasRelatedSynonym ‘produces’ | - |
This creates mappings: {
‘yes’: {‘curie’: ‘METPO:2000202’, ‘label’: ‘produces’}, ‘no’: {‘curie’: ‘METPO:2000222’, ‘label’: ‘does not produce’}
}
The mapping is: - “assay outcome” = “+” maps to “production” = “yes” - “assay outcome” = “-” maps to “production” = “no”
- Returns:
Dictionary mapping production values (“yes”/”no”) to METPO predicate info
- Return type:
Dict[str, Dict[str, str]]
- Raises:
requests.exceptions.HTTPError – If unable to fetch from remote URL
ValueError – If the response content is empty or invalid
- kg_microbe.utils.mapping_file_utils.load_metpo_metabolite_utilization_mappings()
Load METPO metabolite utilization mappings from the METPO properties sheet.
This function parses the “synonym property and value TUPLES” and “assay outcome” columns to extract mappings for metabolite utilization predicates. The mappings are used to convert BacDive’s “kind of utilization tested” values into appropriate METPO predicates.
For example, from the rows: | ID | label | synonym property and value TUPLES | assay outcome | |---------------|——————|---------------------------------------------|—————| | METPO:2000003 | builds acid from | oboInOwl:hasRelatedSynonym ‘builds acid from’ | + | | METPO:2000028 | does not build acid from | oboInOwl:hasRelatedSynonym ‘builds acid from’ | - |
This creates mappings: {
- ‘builds acid from’: {
‘+’: {‘curie’: ‘METPO:2000003’, ‘label’: ‘builds acid from’}, ‘-’: {‘curie’: ‘METPO:2000028’, ‘label’: ‘does not build acid from’}
}
The sign (+ or -) is now directly read from the “assay outcome” column.
- Returns:
Dictionary mapping utilization type synonyms to sign-based predicate info
- Return type:
Dict[str, Dict[str, Dict[str, str]]]
- Raises:
requests.exceptions.HTTPError – If unable to fetch from remote URL
ValueError – If the response content is empty or invalid
- kg_microbe.utils.mapping_file_utils.normalize_biolink_category(category)
Map a deprecated Biolink category URI/CURIE to its recommended replacement.
- Return type:
str
- kg_microbe.utils.mapping_file_utils.uri_to_curie(uri)
Convert a URI to a CURIE using the local prefix map and known ontology aliases.
It also checks if the input uri is already a CURIE, in which case it returns the CURIE value as-is.
>>> uri_to_curie("https://w3id.org/metpo/1000059") 'METPO:1000059'
>>> uri_to_curie("METPO:1000059") 'METPO:1000059'
- Parameters:
uri (
str) – The URI to convert, or a CURIE that’s already in the correct format- Return type:
str- Returns:
The CURIE representation of the URI, or the original input if it’s already a CURIE
kg_microbe.utils.mediadive_bulk_download module
MediaDive bulk download utility.
This module provides functionality to download all MediaDive data in bulk to avoid repeated API calls during transforms.
- kg_microbe.utils.mediadive_bulk_download.download_detailed_media(media_list, max_workers=5, retry_count=3, retry_delay=2.0, requests_per_second=10.0)
Download detailed recipe information for all media.
- Return type:
Dict[str,Dict]
Args:
media_list: List of basic media records max_workers: Number of parallel download threads retry_count: Number of retries on request failure retry_delay: Seconds between retries (overridden by Retry-After on 429) requests_per_second: Reserved for caller compatibility; concurrency is limited by max_workers
Returns:
Dictionary mapping medium_id -> detailed_recipe_data
- kg_microbe.utils.mediadive_bulk_download.download_mediadive_bulk(basic_file, output_dir, max_workers=5, retry_count=3, retry_delay=2.0, ignore_cache=False)
Download all MediaDive data in bulk.
This is the main entry point called from kg_microbe.download.
Args:
basic_file: Path to mediadive.json (basic media list) output_dir: Directory to save bulk data files max_workers: Number of parallel download threads (default: 5, polite for small APIs) retry_count: Number of retries on request failure retry_delay: Seconds between retries (overridden by Retry-After on 429) ignore_cache: If True, discard cached HTTP responses and re-fetch from
the API. Rebuilding the bulk files from a warm cache would otherwise reproduce the old data byte for byte.
- kg_microbe.utils.mediadive_bulk_download.download_medium_strains(media_list, max_workers=5, retry_count=3, retry_delay=2.0, requests_per_second=10.0)
Download strain associations for all media.
- Return type:
Dict[str,List]
Args:
media_list: List of basic media records max_workers: Number of parallel download threads retry_count: Number of retries on request failure retry_delay: Seconds between retries (overridden by Retry-After on 429) requests_per_second: Reserved for caller compatibility; concurrency is limited by max_workers
Returns:
Dictionary mapping medium_id -> list_of_strain_data
- kg_microbe.utils.mediadive_bulk_download.extract_compounds_from_media(detailed_media)
Extract compound data from embedded structure in detailed_media.
Instead of making API calls, extract compound info directly from media_detailed.json.
- Return type:
Dict[str,Dict]
Args:
detailed_media: Dictionary of detailed medium recipes
Returns:
Dictionary mapping compound_id -> compound_data
- kg_microbe.utils.mediadive_bulk_download.extract_solutions_from_media(detailed_media)
Extract solution data from embedded structure in detailed_media.
Instead of making API calls, extract solutions directly from media_detailed.json.
- Return type:
Dict[str,Dict]
Args:
detailed_media: Dictionary of detailed medium recipes
Returns:
Dictionary mapping solution_id -> solution_data
- kg_microbe.utils.mediadive_bulk_download.get_json_from_api(url, retry_count=3, retry_delay=2.0, verbose=False, session=None)
Get JSON data from MediaDive API with retry logic.
Respects Retry-After headers on 429 responses. A requested delay over five minutes ends the request without retrying early; malformed headers use retry_delay instead.
- Return type:
Dict
Args:
url: Full API URL to fetch retry_count: Number of retries on failure retry_delay: Delay in seconds between retries (overridden by Retry-After on 429) verbose: If True, log empty responses (useful for debugging) session: Optional requests Session to reuse (uses this thread’s session if None)
Returns:
Dictionary with API response data (empty dict on failure or empty response)
- kg_microbe.utils.mediadive_bulk_download.load_basic_media_list(basic_file)
Load basic media list from already downloaded file.
- Return type:
List[Dict]
Args:
basic_file: Path to mediadive.json file
Returns:
List of media records with basic info
- kg_microbe.utils.mediadive_bulk_download.save_json_file(data, filepath, description)
Save data to JSON file with logging.
- kg_microbe.utils.mediadive_bulk_download.setup_cache(cache_dir=None, migrate_legacy=False, clear=False)
Enable HTTP caching for subsequent sessions created by this module.
- Return type:
Optional[Path]
Args:
- cache_dir: Directory to hold the cache database. If None, caching is
disabled and plain (uncached) sessions are used.
- migrate_legacy: If True, adopt a cache database left in the current
working directory by older versions of this module (which stored it there). Off by default: this moves a file outside cache_dir, so only the real download entry point should ask for it.
- clear: If True, discard any existing cached responses so every request
goes back to the API. This is what kg download –ignore-cache needs; without it the bulk JSON files are rebuilt from cached HTTP responses and never actually refresh.
Returns:
Path to the cache database, or None if caching was disabled.
kg_microbe.utils.metpo_liveness module
Check that the METPO terms this repo asserts still exist upstream.
A CURIE hardcoded in a constant keeps being emitted long after the ontology
stops declaring it: nothing errors, nothing looks different, and the graph
quietly asserts a term that no longer means anything. That is how
METPO:2000511 came to carry 706,765 shipped edges after being obsoleted, and
METPO:1001000 to sit in 503 node categories (#909).
The pinned metpo.json is the authority. It is pinned (#900), so this check
answers “does the release we build against still declare this”, not “does some
newer release”, which is the question that matters for what we ship.
- kg_microbe.utils.metpo_liveness.KNOWN_DEPRECATED: Dict[str, str] = {'METPO:2000054': '#909', 'METPO:2000508': '#909', 'METPO:2000511': '#909'}
METPO CURIEs this repo still emits despite upstream deprecation, each with the issue that tracks removing it. An entry here is a debt, not an exemption: it says someone looked, found no live replacement, and recorded why. Adding one without an issue defeats the check.
- kg_microbe.utils.metpo_liveness.deprecated_metpo_terms(path=None)
Return every METPO CURIE the pinned release marks deprecated.
- Parameters:
path (
Optional[Path]) – Override for the ontology path, for tests.- Return type:
Set[str]- Returns:
Set of CURIEs; empty when the ontology is not available locally.
- kg_microbe.utils.metpo_liveness.metpo_json_path()
Return the pinned metpo.json this repo builds against.
- Return type:
Path- Returns:
Path to the downloaded ontology, which may not exist.
kg_microbe.utils.metpo_predicates module
Shared METPO predicate → biolink predicate mapping.
- kg_microbe.utils.metpo_predicates.to_biolink_predicate(predicate)
Map METPO or other predicate to biolink predicate; pass through biolink predicates unchanged.
- Return type:
str
kg_microbe.utils.microbial_trait_mappings module
Load curated microbial trait mappings from turbomam/microbial-trait-mappings TSVs.
These mappings provide subject_label -> (object_id, object_source, biolink_predicate, object_category) for metatraits transform edge resolution. Used as the authoritative lookup before METPO fallback.
- kg_microbe.utils.microbial_trait_mappings.canonical_mapping_paths(mappings_dir=None)
Select the current positive canonical TSV inventory for readers and freshness.
- Return type:
tuple[Path,...]
- kg_microbe.utils.microbial_trait_mappings.load_microbial_trait_mappings(mappings_dir=None)
Load all positive mapping TSVs from mappings/canonical/.
Excludes *_negative_mappings.tsv files.
- Parameters:
mappings_dir (
Optional[Path]) – Override default mappings directory- Return type:
Dict[str,Dict[str,str]]- Returns:
Dict mapping subject_label -> {object_id, object_label, object_source, biolink_predicate, object_category}
kg_microbe.utils.ner_utils module
NLP utilities.
- kg_microbe.utils.ner_utils.annotate(df, prefix, exclusion_list, outfile, llm=False, chemical_loader=None)
Annotate dataframe column text using oaklib + llm.
- Parameters:
df (
DataFrame) – Input DataFrameprefix (
str) – Ontology to be used.exclusion_list (
List) – Tokens that can be ignored.chemical_loader (
Optional[ChemicalMappingLoader]) – Optional ChemicalMappingLoader used as a CHEBI fallback when OAK annotation fails. Consults the unified chemical mappings (which already contain the curated chebi_manual_annotation.tsv rows).
kg_microbe.utils.oak_utils module
Description: This file contains utility functions for the OAK client.
- kg_microbe.utils.oak_utils.get_label(oi, curie)
Return the label of a given curie via oaklib.
- kg_microbe.utils.oak_utils.search_by_label(oi, label, limit=5)
Search ontology for entities by label/name.
- Return type:
List[str]
- Args:
oi: OAK OntologyInterface instance label: The label/name to search for limit: Maximum results
- Returns:
List of CURIEs matching the search
kg_microbe.utils.ontology_resolution module
Pure ontology category policies over explicit lookup results and adapters.
- class kg_microbe.utils.ontology_resolution.CategoryAdapter(*args, **kwargs)
Bases:
ProtocolMinimal adapter surface needed by ChEBI category policy.
- ancestors(term_id, predicates)
Return ancestor identifiers along the supplied predicates only.
- label(term_id)
Return the preferred label, if present.
- relationships(term_id, predicates)
Return relationships matching the requested predicates.
- kg_microbe.utils.ontology_resolution.chebi_category(term_id, adapter)
Classify by asserted is-a ancestry, never by a chemical’s has-role links.
- Return type:
str
- kg_microbe.utils.ontology_resolution.foodon_category(term_id)
Apply FOODON’s asserted organism-versus-food boundary consistently across sources.
- Return type:
str
- kg_microbe.utils.ontology_resolution.go_category_for_namespace(namespace)
Map an OBO GO namespace to its Biolink category.
- Return type:
str
- kg_microbe.utils.ontology_resolution.ncbitaxon_category(_term_id)
Return the invariant category for an NCBITaxon term.
- Return type:
str
- kg_microbe.utils.ontology_resolution.pato_category(_term_id)
Return the invariant category for a PATO term (madin_etal’s choice, #1015).
- Return type:
str
- kg_microbe.utils.ontology_resolution.replace_category_by_prefix(line, id_index, category_index)
Replace a TSV category according to its normalized identifier prefix.
- Return type:
str
- kg_microbe.utils.ontology_resolution.replace_deprecated_category_names(category)
Replace Biolink categories removed in Biolink 4.x.
- Return type:
str
- kg_microbe.utils.ontology_resolution.uberon_category(_term_id)
Return the invariant category for an UBERON term.
- Return type:
str
kg_microbe.utils.ontology_utils module
Ontology utilities for category assignment and term processing.
alias of
OntologyDbUnavailableError
- class kg_microbe.utils.ontology_utils.DbEnsureResult(usable: bool, built: bool = False)
Bases:
NamedTupleOutcome of an
_ensure_*_dbcall.usablekeeps the historical truthiness (callers writeif _ensure_chebi_db(...)), whilebuiltsays whether this call produced a new DB. Callers need that distinction and cannot infer it: a restored.prevor an adopted orphan also makes a file appear where there was none, which is what made an earlier fingerprint-based guess misreport a restore as a fresh build (F2).-
built:
bool Alias for field number 1
-
usable:
bool Alias for field number 0
-
built:
- exception kg_microbe.utils.ontology_utils.FatalOntologyError
Bases:
BaseExceptionAn ontology is unusable and no per-item fallback is honest.
Deliberately not an
Exception. Adapters resolve lazily, so the first attribute access — and therefore the failure — lands wherever the transform happens to touch the adapter first, which is almost always inside atrywhoseexcept Exceptionwas written to absorb a per-item lookup miss. A missing DB is not a lookup miss: swallowing it means every ChEBI node gets the default category, every GO term becomes molecular_function, every label becomes a bare numeric ID, and the run exits 0 with a systematically wrong graph. Those handlers are individually reasonable, and patching each one was tried — it regressed three times, because the next broad handler someone writes re-opens the hole.Inheriting
BaseExceptionmakes that structural instead of a convention: noexcept Exception, here or in oaklib/pandas, can turn a fatal ontology failure into wrong data.finallyblocks still run, so cleanup is intact.One constraint this imposes: a
BaseExceptionraised inside amultiprocessing.Poolworker is not caught by the worker loop and can hang the pool. The metatraits pool resolves NCBITaxon eagerly in the parent (_ensure_ncbitaxon_db_ready) through its own module-local adapter, so no worker ever resolves one of these proxies — keep it that way.
- class kg_microbe.utils.ontology_utils.KeptTarget(prev_path: str | None = None, link_target: str | None = None)
Bases:
NamedTupleWhat
_clear_build_target()displaced, so it can be put back.A real file is moved aside to
<db>.prev; a symlink is recorded by its target and recreated on restore. Recording the link matters: pointing at a prebuilt DB with a symlink is a supported way to supply one, and a failed build used to delete it irrecoverably while telling the user to supply a prebuilt DB (F1).-
link_target:
Optional[str] Alias for field number 1
-
prev_path:
Optional[str] Alias for field number 0
-
link_target:
Bases:
FatalOntologyErrorNo usable SemSQL DB could be produced for an ontology.
Distinct from
OntologyVersionMismatchErrorso callers can degrade on “no DB at all” without also swallowing a deliberate*_VERSION_CHECK=strictabort —get_chebi_categoryrelies on exactly that distinction.
- exception kg_microbe.utils.ontology_utils.OntologyVersionMismatchError
Bases:
FatalOntologyErrorA DB and the OWL it must track are stamped with different releases.
Raised only by the
assert_*_version_alignmentgates understrict. The whole point is to abort rather than emit a graph built from two releases, so this must not be catchable asException: under the old plainRuntimeErrora strict ChEBI abort was swallowed per-row byget_chebi_categoryand a strict GO abort byoak_utils.get_label.
- kg_microbe.utils.ontology_utils.assert_chebi_version_alignment(db_path, strict=None)
Guard that the ChEBI lookup DB matches the transform’s OWL release.
chebi.dbdecides each ChEBI node’s Biolink category (SmallMolecule vs ChemicalRole vs macromolecule) via ancestor lookups, while the nodes themselves are emitted fromchebi.owl. A release gap means a term the transform emits may be absent from the DB, so its category silently falls back to the default — the same failure mode the GO gate exists to prevent.Defaults to warn rather than raise: when
semsqlis unavailable the pipeline legitimately falls back to a prebuiltchebi.dbof a different release, and aborting the run would be worse than mis-categorising a handful of terms. SetKG_CHEBI_VERSION_CHECK=strict(or passstrict=True) to fail loud. No-op when either stamp can’t be read.- Parameters:
db_path (
str) – Path tochebi.db.strict (
Optional[bool]) – Override the env var / default strictness.
- Raises:
RuntimeError – On mismatch when strict.
- Return type:
None
- kg_microbe.utils.ontology_utils.assert_go_version_alignment(strict=None, *, raw_dir=None)
Guard that GO’s derived
go.jsonmatches its single sourcego.owl.Since fix 2 (#604) GO is single-source:
go.owlis the only download, andgo.json(transform output) andgo.db(MF/BP/CC aspect map) are both derived from it — the transform regenerates a stale go.json and_ensure_go_dbrebuilds a drifted go.db. This gate is the belt-and-braces check that the derived go.json actually tracks go.owl: a leftover pre-fix-2 go.json, or a conversion that didn’t re-run, would make MF/CC terms silently fall through to thebiological_processdefault. Compare the two releases and, on mismatch, raise (strict) or warn loudly. No-op when either release stamp can’t be read (e.g. a source is absent), so a missed versionIRI never false-alarms — only two readable-but-different stamps trip the gate.strictdefaults to fail-loud (raise). Since the verdict rests on a release-stamp heuristic,KG_GO_VERSION_CHECK=warndowngrades to a warning — an escape hatch if the stamps ever disagree spuriously. An explicitraw_dirselects the same source release as the producer; only legacy callers without that argument useGO_SOURCE.- Return type:
None
- kg_microbe.utils.ontology_utils.assert_ncbitaxon_version_alignment(db_path, strict=None)
Guard that the NCBITaxon lookup DB matches the transform’s OWL release.
The metatraits transform looks taxa up in
ncbitaxon.db(an OAK-fetched prebuilt SemSQL DB whose release is whatever OAK last downloaded) while its nodes are emitted fromncbitaxon.owl. If the two are different releases, lookups can resolve against taxa that differ from those emitted. Compare theowl:versionInfoindb_pathwithncbitaxon.owl’s versionIRI and, on mismatch, warn (default) or raise. No-op when either stamp can’t be read.Unlike the GO gate this defaults to warn — the OAK cache and the pinned
ncbitaxon.owllegitimately drift (OAK auto-refreshes to the latest), and NCBITaxon labels/lineage are stable, so a mismatch is worth surfacing loudly but not aborting. SetKG_NCBITAXON_VERSION_CHECK=strict(or passstrict=True) to fail loud instead.- Return type:
None
- kg_microbe.utils.ontology_utils.get_chebi_adapter()
Return a lazily-resolved guarded ChEBI adapter.
Kept as a named accessor because ChEBI has the most callers; it is now the same lazy proxy the other ontologies use, over the shared cache.
- Returns:
adapter proxy.
- kg_microbe.utils.ontology_utils.get_chebi_category(chebi_term_id, chebi_adapter=None)
Return appropriate Biolink category for ChEBI term.
ChEBI terms can be: - Macromolecules (proteins, nucleic acids, polysaccharides) → biolink:MacromolecularComplex - Roles (e.g., “antioxidant”, “inhibitor”) → biolink:ChemicalRole - Small molecules (default) → CHEBI_CATEGORY (biolink:ChemicalEntity, see constants.py)
- Return type:
str
Args:
chebi_term_id: ChEBI term ID (e.g., “CHEBI:16828”) chebi_adapter: Optional OAK adapter for ChEBI ontology
Returns:
Biolink category string
- kg_microbe.utils.ontology_utils.get_ec_adapter()
Return a lazily-resolved guarded EC adapter. :return: adapter proxy.
- kg_microbe.utils.ontology_utils.get_foodon_category(foodon_term_id)
Return the Biolink category for a FOODON term: always biolink:Food (#1015).
- Return type:
str
- kg_microbe.utils.ontology_utils.get_go_adapter()
Return a lazily-resolved guarded GO adapter. :return: adapter proxy.
- kg_microbe.utils.ontology_utils.get_go_aspect(go_term_id, *, raw_dir=None, authority=None)
Return the OBO namespace/aspect for a GO term (from
go.db).Uses the shared, fingerprinted GO authority. Exact replacements inherit the canonical aspect; retained obsolete terms keep their historical aspect. Missing references raise rather than guessing a namespace.
- Parameters:
go_term_id (
str) – GO CURIE (e.g."GO:0004096").raw_dir (
Optional[Path]) – Producer’s selected raw authority directory, when supplied.authority – Already prepared immutable authority; performs no IO.
- Return type:
str- Returns:
One of
"molecular_function","biological_process","cellular_component". RaisesOntologyDbUnavailableErrorfor an unusable authority orGoReferenceErrorfor an unknown term.
- kg_microbe.utils.ontology_utils.get_go_category_by_aspect(go_term_id, go_adapter=None, *, raw_dir=None, authority=None)
Return Biolink category based on GO aspect (namespace).
GO terms have three aspects (namespaces): - molecular_function → biolink:MolecularActivity - biological_process → biolink:BiologicalProcess - cellular_component → biolink:CellularComponent
- Return type:
str
Args:
go_term_id: GO term ID (e.g., “GO:0004096”) go_adapter: Unused (kept for backward compatibility with existing callers).
Lookup uses the shared fingerprinted authority beside GO_SOURCE.
raw_dir: Producer’s explicit raw directory; overrides the legacy GO_SOURCE root. authority: Prepared immutable authority to use without any database IO.
Returns:
Biolink category string
Examples:
>>> get_go_category_by_aspect("GO:0004096") # catalase activity 'biolink:MolecularActivity'
>>> get_go_category_by_aspect("GO:0006091") # generation of precursor metabolites 'biolink:BiologicalProcess'
- kg_microbe.utils.ontology_utils.get_ncbitaxon_adapter()
Return a lazily-resolved guarded NCBITaxon adapter. :return: adapter proxy.
- kg_microbe.utils.ontology_utils.get_ncbitaxon_category(ncbitaxon_id)
Return appropriate Biolink category for NCBITaxon terms.
NCBITaxon is a taxonomy, so all terms should be OrganismTaxon. This handles edge cases like NCBITaxon:1 (root).
- Return type:
str
Args:
ncbitaxon_id: NCBITaxon term ID (e.g., “NCBITaxon:1”)
Returns:
Biolink category string (always OrganismTaxon for NCBITaxon)
Examples:
>>> get_ncbitaxon_category("NCBITaxon:1") # root 'biolink:OrganismTaxon'
- kg_microbe.utils.ontology_utils.get_ontology_adapter(ontology)
Return an OAK adapter for one ontology, building its DB if needed.
The single entry point transforms should use. Never pass an
.owlpath toget_adapterdirectly: OAK treats that as a request to build, outside every guard this module provides.- Parameters:
ontology (
str) – One ofncbitaxon,chebi,go,ec.- Returns:
OAK adapter over the ontology’s SemSQL DB.
- Raises:
OntologyDbUnavailableError – If no usable DB could be produced.
- kg_microbe.utils.ontology_utils.get_pato_category(pato_term_id)
Return the Biolink category for a PATO term: always biolink:PhenotypicQuality (#1015).
- Return type:
str
- kg_microbe.utils.ontology_utils.get_uberon_category(uberon_term_id)
Return appropriate Biolink category for UBERON anatomical terms.
UBERON is an anatomy ontology, so all terms should be AnatomicalEntity. This handles edge cases where UBERON terms have multiple categories.
- Return type:
str
Args:
uberon_term_id: UBERON term ID (e.g., “UBERON:0000178”)
Returns:
Biolink category string (always AnatomicalEntity for UBERON)
Examples:
>>> get_uberon_category("UBERON:0000178") # blood 'biolink:AnatomicalEntity'
>>> get_uberon_category("UBERON:0001970") # bile 'biolink:AnatomicalEntity'
- kg_microbe.utils.ontology_utils.ontology_db_path(ontology)
Return the SemSQL DB path for an ontology, beside its OWL source.
- Parameters:
ontology (
str) – One ofncbitaxon,chebi,go,ec.- Return type:
str- Returns:
Absolute path to the
.db.- Raises:
KeyError – For an unknown ontology name.
- kg_microbe.utils.ontology_utils.replace_category_ontology(line, id_index, category_index, *, raw_dir=None, authority=None)
Replace node category according to prefix that has already been fixed.
- Parameters:
line (str) – A line from the original triples.
raw_dir (
Optional[Path]) – Selected producer authority root for imported GO references.authority – Optional prepared immutable GO authority.
- kg_microbe.utils.ontology_utils.replace_deprecated_categories(category_str)
Replace deprecated Biolink categories with current equivalents.
- Return type:
str
Args:
category_str: Category string (may be pipe-delimited)
Returns:
Updated category string with deprecated categories replaced
- kg_microbe.utils.ontology_utils.resolve_adapter(adapter)
Force a lazily-resolved adapter to resolve now; pass anything else through.
Use this at the top of a function that is about to open an output file or start a long loop. Resolution is otherwise deferred to the first attribute access, which can be well after work has begun — and a fatal failure there aborts mid-write. The atomic-write helper stops that from leaving a poisoned artifact, but surfacing the failure before any work starts is cheaper and reads better in a log.
- Parameters:
adapter – A lazy proxy, a real OAK adapter, or None.
- Returns:
The resolved adapter (or the argument unchanged).
kg_microbe.utils.optional_consumed_inputs module
Exact producer-read contracts for declared optional repository and effective-raw inputs.
- kg_microbe.utils.optional_consumed_inputs.has_optional_inputs(producer)
Recognize both explicit locator modes without relocating repository inputs.
- kg_microbe.utils.optional_consumed_inputs.optional_input_paths(producer, *, input_dir=None)
Resolve repository roles at the repository and raw roles at the selected lexical directory.
- kg_microbe.utils.optional_consumed_inputs.snapshot_optional_input(transform, name)
Read immutable captured text, or yield None only for explicitly observed optional absence.
- kg_microbe.utils.optional_consumed_inputs.verify_optional_inputs(producer, contract, snapshots, *, admission=None, input_dir=None)
Check declared membership, locator, original bytes or absence, retaining the same guards.
- kg_microbe.utils.optional_consumed_inputs.verify_recorded_optional_inputs(producer, report, *, report_path=None, admission=None)
Use the same contract for standalone, public and diagnostic recorded-source checks.
kg_microbe.utils.pandas_utils module
Pandas utilities.
- kg_microbe.utils.pandas_utils.drop_duplicates(file_path, sort_by_column='subject', dedup_on_sort_column=False)
Read TSV, drop duplicates, and export to the same file without making unnecessary copies.
- Parameters:
file_path (
Path) – Path to the TSV file.sort_by_column (
str) – Column name to sort the DataFrame.dedup_on_sort_column (
bool) – If True, also remove rows with duplicate sort_by_column values (keeping the first occurrence). Use for node files where the same ID may appear with different attributes.
- kg_microbe.utils.pandas_utils.dump_ont_nodes_from(nodes_filepath, target_path, prefix)
Dump CURIEs of an ontology for further processing.
- Parameters:
nodes_filepath (
Path) – Path of the nodes file.target_path (
Path) – Path where this list of CURIEs need to be exported.prefix (
str) – Prefix determines the CURIEs of interest.
- kg_microbe.utils.pandas_utils.establish_transitive_relationship(file_path, subject_prefix, intermediate_prefix, predicate, object_prefixes_list)
Establish transitive relationship given the predicate is the same.
- e.g.: Existent relations:
A => predicate => B
B => predicate => C
This function adds the relation A => predicate => C
- Parameters:
file_path (
Path) – Filepath of the edge file.subject_prefix (
str) – Subject prefix (A in the example)intermediate_prefix (
str) – Intermediate prefix that connects the subject to object (B in the example).predicate (
str) – The common predicate between all relations.object_prefixes_list (
List[str]) – List of Object prefixes (C in the example)
- Return type:
DataFrame- Returns:
Core dataframe with additional deduced rows.
- kg_microbe.utils.pandas_utils.establish_transitive_relationship_multiple(file_path, subject_prefix, intermediate_prefix_list, predicate_list, object_prefixes_lists)
Establish multiple transitive relationships via the establish_transitive_relationship function.
- e.g.: Existent relations:
A => predicate => B
B => predicate => C
C => predicate => D
This function adds the relation A => predicate => D
- Parameters:
file_path (
Path) – Filepath of the edge file.subject_prefix (
str) – Subject prefix (A in the example)intermediate_prefix_list (
list) – List of intermediate prefixes ([B,C] in the example).predicate_list (
list) – List of the common predicate between all relations. Len == intermediate_prefix_list.object_prefixes_list – List of Object prefixes ([[C,D]] in the example). Len == intermediate_prefix_list.
- Return type:
DataFrame- Returns:
Core dataframe with additional deduced rows.
- kg_microbe.utils.pandas_utils.get_ingredients_overlap(file_path, target_path)
Export TSV showing ingredient overlap between solutions and media.
- Parameters:
file_path (
Path) – Edges file pathtarget_path (
Path) – Output path.
kg_microbe.utils.parse_taxon_rank module
kg_microbe.utils.postprocess_artifacts module
Read-only content evidence for postprocess status; never certify freshness from timestamps.
- class kg_microbe.utils.postprocess_artifacts.ArtifactStatus(path, status, detail, archive_sha256=None, members=<factory>)
Bases:
objectExact selected artifact identity plus bounded streaming verification results.
-
archive_sha256:
str|None= None
-
detail:
str
-
members:
dict
-
path:
Path|None
-
status:
str
-
archive_sha256:
- kg_microbe.utils.postprocess_artifacts.inspect_artifact(path)
Inspect archive or loose manifest/pair without extracting or changing any graph files.
- kg_microbe.utils.postprocess_artifacts.review_status(skill, artifact, repo, review_dirs=())
Require archive-bound, report-byte-bound full pass receipts; prose alone is unverified.
- kg_microbe.utils.postprocess_artifacts.select_artifact(repo, *, merged_dir=None, archive=None)
Find graph artifacts themselves, refusing ambiguous siblings instead of using directory mtimes.
kg_microbe.utils.producer_audits module
Preserve mandatory producer disposition reports through finalization and merge.
- kg_microbe.utils.producer_audits.record_producer_audit(transform, name)
Record successful producer output once; finalization must not restamp changed bytes.
- kg_microbe.utils.producer_audits.required_producer_audits(producer)
Allow only distinct local sidecars, never graph or finalizer-owned members.
- kg_microbe.utils.producer_audits.verify_producer_audits(transform, *, output_dir=None)
Verify original producer identities against the original or staged report bytes.
- kg_microbe.utils.producer_audits.verify_recorded_producer_audits(producer, report, report_path, admission=None)
Enforce current registered requirements even if a report omits its own evidence.
kg_microbe.utils.provenance module
Separate scalar information-resource provenance from public record evidence.
Biolink 4.4.2 explicitly includes public web pages in publications. BacDive
record pages belong there, while primary_knowledge_source remains a single
information resource. Legacy source/record list and pipe cells are migrated
without inventing a new primary provider or duplicating ambiguous evidence.
- kg_microbe.utils.provenance.bacdive_record_url(record_id)
Return the public evidence page for a numeric BacDive record or its CURIE.
- Return type:
str
- kg_microbe.utils.provenance.knowledge_source_tokens(value, delimiter='|')
Flatten collections and pipe tokens, keeping first occurrence order.
- Return type:
list[str]
- kg_microbe.utils.provenance.primary_source_and_publications(value, publications=None)
Migrate unambiguous legacy BacDive attribution; refuse pooled primary providers.
- kg_microbe.utils.provenance.serialize_knowledge_sources(*values, delimiter='|')
Serialize explicitly multivalued provenance or publication tokens, never scalar PKS.
- Return type:
str
- kg_microbe.utils.provenance.validate_primary_source_and_publications(value, publications=None)
Validate canonical provider/evidence identity; permit only KGX collection representation.
kg_microbe.utils.review_claims module
Complete-payload and owning-record evidence checks shared by reviewed exports.
- kg_microbe.utils.review_claims.payload_sha256(payload)
Hash the entire reviewed payload, including structured source qualifiers.
- Return type:
str
- kg_microbe.utils.review_claims.validate_review_decision(decision, payload, owner, owner_sha256, verified, proofs, *, require_owner_path=False)
Verify a complete claim against the existing owner/evidence review format.
Callers independently resolve the owning record and verify every input’s digest before supplying its bytes. This verifier is shared by mapping and companion-claim exports; a checksum or receipt alone confers no approval.
- Return type:
None
kg_microbe.utils.robot_utils module
Utility to implement ROBOT over ontology files.
- kg_microbe.utils.robot_utils.convert_to_json(path, ont)
Convert OWL to JSON using ROBOT and the subprocess library.
- Parameters:
path (
str) – Path to ROBOT and the input OWL files.ont (
str) – Ontology
- Returns:
None
- kg_microbe.utils.robot_utils.extract_convert_to_json(path, ont_name, terms, mode)
Extract all children of provided CURIE.
ROBOT Method options:
STAR: The STAR-module contains mainly the terms in the seed and the
inter-relations between them (not necessarily sub- and super-classes).
TOP: The TOP-module contains mainly the terms in the seed, plus all
their sub-classes and the inter-relations between them.
BOT: The BOT, or BOTTOM, -module contains mainly the terms in the seed,
plus all their super-classes and the inter-relations between them.
MIREOT : The MIREOT method preserves the hierarchy of the input ontology
(subclass and subproperty relationships), but does not try to preserve the full set of logical entailments.
- Parameters:
path (
str) – path of file to be convertedont_name (
str) – Name of the ontologyterms (
Union[str,Path]) – Either CURIE or a file of CURIEs listmode (
str) – Method options as listed below.
- Returns:
None
- kg_microbe.utils.robot_utils.initialize_robot(path)
Initialize ROBOT with necessary configuration.
- Parameters:
path (
str) – Path to ROBOT files.- Return type:
list- Returns:
A list consisting of robot shell script name and environment variables.
- kg_microbe.utils.robot_utils.remove_convert_to_json(path, ont_name, terms)
Remove all children of provided CURIE(s).
- Parameters:
path (
str) – path of file to be convertedont_name (
str) – Name of the ontologyterms (
Union[List,Path]) – Either CURIE or a file of CURIEs list.
- Returns:
None
kg_microbe.utils.sanitize_curies module
Sanitize CURIEs and URIs in KGX TSV files.
- kg_microbe.utils.sanitize_curies.robust_fix_uri(val)
Fix URIs by properly encoding path components while preserving structure.
- kg_microbe.utils.sanitize_curies.sanitize_id_or_uri(val)
Sanitize values that could become URIs, including node IDs.
- kg_microbe.utils.sanitize_curies.sanitize_row(row)
Sanitize a single row of TSV data.
- kg_microbe.utils.sanitize_curies.sanitize_tsv(input_file, output_file)
Sanitize a TSV file by fixing line endings and URI encoding.
kg_microbe.utils.source_finalization module
Explicit, audited source normalization before source fingerprints and KGX merge (#1082).
- exception kg_microbe.utils.source_finalization.SourceFinalizationRequired
Bases:
ValueErrorA graph contains noncanonical source data which merge must not repair.
- kg_microbe.utils.source_finalization.finalize_selected_sources(transform, output_dirs, *, fresh_run=False)
Validate all explicitly selected dataset bundles before publishing any finalizer changes.
- kg_microbe.utils.source_finalization.finalize_source(transform, *, file_prefix='', fresh_run=False)
Finalize an explicit source bundle, preserving prior audit evidence on exact repeat calls.
- kg_microbe.utils.source_finalization.graph_rows(path, *, quoting=3)
Read the explicitly selected TSV dialect, preserving literal KGX quotes.
- kg_microbe.utils.source_finalization.snapshot_consumed_input(transform, name, path)
Yield immutable UTF-8 text with literal newlines; stream-copy/hash before parsing.
- kg_microbe.utils.source_finalization.validate_graph_bundle(node_paths, edge_paths, *, require_closure=False, require_contract=False)
Check already finalized records; never remap identities, categories or assertions.
- kg_microbe.utils.source_finalization.validate_identifier(identifier)
Reject known noncanonical representations without guessing unknown URI identities.
- kg_microbe.utils.source_finalization.validate_node_representation(identifier, category)
Enforce representation invariants without loading raw ontology authorities at merge.
- kg_microbe.utils.source_finalization.verify_consumed_inputs(transform, *, native_output_dir=None)
Require every declared producer read and verify its immutable path/digest snapshot.
- kg_microbe.utils.source_finalization.verify_finalized_source_files(paths)
Require exact prepared graph bytes and unchanged authority inputs before public merge.
- kg_microbe.utils.source_finalization.write_merge_validation_report(node_path, edge_path, report_path)
Validate a merged pair without semantic edits; keep the legacy audit member name explicit.
kg_microbe.utils.sssom_identity_policy module
Route SSSOM rows before they can contribute chemical identity lookups.
- kg_microbe.utils.sssom_identity_policy.classify_mapping_row(row, metadata=None)
Return an explicit consumer route and a diagnostic reason for a row.
The legacy unified format has two lexical row shapes and a separately tagged recipe-equivalent hydrate relation. Neither exception authorizes arbitrary closeMatch rows to enter the identity xref index.
- Return type:
tuple[str,str]
kg_microbe.utils.string_coding module
Decoding and encoding strings such that everything is utf8.
- kg_microbe.utils.string_coding.process_and_decode_label(label)
Process and decode a label string.
- Parameters:
label – A string to process and decode.
- Returns:
A processed and decoded string.
- kg_microbe.utils.string_coding.remove_nextlines(input_str)
Clean string by removing nextlines.
kg_microbe.utils.stub_curie_collection module
Collect stub-prefix CURIEs referenced anywhere in the mapping TSVs.
KG-Microbe deliberately does NOT load the full NCIT or MESH ontologies (those
belong to the sibling kg-microbe-biomedical pipeline), but the chemical-mapping
consolidator and the BacDive isolation-source mapper reference a small handful
of NCIT and MESH IDs. This collector finds every such CURIE so that the
downstream OntologiesStubsTransform
can fetch a labelled stub node for each one.
It scans a fixed set of mapping files at the repo root (no glob magic — wrong edits silently change the import set, so the file list is explicit and auditable):
mappings/kgmicrobe_unified_entity_mappings.sssom.tsv.gz— unified chemical/anatomy/environment mappings (object_id, subject_id columns).mappings/isolation_source_to_ontology.tsv— BacDive isolation-source mappings (object_id column).mappings/ingredient_mappings.sssom.tsv— vendored MIM SSSOM (object_id, subject_id).mappings/canonical/*.tsv— chemical/enzyme/pathway/phenotype canonical exports (object_id).
Returned dict shape: {normalized_prefix: {curie, curie, ...}} where
normalized_prefix matches the case used in
STUB_ONTOLOGY_PREFIXES
(e.g. "NCIT" is uppercase, "mesh" is lowercase). Inputs in any case
are accepted and normalized.
- kg_microbe.utils.stub_curie_collection.collect_stub_curies(prefixes, mapping_paths=None)
Scan the mapping TSVs and return the set of CURIEs that match each requested prefix.
- Parameters:
prefixes (
Iterable[str]) – Iterable of CURIE prefixes to collect. Case-insensitive on input; the returned dict’s keys preserve the case as given here, so callers should pass them in the canonical form they want ("NCIT","mesh", …).mapping_paths (
Optional[Iterable[Path]]) – Override the file list (mainly for tests). Defaults toDEFAULT_MAPPING_PATHS.
- Return type:
Dict[str,Set[str]]- Returns:
{canonical_prefix: {curie, ...}}for every prefix inprefixes, with the empty set as default for prefixes that have no references in any mapping file.
kg_microbe.utils.transform_fingerprint module
Record which code and curation data produced a transform’s output.
Freshness detection has been timestamp-based, and timestamps do not survive routine git operations. Two failure modes, both observed:
git checkoutrewrites a tracked file’s mtime with no content change, so visiting another branch flips the verdict (#797). Moving to commit time fixed that one.A squash merge mints a new commit for content that already existed, so commit time jumps forward while the bytes stay identical (#836). #832 squashed at 19:30 for code the gold transform had already run against at 19:06, and the guard reported stale output that was byte-for-byte current.
Content is the only signal immune to both. A transform writes a fingerprint of its inputs beside its output; anything comparing them asks whether the bytes match rather than which timestamp is larger.
Code and data are fingerprinted separately so a stale output can still say
why — the distinction kgm-freshness-check reports as STALE_VS_CODE
versus STALE_VS_DATA.
- kg_microbe.utils.transform_fingerprint.FINGERPRINT_FILE = 'source_fingerprint.json'
Filename written beside
nodes.tsv/edges.tsv.
- kg_microbe.utils.transform_fingerprint.FINGERPRINT_VERSION = 3
Bumped when the hashing scheme changes, so an old marker is treated as absent rather than silently compared under different rules.
- kg_microbe.utils.transform_fingerprint.SCHEMA_FILES = (PosixPath('data/raw/biolink-model.yaml'), PosixPath('data/raw/attributes.yaml'), PosixPath('data/raw/predicate_mapping.yaml'))
The pinned Biolink schema every transform validates against, relative to the repository root. All three move together (see CLAUDE.md).
- kg_microbe.utils.transform_fingerprint.SHARED_CODE = (PosixPath('kg_microbe/utils'), PosixPath('kg_microbe/transform_utils/constants.py'), PosixPath('kg_microbe/transform_utils/transform.py'), PosixPath('kg_microbe/transform.py'), PosixPath('kg_microbe/merge_utils/external_node_closure.py'), PosixPath('kg_microbe/merge_utils/local_context.py'))
First-party code every transform runs through besides its own package. A change here changes outputs just as much as a change in the package, and most behaviour-changing PRs land here (#1002).
- kg_microbe.utils.transform_fingerprint.bounded_file_fingerprint(path)
Return a content-sensitive fingerprint with bounded IO for large files.
Small files are hashed in full. For larger files, the digest covers the size and eight evenly spaced 128 KiB windows (including both ends). This is intentionally stronger than size/mtime metadata while keeping database startup independent of graph size. It is a cache-invalidation signal, not a cryptographic proof that two multi-gigabyte files are identical.
- Parameters:
path (
Path) – File to fingerprint.- Return type:
str- Returns:
Versioned SHA-256 digest string.
- kg_microbe.utils.transform_fingerprint.code_fingerprint(code_dir, repo_root=None, code_inputs=())
Fingerprint every Python file in a transform’s package directory.
Directory rather than the single module, matching what the freshness check already treats as “this transform’s code”: several transforms are split across helper modules in the same package, and a change to one of those changes the output just as much.
Names are folded in repo-relative, so the same package under two checkouts is one fingerprint (#983).
- Parameters:
code_dir (
Path) – e.g.kg_microbe/transform_utils/gold.repo_root (
Optional[Path]) – Repository root; inferred from this module when omitted.code_inputs (
Iterable[str]) – Additional repo-relative Python packages/files used only by this producer.
- Return type:
str- Returns:
Hex digest, or the digest of nothing when the directory is absent.
- kg_microbe.utils.transform_fingerprint.data_fingerprint(repo_root, data_inputs, input_dir=None)
Fingerprint a transform’s declared curation inputs.
- Parameters:
repo_root (
Path) – Repository root, whichDATA_INPUTSare relative to.data_inputs (
Iterable[str]) – Repo-relative paths fromTransform.DATA_INPUTS.input_dir (
Optional[Path]) – Effective raw input directory fordata/rawdeclarations.
- Return type:
str- Returns:
Hex digest.
- kg_microbe.utils.transform_fingerprint.declared_data_inputs(transform_class, repo_root)
Resolve static and discovered inputs now, never freezing a glob at class import.
- Return type:
tuple[str,...]
- kg_microbe.utils.transform_fingerprint.finalization_inputs_current(recorded, repo_root)
Verify exact consumed authorities; missing or changed bytes invalidate finalized output.
- Return type:
bool
- kg_microbe.utils.transform_fingerprint.migrate_markers(transformed_dir, repo_root, sources)
Rewrite scheme-2 markers that still vouch for their output as scheme 3.
Two phases, because upstream digests read other markers: first judge every scheme-2 marker against its inputs under scheme 2, while all markers are still scheme 2; then rewrite the ones that were FRESH, in dependency order, so each downstream’s new upstream digest sees its upstreams’ new markers.
- Parameters:
transformed_dir (
Path) –data/transformed.repo_root (
Path) – Repository root.sources (
Iterable[dict]) – One dict per registered source:name,output_dir(the directory undertransformed_dir),code_dir,data_inputs,transform_inputs; optionalrequires_dependency_rebuildrefuses migration when older evidence did not bind discovered/inherited inputs.
- Return type:
Dict[str,str]- Returns:
{source name: "migrated" | "left: <reason>"}.
- kg_microbe.utils.transform_fingerprint.read_fingerprint(output_dir)
Read a recorded fingerprint, if one is present and readable.
A marker from a different scheme version, or one that will not parse, reads as absent. Callers fall back to their timestamp comparison in that case, which is weaker but defined — better than asserting a mismatch on a marker we cannot interpret.
- Parameters:
output_dir (
Path) – Directory holding the transform’s output.- Return type:
Optional[dict]- Returns:
The payload, or None.
- kg_microbe.utils.transform_fingerprint.resolve_data_input(repo_root, declaration, input_dir=None)
Resolve raw declarations under the effective CLI raw directory; keep curation repo-relative.
- Return type:
Path
- kg_microbe.utils.transform_fingerprint.schema_fingerprint(repo_root)
Identify the Biolink schema a transform ran against.
The pinned model is a real pipeline input –
prepare_kgxmakes it the default schema for every KGXToolkit– but no transform declares it, so swapping it (as #941 did, 4.3.6 -> 4.4.2) marked nothing stale and a partial rerun could mix schema versions in one merged graph (#943). Record it in the marker so the artifact says which schema produced it.- Parameters:
repo_root (
Path) – Repository root.- Return type:
Optional[dict]- Returns:
{"version": ..., "digest": ...}, or None when no schema file is on disk – unknown provenance is recorded as unknown, not invented.
Fingerprint the first-party code every transform shares.
code_fingerprintsees only the transform’s own package, so a change inkg_microbe/utils/– the mapping loaders, the ontology adapters, the chemical-mapping utils – or inconstants.pymarked nothing stale: #999 changed three transforms’ predicates and the freshness table stayed FRESH (#1002). Blunt by design: an edit here reads as “every output may differ”, which is the honest answer.- Parameters:
repo_root (
Optional[Path]) – Repository root; inferred from this module when omitted.- Return type:
str- Returns:
Hex digest.
- kg_microbe.utils.transform_fingerprint.upstream_fingerprint(output_base_dir, transform_inputs)
Fingerprint the recorded state of every upstream transform.
Folds in each upstream’s own marker rather than its output TSVs, which can be hundreds of megabytes. If an upstream re-runs, its marker changes and every downstream goes stale — which is the point (#845).
An upstream with no marker folds in as absent, so a downstream goes stale exactly once when that upstream first records one. That is the correct direction to fail: it prompts a re-run rather than asserting currency that was never established.
- Parameters:
output_base_dir (
Path) –data/transformed.transform_inputs (
Iterable[str]) – Registered source names read by this transform.
- Return type:
str- Returns:
Hex digest.
- kg_microbe.utils.transform_fingerprint.write_fingerprint(output_dir, code_dir, repo_root, data_inputs, transform_inputs=(), input_dir=None, finalization_inputs=(), verify_inputs=None, code_inputs=())
Record the fingerprint of a completed run.
Call after the outputs are written, so a run that dies partway leaves no marker claiming its output matches the current inputs. Written through
atomic_writefor the same reason a torn marker would be worse than none.- Parameters:
output_dir (
Path) – Where the transform wrote its TSVs.code_dir (
Path) – The transform’s package directory.repo_root (
Path) – Repository root.data_inputs (
Iterable[str]) – Repo-relative curation paths.transform_inputs (
Iterable[str]) – Registered sources whose output this one reads.input_dir (
Optional[Path]) – Effective raw directory read by the completed transform.finalization_inputs (
Iterable[str]) – Exact authority/dependency paths consumed by source finalization.verify_inputs (
Optional[Callable[[],None]]) – Optional producer-time snapshot verifier. Raises if consumed bytes changed; checked before hashing and immediately before atomic marker publication.code_inputs (
Iterable[str]) – Producer-specific inherited Python code, using the same AST semantics.
- Return type:
dict- Returns:
The recorded payload.
kg_microbe.utils.trembl_utils module
kg_microbe.utils.tsv_io module
One place that decides how KG-Microbe writes a TSV.
csv.writer’s default lineterminator is "\r\n" on every platform,
and it is emitted literally regardless of the newline argument to
open(). Four transforms and both merge normalizers used that default, so
every line of the shipped merged graph ended with a carriage return and the
last column of every row carried it as data – an “empty” trailing field was
"\r", which inverts a truthiness test on it (#1041, and the corruption
open issue #541 describes in the DuckDB loader).
Use tsv_writer() instead of constructing csv.writer directly for any
file another tool will parse. tests/test_tsv_line_endings.py fails on a
graph-writing module that goes back to the bare constructor.
- kg_microbe.utils.tsv_io.LINE_TERMINATOR = '\n'
The only line terminator KG-Microbe writes.
- kg_microbe.utils.tsv_io.tsv_dict_writer(handle, fieldnames, **kwargs)
Return a tab-delimited
csv.DictWriterterminating lines with\n.- Parameters:
handle (
Any) – A file opened withnewline="".fieldnames – Column names, in order.
kwargs (
Any) – Passed through tocsv.DictWriter.
- Returns:
The configured writer.
- kg_microbe.utils.tsv_io.tsv_writer(handle, **kwargs)
Return a tab-delimited
csv.writerthat terminates lines with\n.- Parameters:
handle (
Any) – A file opened withnewline=""(so the text layer does not translate what the writer emits).kwargs (
Any) – Passed through tocsv.writer;delimiterandlineterminatordefault to tab and newline.
- Returns:
The configured writer.
kg_microbe.utils.unipathways_utils module
Unipathways utilities.
- kg_microbe.utils.unipathways_utils.check_wanted_pairs(line, subject_index, object_index)
Check if subject object pair should be included.
- Parameters:
line (str) – A line from the original triples.
subject_index (int) – The index of the tab delimited line with the triple subject.
object_index (int) – The index of the tab delimited line with the triple object.
- kg_microbe.utils.unipathways_utils.create_df_from_pair(df, pair, subject_node=None)
Create a dataframe from a given dataframe according to substrings in a given pair.
- Parameters:
df (pd.DataFrame) – A dataframe that contains all triples.
pair (List) – A list of the subject, object of the desired triple pattern.
subject_node (str) – Optional, a specific subject node ID to base the search on.
- kg_microbe.utils.unipathways_utils.get_key_from_value(dictionary, value)
Extract a key from a dictionary with the corresponding value.
- kg_microbe.utils.unipathways_utils.get_unipathways_prefix(id)
Get unipathways prefix of a given node ID if available.
- Parameters:
id – The node ID.
- kg_microbe.utils.unipathways_utils.project_onto_header(parts, source_header, node_header)
Reshape one row onto
node_header, matching columns by name.The caller writes
node_headeras the file’s header, so every row has to be exactly that wide. Padding positionally is not enough: KGX leakssubsets,metaandirionto node rows (the columns_normalize_schemastrips), and a row wider than the header made[""] * (len(node_header) - len(parts))evaluate to[]– so the extra fields survived and the file became unreadable (#1033). Matching by name rather than truncating also means an extra, missing or reordered upstream column cannot shift a value into the wrong field.- Parameters:
parts – The row’s fields.
source_header – Column names of the row, in order.
Nonefalls back to positional pad-or-truncate, for callers that never saw a header.node_header – Canonical column names to emit.
- Returns:
Exactly
len(node_header)fields.
- kg_microbe.utils.unipathways_utils.remove_unwanted_prefixes_from_edges(df)
Remove unwanted prefixes that exist in all triple.
- Parameters:
df (pd.DataFrame) – A dataframe of all triples.
- kg_microbe.utils.unipathways_utils.remove_unwanted_prefixes_from_node_xrefs(line, xref_index)
Remove unwanted prefixes that exist in xrefs for given node.
- Parameters:
line (str) – A line from the original triples.
xref_index (int) – The index of the tab delimited line with the triple xref.
- kg_microbe.utils.unipathways_utils.replace_category_for_unipathways(line, id_index, category_index, node_header, source_header=None)
Replace category of a given node.
- Parameters:
line (str) – A line from the original triples.
id_index (int) – The index of the tab delimited line with the node id.
category_index (int) – The index of the tab delimited line with the node category.
source_header (list) – Column names of
line, so fields are matched by name rather than by position (#1033).
- kg_microbe.utils.unipathways_utils.replace_id_with_xref(line, xref_index, id_index, category_index, nodes_dictionary, node_header, source_header=None)
Replace node ID with corresponding xref.
- Parameters:
line (str) – A line from the original triples.
xref_index (int) – The index of the tab delimited line with the node xref.
id_index (int) – The index of the tab delimited line with the node id.
category_index (int) – The index of the tab delimited line with the node category.
node_header (list) – List of all values in nodes file header.
source_header (list) – Column names of
line, so fields are matched by name rather than by position (#1033).
- kg_microbe.utils.unipathways_utils.replace_triples_with_labels(line, subject_index, object_index, predicate_index, relation_index, nodes_dictionary)
Replace triples labels according to a dictionary lookup. Also replace the predicate and relation.
- Parameters:
line (str) – A line from the original triples.
subject_index (int) – The index of the tab delimited line with the triple subject.
object_index (int) – The index of the tab delimited line with the triple object.
predicate_index (int) – The index of the tab delimited line with the triple predicate.
relation_index (int) – The index of the tab delimited line with the triple relation.
kg_microbe.utils.uniprot_utils module
Uniprot utilities.
- kg_microbe.utils.uniprot_utils.check_string_in_tar(tar_file, uniprot_relevant_file_list, progress_class, regex_pattern='UP\\d+: (Chromosome|Plasmid .+)', min_line_count=1000)
Look for a specific string in tsvs in the tarfile and return the content of matching members.
- Parameters:
tar_file – The path to the tarfile containing the tsv files.
progress_class – The class to use for progress tracking. (tqdm or dummy)
regex_pattern – The regex pattern to search for in the tsv files.
min_line_count – The minimum number of lines that must match the pattern.
- kg_microbe.utils.uniprot_utils.convert_omim_diseases(omim_list, mondo_xref_dict)
Convert OMIM ID’s to a MONDO ID using MONDO xref dictionary.
This method uses the MONDO ontology xrefs to convert OMIM to MONDO.
- Parameters:
omim_list (list) – A list containing OMIM IDs.
mondo_xref_dict (dict) – A dictionary of MONDO IDs and their xrefs.
- Returns:
A list of MONDO curies.
- Return type:
list
- kg_microbe.utils.uniprot_utils.create_pool(source_name, tar_file, n_workers, chunk_size_denominator, show_status, node_header, edge_header, output_node_file, output_edge_file, go_category_trees_dict, mondo_xrefs_dict, mondo_gene_dict, obsolete_terms_csv_file, uniprot_relevant_file_list, uniprot_tmp_ne_dir)
Process tar file in parallel threads.
- kg_microbe.utils.uniprot_utils.get_go_category_trees(go_oi)
Extract category of all GO terms using oak, and write to file.
Written atomically, and the adapter is resolved before the file is opened. Both matter because the only guard on this cache is a bare
.exists()in the two UniProt transforms: the previous in-place write emitted the header, then resolved the adapter on the firstdescendantscall, so an unusable go.db left a header-only file behind. That file exists, so it is never regenerated, andprepare_go_dictionaryreads it as{}— dropping every protein→GO edge and logging every GO term as obsolete on every run thereafter, silently. A partial write is worse still: MF/CC survive while every BP term is reclassified as obsolete.- Parameters:
go_oi (oaklib sql_implementation class) – A oaklib sql_implementation class to access GO information.
- kg_microbe.utils.uniprot_utils.get_go_relation_and_obsolete_terms(term_id, uniprot_id, go_category_trees_dictionary, obsolete_terms_csv_file)
Extract category of GO term and handle obsolete terms according to oak.
- Parameters:
term_id (str) – A string containing the GO ID.
uniprot_id (str) – A string containing the Uniprot protein ID.
go_category_dictionary (dict) – Dictionary of all GO Terms with their corresponding category.
- Returns:
The appropriate predicate for the GO ID as it relates to the protein ID, or None if obsolete.
- Return type:
str or None
- kg_microbe.utils.uniprot_utils.get_nodes_and_edges(source_name, uniprot_df, go_category_trees_dictionary, mondo_xrefs_dict, mondo_gene_dict, obsolete_terms_csv_file)
Process UniProt entries and writes organism-enzyme relationship data to CSV files.
This method iterates over a list of UniProt entries, extracts relevant information, and writes it to two separate CSV files using the provided CSV writers. One file contains edges representing relationships between organisms and enzymes, and the other contains nodes representing enzymes. It also handles binding site information by calling parse_binding_site method if available in the entry.
- Parameters:
uniprot_df (dataframe) – A dataframe where each row represents a UniProt entry.
- kg_microbe.utils.uniprot_utils.go_category_trees_is_complete(path=None)
Report whether the GO category-trees cache is present and has content.
- Parameters:
path – Cache path; defaults to GO_CATEGORY_TREES_FILE.
- Return type:
bool- Returns:
True if the cache exists and holds at least one term.
- kg_microbe.utils.uniprot_utils.is_float(entry)
Determine if value is float, returns True/False.
- kg_microbe.utils.uniprot_utils.parse_binding_site(binding_site_entry)
Extract chemical identifiers from a binding site entry.
This method uses regular expressions to find all occurrences of ligand IDs within a given binding site entry string. It specifically looks for ChEBI identifiers and returns a list of these identifiers found in the entry.
- Parameters:
binding_site_entry (str) – A string containing the binding site information.
- Returns:
A list of ChEBI ligand identifiers extracted from the binding site entry.
- Return type:
list
- kg_microbe.utils.uniprot_utils.parse_disease(disease_entry, mondo_xref_dict)
Extract OMIM ID’s from a disease entry.
This method uses regular expressions to find all occurrences of OMIM IDs.
- Parameters:
disease_entry (str) – A string containing the disease information.
- Returns:
A list of OMIM curies.
- Return type:
list
- kg_microbe.utils.uniprot_utils.parse_ec(ec_entry)
Extract ec identifiers from a EC Number entry.
This method finds all EC IDs for the corresponding entry.
- Parameters:
ec_entry (str) – A string containing the ec information.
- Returns:
A list of EC identifiers extracted from the EC entry.
- Return type:
list
- kg_microbe.utils.uniprot_utils.parse_gene(gene_entry, mondo_gene_dict)
Get gene ID from gene name entry.
This method uses the MONDO ontology transform to get the HGNC id of a gene.
- Parameters:
gene_entry (str) – A string containing the gene name.
- Returns:
A gene id.
- Return type:
str
- kg_microbe.utils.uniprot_utils.parse_go_entry(go_entry)
Extract chemical identifiers from a binding site entry.
This method uses regular expressions to find all occurrences of ligand IDs within a given binding site entry string. It specifically looks for ChEBI identifiers and returns a list of these identifiers found in the entry.
- Parameters:
binding_site_entry (str) – A string containing the binding site information.
- Returns:
A list of ChEBI ligand identifiers extracted from the binding site entry.
- Return type:
list
- kg_microbe.utils.uniprot_utils.parse_rhea_entry(rhea_entry)
Extract rhea identifiers from a rhea ID entry.
This method finds all RHEA IDs for the corresponding entry.
- Parameters:
rhea_entry (str) – A string containing the rhea information.
- Returns:
A list of RHEA identifiers extracted from the rhea entry.
- Return type:
list
- kg_microbe.utils.uniprot_utils.prepare_go_dictionary()
Generate dictionary of GO categories for each GO term.
- kg_microbe.utils.uniprot_utils.prepare_mondo_dictionary()
Generate dictionaries for disease xrefs and gene names from MONDO.
- kg_microbe.utils.uniprot_utils.process_lines(source_name, all_lines, headers, node_header, edge_header, node_filename, edge_filename, progress_class, go_category_dictionary, mondo_xrefs_dict, mondo_gene_dict, obsolete_terms_csv_file)
Process a member of a tarfile containing UniProt data.
This method reads the content of a member file, processes the data, and writes the resulting node and edge data to the provided CSV files. It uses a lock to ensure that the files are written to correctly and that no data is lost.
- Parameters:
all_lines – A list of lines from the member file.
headers – A list of headers for the data.
node_header – A list of headers for the node data.
edge_header – A list of headers for the edge data.
node_filename – The name of the file to write node data to.
edge_filename – The name of the file to write edge data to.
progress_class – The class to use for progress tracking. (tqdm or dummy)
go_category_dictionary – Dictionary of all GO Terms with their corresponding category.
- kg_microbe.utils.uniprot_utils.write_obsolete_file_header(obsolete_terms_csv_file)
Write obsolete header to file.
kg_microbe.utils.validation_utils module
Validation utilities for knowledge graph construction.
- kg_microbe.utils.validation_utils.get_curie_id(curie)
Extract ID from a CURIE.
- Parameters:
curie (
str) – CURIE string- Return type:
Optional[str]- Returns:
ID (e.g., “12345” from “CHEBI:12345”) or None if invalid
- kg_microbe.utils.validation_utils.get_curie_prefix(curie)
Extract prefix from a CURIE.
- Parameters:
curie (
str) – CURIE string- Return type:
Optional[str]- Returns:
Prefix (e.g., “CHEBI” from “CHEBI:12345”) or None if invalid
- kg_microbe.utils.validation_utils.load_valid_kgm_terms(custom_curies_path=None)
Load valid KG-Microbe custom terms from custom_curies.yaml.
- Parameters:
custom_curies_path (
Optional[Path]) – Path to custom_curies.yaml file If None, uses default path relative to this file- Return type:
Set[str]- Returns:
Set of valid CURIEs (e.g., {“kgmicrobe.trait:voges_proskauer_test_positive”})
- kg_microbe.utils.validation_utils.validate_curie(curie)
Validate that a string follows CURIE format (PREFIX:ID).
Valid CURIEs must: - Start with a letter - Have a prefix of alphanumeric characters and underscores - Have a colon separator - Have a non-empty ID part
- Parameters:
curie (
str) – String to validate- Return type:
bool- Returns:
True if valid CURIE format, False otherwise
- kg_microbe.utils.validation_utils.validate_curie_prefix(curie, allowed_prefixes)
Validate that a CURIE uses an allowed prefix.
- Parameters:
curie (
str) – CURIE to validateallowed_prefixes (
Set[str]) – Set of allowed prefixes (e.g., {“CHEBI”, “GO”, “METPO”})
- Return type:
bool- Returns:
True if CURIE prefix is in allowed set, False otherwise
- kg_microbe.utils.validation_utils.validate_kgm_term(curie, valid_kgm_terms=None)
Validate that a KG-Microbe custom term exists in custom_curies.yaml.
- Parameters:
curie (
str) – CURIE to validate (e.g., “kgmicrobe.trait:voges_proskauer_test_positive”)valid_kgm_terms (
Optional[Set[str]]) – Pre-loaded set of valid terms If None, will load from custom_curies.yaml
- Return type:
bool- Returns:
True if the term is valid, False otherwise
Module contents
ROBOT utility.