kg_microbe.transform_utils package

Subpackages

Submodules

kg_microbe.transform_utils.constants module

Constants for transform_utilities.

kg_microbe.transform_utils.constants.GTDB_NCBI_POOLING_REPORT = 'gtdb_ncbi_pooling_report.tsv'

Report of NCBI taxa that several GTDB taxa map onto (#883). Written on every gtdb run, empty or not.

kg_microbe.transform_utils.transform module

Transform utility module.

class kg_microbe.transform_utils.transform.Transform(source_name, input_dir=None, output_dir=None, nlp=False)

Bases: object

Parent class for transforms, that sets up a lot of default file info.

CODE_INPUTS: tuple = ()

Additional repo-relative Python packages/files whose behavior this producer inherits. This is code provenance, never a dependency on another producer’s graph outputs.

DATA_DIR = PosixPath('/home/runner/work/kg-microbe/kg-microbe/kg_microbe/transform_utils/data')
DATA_INPUTS: tuple = ()

Repo-relative curation files this transform reads, beyond its own data/raw/ download.

Declared so freshness tooling can tell that an output is stale against its data rather than only its code. Without it a mapping correction lands, every consumer keeps reporting FRESH, and a re-merge silently ships the old groundings: #778 corrected 16 isolation-source ids and #786 rewrote the unified chemical SSSOM, and the merged KG built afterwards still asserted 75 organisms isolated from a “Cell Line”, because nothing re-ran the transforms that read those files (#812).

Paths are relative to the repo root. Keep them tracked in git — the freshness check uses commit time, not mtime, because git checkout rewrites mtimes without changing content (#797).

List every curation file read, not a representative one. A partial declaration fails silently and looks identical to a complete one: ontologies_stubs declared 1 of the 11 files it read and was reported fresh after changes to the other ten (#839). Where the set comes from a constant, derive this from it rather than restating it.

DEFAULT_INPUT_DIR = PosixPath('/home/runner/work/kg-microbe/kg-microbe/kg_microbe/transform_utils/data/raw')
DEFAULT_OUTPUT_DIR = PosixPath('/home/runner/work/kg-microbe/kg-microbe/kg_microbe/transform_utils/data/transformed')
OPTIONAL_CONSUMED_INPUTS: tuple = ()

Named optional repository-relative reads; absence is evidence, never a required file. Unlike DATA_INPUTS these locators do not relocate to input_base_dir.

OPTIONAL_RAW_CONSUMED_INPUTS: tuple = ()

Named optional reads relative to the effective input_base_dir, including explicit absence. Kept separate so repository-relative optional inputs never silently relocate.

REQUIRED_CONSUMED_INPUTS: tuple = ()

Named generated inputs that must actually be read before fresh finalization. Unlike DATA_INPUTS, these are checked after upstream producers have run.

TRANSFORM_INPUTS: tuple = ()

Registered source names whose output this transform reads.

DATA_INPUTS covers curation files under mappings/. It does not cover a dependency on another transform’s output, and eight sources have one: gold reads ontologies/ncbitaxon_nodes.tsv (and refuses to run without it) plus ontologies_stubs/po_nodes.tsv, lpsn reads gtdb/nodes.tsv, lpsn_api and microbedecoder read lpsn/nodes.tsv, prego reads ontologies/, and bacdive, mediadive and metatraits reach ontologies/ through the NCBITAXON_NODES_FILE / CHEBI_NODES_FILE constants (metatraits_gtdb inherits metatraits’). Those last three went undeclared for months precisely because the path is spelled in constants.py rather than in the transform, which the guard below could not see until #1035.

Undeclared, re-running an upstream leaves every downstream genuinely stale while all three freshness signals report fresh — the #812 shape, across transforms rather than within one (#845). The ordering was already encoded in DATA_SOURCES comments (“Run gold after ontologies…”); this makes it machine-readable so the fingerprint can fold the upstream in.

tests/test_cross_transform_inputs.py derives the real dependency map from the source and fails on anything undeclared, because every previous version of this contract was opt-in and was forgotten (#812, #839, #876).

begin_consumed_inputs()

Start a real producer run with no inherited input-consumption claims.

begin_dependency_admission()

Bind discovered curation and inherited code before a producer loads either.

consume_input(name, path)

Read a generated UTF-8 input from an immutable, exact-byte tracked snapshot.

consume_optional_input(name)

Read one declared optional input immutably, retaining its locator and explicit absence.

property consumed_input_snapshots

Return copied named path/digest snapshots so callers cannot mutate recorded evidence.

finalize(*, file_prefix='', fresh_run=False)

Validate and finalize produced TSVs before publication as a current source.

The CLI invokes this after run and before writing the source fingerprint. Direct Python callers must invoke it explicitly after run; producer writes themselves are not a bundle transaction.

property optional_consumed_inputs

Return independent serialized state without replacing original read evidence.

pass_through(nodes_file, edges_file)

Copy nodes and edges files to output directory.

Parameters:
  • nodes_file (str) – nodes files to take from raw directory and put in transform directory

  • edges_file (str) – edges files to take from raw directory and put in transform directory

Return type:

None

run(data_file=None)

Run the transform.

Parameters:

data_file (Union[Path, None, str]) – Input data file, defaults to None

verify_consumed_inputs()

Reject missing required reads or changed/deleted bytes consumed by this run.

verify_declared_dependencies()

Refuse drift without replacing the original producer-time dependency snapshot.

Module contents

Transform utilities module.