Skip to content
Documentation for triage-pg 1.1.4 — the current stable release. Release notes

FAQ

The questions that come up over and over, and where each answer lives in the design. Every answer here traces to a concrete behaviour — an ADR, a CLI flag, or a config rule — not folklore.


Why did my run produce 0-entity (empty) matrices?

Section titled “Why did my run produce 0-entity (empty) matrices?”

Almost always a cache-identity leak: the cohort and labels artifacts must include the as_of_dates they were built for in their identity. If those dates don’t enter the hash, a later run over a different temporal grid can silently cache-hit a cohort/labels artifact that was populated at other dates. The matrix then inner-joins your entities against a cohort that is empty for the dates you asked about — and you get zero rows.

This is exactly why forward scoring and retraining fold the single scoring date back into the cohort/label config before building (see triage.adapters.forward.dated_config): the experiment path pins its dates upstream in the temporal config and deliberately excludes them from cohort/label identity, but a one-off date is not derivable from the config, so it must enter identity or the inner join comes back empty. If you’re assembling artifacts outside the normal run path, make sure the dates are part of what you hash.

See identity and caching and point-in-time correctness.


Because predictions are append-only. Every scoring run inserts rows carrying a scored_at wall-clock timestamp; nothing is ever overwritten, and the table is time-partitioned. Re-scoring the same model at the same as_of_date doesn’t replace the old row — it adds a new one. So (model, entity, as_of_date) is not unique, and a naive read gives you whichever row the planner returned, not the newest.

To read “current”, pick the maximum scored_at per key:

select distinct on (model_id, entity_id, as_of_date)
model_id, entity_id, as_of_date, score, scored_at
from triage.predictions
order by model_id, entity_id, as_of_date, scored_at desc

This is a concept to teach, not a bug: “a score is not the latest score.” The payoff is that prediction history, drift, and trajectories are already recorded and become a GROUP BY later, with no migration. See the data model.


I get ValueError: Cannot resolve a distribution for estimator module 'triage'

Section titled “I get ValueError: Cannot resolve a distribution for estimator module 'triage'”

The estimator library’s version enters model identity, so engine_versions_for('model', …) has to reverse-map the estimator’s import package to its installed distribution. It does that via importlib.metadata.packages_distributions(), which reads top_level.txt / RECORD — and a PEP 660 editable install (a plain uv sync of this project, as in CI) frequently omits those, so triage-pg’s own triage.* estimators resolve to nothing even though the distribution is installed.

Current triage-pg handles this: when the reverse-map is empty it falls back to the module name as the distribution name (which holds for triage, sklearn, sksurv, …), producing the same (name, version) a populated reverse-map would. If you hit this error, you’re on an older build — update and re-run uv sync. See identity and caching.


Put its dotted class_path under grid_config, with each hyperparameter as a list — the run sweeps the cartesian product. Any sklearn.* estimator class works directly, and triage-pg ships triage.component.catwalk.estimators.classifiers.ScaledLogisticRegression (a minmax-scaler + logistic regression, so the persisted coefficients sit on comparable [0, 1]-scaled features):

grid_config:
'sklearn.ensemble.RandomForestClassifier':
n_estimators: [10, 100]
max_depth: [3, 5]
'triage.component.catwalk.estimators.classifiers.ScaledLogisticRegression':
C: [0.01, 1.0]
penalty: ['l2']
max_iter: [1000]

Any importable class with a scikit-learn-style fit/predict interface is fair game — the class path is resolved at train time. See the configuration reference.


feature_groups must be nested under feature_config. A top-level feature_groups: key at the root of the experiment config is not read by the adapter — it’s silently ignored, so you get today’s default behaviour (one implicit group, one Run) with no error. The config validator now emits a warning when it sees a stray top-level feature_groups, but the fix is to move it inside feature_config:

feature_config:
# … your feature definitions …
feature_groups:
group_by: source_entity
strategies: [all, leave-one-out, leave-one-in]

The adapter strips feature_groups back out before featurizer (or the feature_group node identity) ever sees it — featurizer stays group-agnostic ; grouping is a triage-pg concern over featurizer’s columns. See the configuration reference.


My run aborts on a feature column ending in ~1c0ca6e8.

Section titled “My run aborts on a feature column ending in ~1c0ca6e8.”

The message reads “feature column …frecuen~1c0ca6e8 matches no feature_groups.definitions glob”, and its advice — add a glob or widen one — looks impossible to follow. It is: that suffix is a content hash you cannot predict or type.

What happened is that the generated feature name was longer than PostgreSQL’s 63-byte identifier limit, so featurizer hash-truncated it from the tail. A glob aimed at a fragment that sits past the cut ("*frecuencia_cardiaca*") still describes the feature correctly, but no longer appears in the physical column name.

You do not need to work around it. Globs are matched against each column’s full, untruncated label as well as its physical name, so write the glob against the label and it will match. To confirm before running anything:

Terminal window
$ triage analyze-config experiment.yaml --features '*frecuencia_cardiaca*'

That prints exactly the columns partitioning would group (same matching rule), showing truncated ones as label → column. --features '*' lists every column. Neither needs a database.

Two related notes. group_by: source_entity — the default — is never affected: truncation keeps the head of the name, so the leading <alias>. token always survives. And a glob you previously hand-wrote against a truncated name still works; matching the label is additional, not a replacement.


triage score scattered Parquet files into my cron directory.

Section titled “triage score scattered Parquet files into my cron directory.”

Fixed. --project-path (the matrix output root) now defaults to the model’s own artifact root — the parent of its recorded artifact_uri — so a bare scheduled line like triage score --model-id 42 --as-of-date 2026-07-01 writes its production matrix beside the model’s existing artifacts, not into whatever CWD the scheduler happened to hand it. The same default applies to retrainpredict, so retrained artifacts land next to the originals.

Pass --project-path only when you actually want to redirect output somewhere else. If you’re seeing Parquets in your cron directory, you’re on an older build where the default was the CWD — update.


The tutorial fails at db upgrade / can’t connect.

Section titled “The tutorial fails at db upgrade / can’t connect.”

Two usual causes:

  1. No connection file. The CLI needs to know how to reach the tutorial DB. Create the git-ignored dirtyduck-database.yaml and pass it with --dbfile (the Dirty Duckling tutorial has the exact contents). Without it, the CLI falls back to your ambient PG* env and connects to the wrong place — or nowhere.

  2. The DB is still loading. The tutorial container does a first-boot data load that takes a few minutes; connecting before it finishes fails. Wait until Postgres is actually accepting connections:

    Terminal window
    pg_isready -h 127.0.0.1 -p 5440
    # 127.0.0.1:5440 - accepting connections

    Only then run uv run triage --dbfile dirtyduck-database.yaml db upgrade. If port 5440 is taken, pick another and adjust dirtyduck-database.yaml to match.


Can triage-pg share a database with a project that already has schemas?

Section titled “Can triage-pg share a database with a project that already has schemas?”

Yes. Everything it creates lives in the triage schema — the 41 tables, the views, the enums, the metric functions — including alembic’s stamp table, triage.results_schema_versions. Nothing is written to public, and nothing is written to your schemas. Drop the schema and triage-pg is gone.

That last part was not true before v1.1.2: the stamp table was created unqualified, so PostgreSQL resolved it through the connecting role’s search_path and it landed in whatever schema came first there. On a role carrying search_path = raw, clean, ontology, … that meant a raw-schema table triage-pg had no business creating. Upgrading is automatic — triage db upgrade moves an old stamp into triage before running anything, so a database already at head stays at head and no migration replays. The manual equivalent, if you would rather do it yourself first:

alter table public.results_schema_versions set schema triage;

The registry control-plane database gets the same treatment in its own schema (registry.registry_schema_versions).


Why did my command connect to the tutorial database instead of my own?

Section titled “Why did my command connect to the tutorial database instead of my own?”

Because a database.yaml in the current directory outranks PG* / DATABASE_URLdocumented precedence, and the repo root ships example connection files. Run a command from there with a production environment loaded and the file still wins.

Since v1.1.2 that is at least loud: when a database.yaml is picked up while the environment names a different database, the CLI logs a warning naming both. To use the environment, run from a directory without a database.yaml — or name the file you mean with --dbfile, which is always honoured without complaint.