Daily HealthCheck
LIVE DAG STATUSLoading current status…/S3 update time unavailable

THE CLIENT VIEW

What needs attention?

Current job stateRead from the S3 status feed every two minutes.

Client environments

Only environments evidenced by current DAG IDs are shown. Shared jobs are repeated where they may affect the same flows.

Ordered by current status

Running and successful jobs are current healthy signals. Failed jobs remain red until the S3 feed reports another state. “No runs” is shown separately and is not treated as a failure.

How to read this report
Failed

At least one DAG currently reports failure. Shared failures apply to every client environment.

Running

The DAG is currently executing and no longer shown as failed.

Healthy

Every listed job is successful or running.

Needs mapping

A job has no runs or lacks enough mapping data for a clear impact.

Job → flow → environment. The full DAG ID determines the environment. A DAG without a client environment, or with the `common` marker, is shared. It appears in the Shared jobs section and under every active client environment.

S3 is the current-state source. A job changes immediately when the status feed changes. Failure first detected and later transitions are retained in the local history database.

ClickUp tracks follow-through. Current failures link to one incident per environment and Primary Owner. When jobs recover, an open ticket stays in the follow-up queue until its owner documents the fix, verifies any backfill, and closes it.

Slack adds diagnostics. One minute after a new failure, the collector looks for the exact DAG alert and records its task, execution time, owner and error category. Raw error text remains local and in the private ClickUp incident.

Missing mappings stay explicit. Shared failures still appear everywhere when their description or functional tag is incomplete, but their user-facing impact remains marked as mapping pending. A later workbook import can append the missing information without resetting history.