Skip to content

k4bench.regression.report_builder

k4bench.regression.report_builder

Assemble the nightly regression report from the EOS run history.

Walks every (detector, platform, sample) triple found under the WebEOS data URL (the same hierarchy the dashboard's sidebar cascades through — these triples have independent baselines and are never pooled), pulls a trailing window of runs into the local cache, rebuilds the trend frames with :mod:k4bench.analysis.trend, attaches per-run reliability verdicts with :mod:k4bench.results.reliability_evidence, and runs the step detector in :mod:k4bench.regression.engine over every metric series.

PredecessorRuns dataclass

PredecessorRuns(platform: str, results_df: DataFrame | None, event_df: DataFrame | None, reliability: dict[str, bool | None], hosts: dict[str, HostFact] = dict(), run_dirs: tuple[str, ...] = ())

One run group's predecessor-platform frames, whose series this group's series continue (see :mod:k4bench.regression.lineage).

Kept apart from the group's own frames and joined per metric series only: everything else read off those frames — tonight's release, the CI run, the config roster, failures, region timings — is a statement about this platform, which a predecessor row would answer wrongly.

predecessor_runs

predecessor_runs(platform: str, run_dirs: tuple[str, ...]) -> PredecessorRuns | None

Build :class:PredecessorRuns from the predecessor's run directories.

Failed configs are dropped as they are for the group's own runs: a gap must not enter a continued series any more than the group's own.

Source code in k4bench/regression/report_builder.py
def predecessor_runs(platform: str, run_dirs: tuple[str, ...]) -> PredecessorRuns | None:
    """Build :class:`PredecessorRuns` from the predecessor's run directories.

    Failed configs are dropped as they are for the group's own runs: a gap must
    not enter a continued series any more than the group's own.
    """
    if not run_dirs:
        return None
    results_df = build_results_trend(run_dirs)
    event_df = build_event_timing_trend(run_dirs)
    machine_df = build_machine_info_trend(run_dirs)
    return PredecessorRuns(
        platform=platform,
        results_df=judgeable_config_rows(results_df, results_df),
        event_df=judgeable_config_rows(event_df, results_df),
        reliability=run_reliability_map(results_df, machine_df),
        hosts=host_facts(machine_df),
        run_dirs=tuple(run_dirs),
    )

unjudged_value_verdicts

unjudged_value_verdicts(*, detector: str, platform: str, sample: str, results_df: DataFrame | None, event_df: DataFrame | None, tonight: str, already: set[tuple[str, str]]) -> list[MetricVerdict]

Raw metric values for tonight's run as unjudged UNKNOWN verdicts.

Two different things end up here. The engine skips unreliable runs (they must not pollute baselines or flags), so their metrics get no verdict and their values would never reach the report the dashboard's Overview tab reads — leaving that tab unable to plot them even with "Exclude unreliable runs" off. And :data:REPORTED_ONLY_METRICS and :data:EVENT_REPORTED_ONLY_METRICS are never judged on any night by design, but are still worth being able to look up.

Either way this records tonight's raw value for every (label, metric) not already judged, marked UNKNOWN (never a flag), so the value is preserved for display. A normally-judged run is already covered.

Source code in k4bench/regression/report_builder.py
def unjudged_value_verdicts(
    *,
    detector: str,
    platform: str,
    sample: str,
    results_df: pd.DataFrame | None,
    event_df: pd.DataFrame | None,
    tonight: str,
    already: set[tuple[str, str]],
) -> list[MetricVerdict]:
    """Raw metric values for *tonight*'s run as unjudged ``UNKNOWN`` verdicts.

    Two different things end up here. The engine skips unreliable runs (they
    must not pollute baselines or flags), so their metrics get no verdict and
    their values would never reach the report the dashboard's Overview tab
    reads — leaving that tab unable to plot them even with "Exclude unreliable
    runs" off. And :data:`REPORTED_ONLY_METRICS` and
    :data:`EVENT_REPORTED_ONLY_METRICS` are never judged on any night
    by design, but are still worth being able to look up.

    Either way this records tonight's raw value for every ``(label, metric)``
    not *already* judged, marked ``UNKNOWN`` (never a flag), so the value is
    preserved for display. A normally-judged run is already covered.
    """
    out: list[MetricVerdict] = []

    def _emit(df: pd.DataFrame | None, metrics: dict[str, str]) -> None:
        if df is None or df.empty:
            return
        tonight_rows = df[df["run_id"] == tonight]
        for label in sorted(tonight_rows["label"].dropna().unique()):
            row = tonight_rows[tonight_rows["label"] == label]
            for metric, family in metrics.items():
                if metric not in row.columns or (str(label), metric) in already:
                    continue
                val = row[metric].iloc[0]
                if pd.isna(val) or not math.isfinite(float(val)):
                    continue
                reported_only = (
                    metric in REPORTED_ONLY_METRICS or metric in EVENT_REPORTED_ONLY_METRICS
                )
                out.append(
                    MetricVerdict(
                        detector=detector,
                        platform=platform,
                        sample=sample,
                        label=str(label),
                        metric_family=family,
                        metric=metric,
                        sub_detector=None,
                        run_id=tonight,
                        run_date=tonight,
                        value=float(val),
                        baseline_median=None,
                        baseline_mad=None,
                        pct_change=None,
                        z_score=None,
                        severity=Severity.UNKNOWN,
                        direction=Direction.NONE,
                        unjudged=(
                            Unjudged.REPORTED_ONLY if reported_only else Unjudged.UNRELIABLE_HOST
                        ),
                        reason=(
                            _REPORTED_ONLY_REASONS.get(metric, REPORTED_ONLY_REASON)
                            if reported_only
                            else UNRELIABLE_HOST_REASON
                        ),
                    )
                )

    _emit(results_df, RUN_VALUE_METRICS)
    _emit(event_df, EVENT_VALUE_METRICS)
    return out

evaluate_group_series

evaluate_group_series(*, detector: str, platform: str, sample: str, results_df: DataFrame | None, event_df: DataFrame | None, reliability: dict[str, bool | None], hosts: dict[str, HostFact] | None = None, predecessor: PredecessorRuns | None = None) -> dict[SeriesId, list[MetricVerdict]]

Run the step detector over every run/event metric series of one run group. Region timings are not walked.

Returns the full verdict series per :class:SeriesId — the nightly report takes each series' verdict for the report night, while the dashboard drill-down and the retrospective threshold validation consume the whole walk.

predecessor (from :func:predecessor_runs) is the replaced platform this one succeeds: each series is walked with the predecessor's matching series in front of it, so a migration step is judged like any other. Only verdicts for this group's own runs are returned. Omitted — the normal case — each series is its own history alone.

hosts (from :func:~k4bench.regression.history.host_facts) names the machine behind each run, and only reaches the history tails attached to confirmed verdicts; it never enters the judgement itself. Omitted, the tails simply carry no host.

Every configuration is judged on what it measured. A night where the whole run group moved together therefore reports each config's own move, rather than one synthetic finding standing in for all of them.

Source code in k4bench/regression/report_builder.py
def evaluate_group_series(
    *,
    detector: str,
    platform: str,
    sample: str,
    results_df: pd.DataFrame | None,
    event_df: pd.DataFrame | None,
    reliability: dict[str, bool | None],
    hosts: dict[str, HostFact] | None = None,
    predecessor: PredecessorRuns | None = None,
) -> dict[SeriesId, list[MetricVerdict]]:
    """Run the step detector over every run/event metric series of one run
    group. Region timings are not walked.

    Returns the **full verdict series** per :class:`SeriesId` — the nightly
    report takes each series' verdict for the report night, while the
    dashboard drill-down and the retrospective threshold validation consume
    the whole walk.

    *predecessor* (from :func:`predecessor_runs`) is the replaced platform this
    one succeeds: each series is walked with the predecessor's matching series in
    front of it, so a migration step is judged like any other. Only verdicts for
    this group's own runs are returned. Omitted — the normal case — each series
    is its own history alone.

    *hosts* (from :func:`~k4bench.regression.history.host_facts`) names the
    machine behind each run, and only reaches the history tails attached to
    confirmed verdicts; it never enters the judgement itself. Omitted, the tails
    simply carry no host.

    Every configuration is judged on what it measured. A night where the whole
    run group moved together therefore reports each config's own move, rather
    than one synthetic finding standing in for all of them.
    """
    out: dict[SeriesId, list[MetricVerdict]] = {}

    all_hosts = {**(predecessor.hosts if predecessor else {}), **(hosts or {})}

    def _walk(
        df: pd.DataFrame, metrics: dict[str, str], before_df: pd.DataFrame | None,
    ) -> None:
        labels = sorted(df["label"].dropna().unique())

        for metric, family in metrics.items():
            if metric not in df.columns:
                continue
            for label in labels:
                name = str(label)
                sid = SeriesId(detector, platform, sample, name, family, metric)
                own = _series_history(df, df["label"] == label, metric, reliability)
                history = _continued_history(own, predecessor, before_df, name, metric)
                verdicts = _with_history(
                    history, evaluate_series(history, series=sid), all_hosts,
                )
                own_runs = set(own["run_id"])
                verdicts = [
                    _with_endpoint_platforms(v, own_runs, predecessor)
                    for v in verdicts if v.run_id in own_runs
                ]
                if verdicts:
                    out[sid] = verdicts

    if results_df is not None and not results_df.empty:
        _walk(
            results_df, RUN_METRICS,
            predecessor.results_df if predecessor is not None else None,
        )

    if event_df is not None and not event_df.empty:
        _walk(
            event_df, EVENT_METRICS,
            predecessor.event_df if predecessor is not None else None,
        )

    return out

build_group_report

build_group_report(data_url: str, cache_dir: str | None, detector: str, platform: str, sample: str, *, fetch_window_runs: int = FETCH_WINDOW_RUNS, as_of: str | None = None) -> RunGroupReport | None

Build one triple's report from its trailing run window, or None when the triple has no fetchable runs at all.

as_of (a YYYY-MM-DD night) truncates the run history to runs on or before that night before the trailing window is taken, reproducing the report that night's runs would have produced — the seam the historical backfill drives. None judges the full history (the nightly CI case).

Source code in k4bench/regression/report_builder.py
def build_group_report(
    data_url: str,
    cache_dir: str | None,
    detector: str,
    platform: str,
    sample: str,
    *,
    fetch_window_runs: int = FETCH_WINDOW_RUNS,
    as_of: str | None = None,
) -> RunGroupReport | None:
    """Build one triple's report from its trailing run window, or ``None``
    when the triple has no fetchable runs at all.

    *as_of* (a ``YYYY-MM-DD`` night) truncates the run history to runs on or
    before that night before the trailing window is taken, reproducing the
    report that night's runs would have produced — the seam the historical
    backfill drives. ``None`` judges the full history (the nightly CI case).
    """
    run_dirs = _fetch_run_dirs(
        data_url, cache_dir, detector, platform, sample,
        fetch_window_runs=fetch_window_runs, as_of=as_of,
    )
    if not run_dirs:
        return None
    predecessor = predecessor_of(platform)
    return group_report_from_run_dirs(
        detector, platform, sample, run_dirs,
        predecessor=None if predecessor is None else partial(
            _fetch_predecessor, data_url, cache_dir, detector, predecessor, sample,
            fetch_window_runs=fetch_window_runs, as_of=as_of,
        ),
    )

group_report_from_run_dirs

group_report_from_run_dirs(detector: str, platform: str, sample: str, run_dirs: tuple[str, ...], *, predecessor: Callable[[], PredecessorRuns | None] | None = None) -> RunGroupReport | None

Build one triple's report from already-local run directories (ordered oldest → newest; each directory's name is its nightly date).

predecessor loads the platform this one replaced, whose series this group's continue; it never contributes a verdict, a release, a failure or a timing. It is loaded for every replay, because a series' detection state can depend on the nights it continues however long ago they were.

Source code in k4bench/regression/report_builder.py
def group_report_from_run_dirs(
    detector: str,
    platform: str,
    sample: str,
    run_dirs: tuple[str, ...],
    *,
    predecessor: Callable[[], PredecessorRuns | None] | None = None,
) -> RunGroupReport | None:
    """Build one triple's report from already-local run directories (ordered
    oldest → newest; each directory's name is its nightly date).

    *predecessor* loads the platform this one replaced, whose series
    this group's continue; it never contributes a verdict, a release, a failure
    or a timing. It is loaded for every replay, because a series' detection
    state can depend on the nights it continues however long ago they were.
    """
    if not run_dirs:
        return None
    tonight = max(Path(d).name for d in run_dirs)
    tonight_meta = parse_run_dir(
        next(Path(d) for d in run_dirs if Path(d).name == tonight)
    )
    results_df = build_results_trend(run_dirs)
    event_df = build_event_timing_trend(run_dirs)
    machine_df = build_machine_info_trend(run_dirs)
    reliability = run_reliability_map(results_df, machine_df)
    before = predecessor() if predecessor is not None else None
    group = _group_report_from_frames(
        detector, platform, sample,
        results_df=results_df, event_df=event_df,
        reliability=reliability, tonight=tonight,
        hosts=host_facts(machine_df),
        configured_labels=tonight_meta["configured_labels"],
        predecessor=before,
    )
    if group is None:
        return None
    # A night that wrote no result CSV has no release in its (absent) rows;
    # run_info still names the stack that failed.
    if not group.k4h_release:
        group.k4h_release = tonight_meta["k4h_release"] or ""
    return _with_region_deltas(
        group, run_dirs, judgeable_config_keys(results_df), predecessor=before,
    )

build_nightly_report

build_nightly_report(data_url: str, cache_dir: str | None = None, *, fetch_window_runs: int = FETCH_WINDOW_RUNS, as_of: str | None = None, night: str | None = None) -> NightlyReport

Build the cross-detector report for the most recent nightly.

The report night is the newest run date seen across all triples. A triple dated earlier is still reported normally when its CI run says it came from the report night's own batch, whatever the gap between the two dates (see :func:_same_batch); for a night whose runs carry no CI run at all, a lag of up to :data:SAME_BATCH_LAG_DAYS stands in for that. Anything else gets a missing run job failure (a hard crash uploads nothing, so absence is itself the failure signal) — unless it is stale by more than :data:MISSING_RUN_GRACE_DAYS, in which case it is treated as retired and dropped.

as_of truncates every triple's history to runs on or before that night (see :func:build_group_report), making the report night the newest run ≤ as_of — the historical-backfill seam. night reports a night newer than every run found, as one missing run per triple (see :func:_finalize_report) — the outage seam.

Source code in k4bench/regression/report_builder.py
def build_nightly_report(
    data_url: str,
    cache_dir: str | None = None,
    *,
    fetch_window_runs: int = FETCH_WINDOW_RUNS,
    as_of: str | None = None,
    night: str | None = None,
) -> NightlyReport:
    """Build the cross-detector report for the most recent nightly.

    The report night is the newest run date seen across all triples. A triple
    dated earlier is still reported normally when its CI run says it came from
    the report night's own batch, whatever the gap between the two dates (see
    :func:`_same_batch`); for a night whose runs carry no CI run at all, a lag
    of up to :data:`SAME_BATCH_LAG_DAYS` stands in for that. Anything else gets
    a *missing run* job failure (a hard crash uploads nothing, so absence is
    itself the failure signal) — unless it is stale by more than
    :data:`MISSING_RUN_GRACE_DAYS`, in which case it is treated as retired and
    dropped.

    *as_of* truncates every triple's history to runs on or before that night
    (see :func:`build_group_report`), making the report night the newest run
    ≤ *as_of* — the historical-backfill seam. *night* reports a night newer
    than every run found, as one missing run per triple (see
    :func:`_finalize_report`) — the outage seam.
    """
    groups: list[RunGroupReport] = []
    for detector in list_detectors(data_url):
        for platform in list_platforms(data_url, detector):
            stack_samples = scan_stack_samples(data_url, detector, platform)
            samples = sorted({s for ss in stack_samples.values() for s in ss})
            for sample in samples:
                try:
                    group = build_group_report(
                        data_url, cache_dir, detector, platform, sample,
                        fetch_window_runs=fetch_window_runs, as_of=as_of,
                    )
                except Exception:
                    _log.exception(
                        "build_nightly_report: failed for %s/%s/%s",
                        detector, platform, sample,
                    )
                    continue
                if group is not None:
                    groups.append(group)

    return _finalize_report(groups, night=night)

build_nightly_report_local

build_nightly_report_local(data_dir: str, *, fetch_window_runs: int = FETCH_WINDOW_RUNS, as_of: str | None = None, night: str | None = None) -> NightlyReport

Like :func:build_nightly_report, but over a local directory tree with the same {detector}/{platform}/{stack}/{sample}/{date} layout as EOS (used by the integration test and for offline dry-runs; no network). as_of truncates each sample's runs the same way, and night reports an outage night the same way.

Source code in k4bench/regression/report_builder.py
def build_nightly_report_local(
    data_dir: str,
    *,
    fetch_window_runs: int = FETCH_WINDOW_RUNS,
    as_of: str | None = None,
    night: str | None = None,
) -> NightlyReport:
    """Like :func:`build_nightly_report`, but over a local directory tree with
    the same ``{detector}/{platform}/{stack}/{sample}/{date}`` layout as EOS
    (used by the integration test and for offline dry-runs; no network).
    *as_of* truncates each sample's runs the same way, and *night* reports an
    outage night the same way."""
    root = Path(data_dir)
    groups: list[RunGroupReport] = []

    def _window(run_paths: list[Path]) -> tuple[str, ...]:
        return tuple(
            str(p) for p in sorted(run_paths, key=lambda p: p.name)
            if as_of is None or p.name <= as_of
        )[-fetch_window_runs:]

    for det_dir in sorted(p for p in root.iterdir() if p.is_dir()):
        if det_dir.name.startswith(("_", ".")):
            continue
        # Every platform first: a successor continues its predecessor's runs.
        per_platform: dict[str, dict[str, list[Path]]] = {}
        for plat_dir in sorted(p for p in det_dir.iterdir() if p.is_dir()):
            # Collect each sample's run dirs across all stacks.
            per_sample: dict[str, list[Path]] = {}
            for stack_dir in sorted(p for p in plat_dir.iterdir() if p.is_dir()):
                for sample_dir in sorted(p for p in stack_dir.iterdir() if p.is_dir()):
                    per_sample.setdefault(sample_dir.name, []).extend(
                        p for p in sample_dir.iterdir() if p.is_dir()
                    )
            per_platform[plat_dir.name] = per_sample

        for platform, per_sample in per_platform.items():
            predecessor = predecessor_of(platform)
            for sample, run_paths in sorted(per_sample.items()):
                predecessor_dirs = _window(
                    per_platform.get(predecessor, {}).get(sample, [])
                ) if predecessor is not None else ()
                group = group_report_from_run_dirs(
                    det_dir.name, platform, sample, _window(run_paths),
                    predecessor=None if predecessor is None else partial(
                        predecessor_runs, predecessor, predecessor_dirs,
                    ),
                )
                if group is not None:
                    groups.append(group)
    return _finalize_report(groups, night=night)

report_covers_run

report_covers_run(report: NightlyReport, run_id: str) -> bool

Whether run_id's benchmark jobs produced any of this report's runs.

The question a nightly's report job has to answer before publishing: a fan-out where every job failed uploads nothing, and the newest run on EOS is then still the previous night's — a report built from it would republish and re-mail verdicts that night already sent. One group naming the run is enough, whatever its date: a batch is stamped per job start, so its own runs can straddle midnight.

Compares run ids rather than URLs, so a re-run's /attempts/2 suffix still matches (see :func:_batch_key).

Source code in k4bench/regression/report_builder.py
def report_covers_run(report: NightlyReport, run_id: str) -> bool:
    """Whether *run_id*'s benchmark jobs produced any of this report's runs.

    The question a nightly's report job has to answer before publishing: a
    fan-out where every job failed uploads nothing, and the newest run on EOS
    is then still the previous night's — a report built from it would republish
    and re-mail verdicts that night already sent. One group naming the run is
    enough, whatever its date: a batch is stamped per job start, so its own
    runs can straddle midnight.

    Compares run ids rather than URLs, so a re-run's ``/attempts/2`` suffix
    still matches (see :func:`_batch_key`).
    """
    if not run_id:
        return False
    keys = {_batch_key(g.github_run_url) for g in report.groups if g.github_run_url}
    return any(k.rpartition("/actions/runs/")[2] == run_id for k in keys)