Skip to content

k4bench.regression.regions

k4bench.regression.regions

Where inside the detector a timing step landed.

A confirmed run-level regression is one number, and one number names no mechanism: "ALLEGRO got 21% slower" and "the HCAL barrel got fourteen times slower while everything else stood still" are the same measurement, but only the second can be matched against a diff. The k4BenchRegionTimingAction plugin records per-event time per top-level detector region on every run, so the second form is already measured — it has simply never been read by anything that attributes a regression.

This module reads it: for one benchmark configuration and one change window, how each region's per-event time differs between the two ends the change entered between — the two releases, or the window's two runs when one release holds both. It judges nothing (the engine has already decided that the metric stepped) and it introduces no thresholds of its own; it reports the decomposition, largest movement first, and leaves the reading to whoever asked.

Two costs shape the implementation. Region files are per configuration and hold per-event arrays, so loading a whole trend window across every label is expensive — this loads exactly the two ends of one window, for one label, and only when something actually regressed there. And a release that recorded no region file is absent, never zero: a region that appears on one side of a window only is a real event (a detector added, removed or renamed) and must stay distinguishable from one that stood still.

The same files also carry each event's wall time and how much of it was spent in stepping, which is the other half of the question: whether the step is in the typical event or in a handful of long ones (:func:region_evidence). Both readings come from one read of each file.

dirs_by_release

dirs_by_release(run_dirs: Sequence[str]) -> dict[str, list[Path]]

Group run directories by the release they measured, keyed exactly as the engine keys releases (:func:~k4bench.regression.engine.release_key), so a window's ends match the verdict that named them.

Source code in k4bench/regression/regions.py
def dirs_by_release(run_dirs: Sequence[str]) -> dict[str, list[Path]]:
    """Group run directories by the release they measured, keyed exactly as the
    engine keys releases (:func:`~k4bench.regression.engine.release_key`), so a
    window's ends match the verdict that named them."""
    grouped: dict[str, list[Path]] = {}
    for path in run_dirs:
        run_dir = Path(path)
        if not run_dir.is_dir():
            continue
        meta = parse_run_dir(run_dir)
        # `pd.NaT` is *truthy*, so an `or` here would keep the missing release
        # date instead of falling back to the run date — the same fallback
        # `x_date` makes for the frame the engine walked. Test for absence
        # explicitly, or a run predating release-date capture keys on something
        # the verdict's window never names, and its regions read as unmeasured.
        release_date = meta.get("k4h_release_date")
        if release_date is None or pd.isna(release_date):
            release_date = meta.get("run_date")
        key = release_key(release_date, run_dir.name)
        grouped.setdefault(key, []).append(run_dir)
    return grouped

region_evidence

region_evidence(run_dirs: Sequence[str], *, label: str, base_release: str, onset_release: str, base_run_id: str | None = None, onset_run_id: str | None = None, limit: int = MAX_REGIONS, judgeable_configs: set[tuple[str, str]] | None = None) -> tuple[tuple[RegionDelta, ...], EventProfile | None]

How each region's per-event time moved across (base, onset], largest movement first, and the per-event wall times at both ends (:class:~k4bench.regression.models.EventProfile).

Returns ((), None) when either end recorded no region timing at all — with only one side measured there is no comparison to make, and inventing one (treating the missing side as zero) would report every region of the detector as newly appearing. The profile is likewise None unless both ends recorded per-event wall times. When judgeable_configs is supplied, failed or orphaned config-nights absent from that set are gaps and do not enter either end.

A window whose ends name one release is two runs of that release, and base_run_id / onset_run_id are what tell them apart — the same pair the verdict, the email's window token and the blame range are identified by. The ends are then those two runs rather than the release's pool, because one pool measured against itself is not a comparison and would report every region as having stood still. Without a resolvable run on each side there is again nothing to compare.

Source code in k4bench/regression/regions.py
def region_evidence(
    run_dirs: Sequence[str],
    *,
    label: str,
    base_release: str,
    onset_release: str,
    base_run_id: str | None = None,
    onset_run_id: str | None = None,
    limit: int = MAX_REGIONS,
    judgeable_configs: set[tuple[str, str]] | None = None,
) -> tuple[tuple[RegionDelta, ...], EventProfile | None]:
    """How each region's per-event time moved across ``(base, onset]``, largest
    movement first, and the per-event wall times at both ends
    (:class:`~k4bench.regression.models.EventProfile`).

    Returns ``((), None)`` when either end recorded no region timing at all —
    with only one side measured there is no comparison to make, and inventing
    one (treating the missing side as zero) would report every region of the
    detector as newly appearing. The profile is likewise ``None`` unless both
    ends recorded per-event wall times. When *judgeable_configs* is supplied,
    failed or orphaned config-nights absent from that set are gaps and do not
    enter either end.

    A window whose ends name one release is two runs of that release, and
    *base_run_id* / *onset_run_id* are what tell them apart — the same pair the
    verdict, the email's window token and the blame range are identified by. The
    ends are then those two runs rather than the release's pool, because one
    pool measured against itself is not a comparison and would report every
    region as having stood still. Without a resolvable run on each side there is
    again nothing to compare.
    """
    grouped = dirs_by_release(run_dirs)
    base_dirs, onset_dirs = grouped.get(base_release, []), grouped.get(onset_release, [])
    if not base_dirs or not onset_dirs:
        return (), None

    if base_release == onset_release:
        base_dir = _run_dir_named(base_dirs, base_run_id)
        onset_dir = _run_dir_named(onset_dirs, onset_run_id)
        if base_dir is None or onset_dir is None or base_dir == onset_dir:
            return (), None
        base_dirs, onset_dirs = [base_dir], [onset_dir]

    base_nights = _read_nights(base_dirs, label, judgeable_configs)
    onset_nights = _read_nights(onset_dirs, label, judgeable_configs)
    base = _release_medians(base_nights)
    onset = _release_medians(onset_nights)
    if not base or not onset:
        return (), None

    deltas = []
    for region in sorted(set(base) | set(onset)):
        before, after = base.get(region), onset.get(region)
        # A region present on one side only genuinely appeared or disappeared;
        # its whole time is the movement, and saying so is the point.
        delta = (after or 0.0) - (before or 0.0)
        if not math.isfinite(delta):
            continue
        deltas.append(RegionDelta(region=region, base=before, onset=after, delta=delta))
    deltas.sort(key=lambda d: (-abs(d.delta), d.region))

    base_sample, onset_sample = _event_sample(base_nights), _event_sample(onset_nights)
    profile = None
    if base_sample is not None and onset_sample is not None:
        events = dict.fromkeys(
            e.event for e in (*base_sample.longest, *onset_sample.longest)
        )
        matched = sorted(
            (
                MatchedEvent(
                    event=event,
                    base=_event_time(base_nights, event),
                    onset=_event_time(onset_nights, event),
                )
                for event in events
            ),
            key=lambda m: (-max(m.base or 0.0, m.onset or 0.0), m.event),
        )
        profile = EventProfile(
            base=base_sample, onset=onset_sample, matched=tuple(matched),
        )
    return tuple(deltas[:limit]), profile

region_deltas

region_deltas(run_dirs: Sequence[str], *, label: str, base_release: str, onset_release: str, base_run_id: str | None = None, onset_run_id: str | None = None, limit: int = MAX_REGIONS, judgeable_configs: set[tuple[str, str]] | None = None) -> tuple[RegionDelta, ...]

The region half of :func:region_evidence alone.

Source code in k4bench/regression/regions.py
def region_deltas(
    run_dirs: Sequence[str],
    *,
    label: str,
    base_release: str,
    onset_release: str,
    base_run_id: str | None = None,
    onset_run_id: str | None = None,
    limit: int = MAX_REGIONS,
    judgeable_configs: set[tuple[str, str]] | None = None,
) -> tuple[RegionDelta, ...]:
    """The region half of :func:`region_evidence` alone."""
    deltas, _profile = region_evidence(
        run_dirs, label=label,
        base_release=base_release, onset_release=onset_release,
        base_run_id=base_run_id, onset_run_id=onset_run_id,
        limit=limit, judgeable_configs=judgeable_configs,
    )
    return deltas