k4bench.regression.report_builder¶
k4bench.regression.report_builder ¶
Assemble the nightly regression report from the EOS run history.
Walks every (detector, platform, sample) triple found under the WebEOS
data URL (the same hierarchy the dashboard's sidebar cascades through — these
triples have independent baselines and are never pooled), pulls a trailing
window of runs into the local cache, rebuilds the trend frames with
:mod:k4bench.analysis.trend, attaches per-run reliability verdicts with
:mod:k4bench.results.reliability_evidence, and runs the step detector in
:mod:k4bench.regression.engine over every metric series.
PredecessorRuns
dataclass
¶
PredecessorRuns(platform: str, results_df: DataFrame | None, event_df: DataFrame | None, reliability: dict[str, bool | None], hosts: dict[str, HostFact] = dict(), run_dirs: tuple[str, ...] = ())
One run group's predecessor-platform frames, whose series this group's
series continue (see :mod:k4bench.regression.lineage).
Kept apart from the group's own frames and joined per metric series only: everything else read off those frames — tonight's release, the CI run, the config roster, failures, region timings — is a statement about this platform, which a predecessor row would answer wrongly.
predecessor_runs ¶
predecessor_runs(platform: str, run_dirs: tuple[str, ...]) -> PredecessorRuns | None
Build :class:PredecessorRuns from the predecessor's run directories.
Failed configs are dropped as they are for the group's own runs: a gap must not enter a continued series any more than the group's own.
Source code in k4bench/regression/report_builder.py
unjudged_value_verdicts ¶
unjudged_value_verdicts(*, detector: str, platform: str, sample: str, results_df: DataFrame | None, event_df: DataFrame | None, tonight: str, already: set[tuple[str, str]]) -> list[MetricVerdict]
Raw metric values for tonight's run as unjudged UNKNOWN verdicts.
Two different things end up here. The engine skips unreliable runs (they
must not pollute baselines or flags), so their metrics get no verdict and
their values would never reach the report the dashboard's Overview tab
reads — leaving that tab unable to plot them even with "Exclude unreliable
runs" off. And :data:REPORTED_ONLY_METRICS and
:data:EVENT_REPORTED_ONLY_METRICS are never judged on any night
by design, but are still worth being able to look up.
Either way this records tonight's raw value for every (label, metric)
not already judged, marked UNKNOWN (never a flag), so the value is
preserved for display. A normally-judged run is already covered.
Source code in k4bench/regression/report_builder.py
evaluate_group_series ¶
evaluate_group_series(*, detector: str, platform: str, sample: str, results_df: DataFrame | None, event_df: DataFrame | None, reliability: dict[str, bool | None], hosts: dict[str, HostFact] | None = None, predecessor: PredecessorRuns | None = None) -> dict[SeriesId, list[MetricVerdict]]
Run the step detector over every run/event metric series of one run group. Region timings are not walked.
Returns the full verdict series per :class:SeriesId — the nightly
report takes each series' verdict for the report night, while the
dashboard drill-down and the retrospective threshold validation consume
the whole walk.
predecessor (from :func:predecessor_runs) is the replaced platform this
one succeeds: each series is walked with the predecessor's matching series in
front of it, so a migration step is judged like any other. Only verdicts for
this group's own runs are returned. Omitted — the normal case — each series
is its own history alone.
hosts (from :func:~k4bench.regression.history.host_facts) names the
machine behind each run, and only reaches the history tails attached to
confirmed verdicts; it never enters the judgement itself. Omitted, the tails
simply carry no host.
Every configuration is judged on what it measured. A night where the whole run group moved together therefore reports each config's own move, rather than one synthetic finding standing in for all of them.
Source code in k4bench/regression/report_builder.py
build_group_report ¶
build_group_report(data_url: str, cache_dir: str | None, detector: str, platform: str, sample: str, *, fetch_window_runs: int = FETCH_WINDOW_RUNS, as_of: str | None = None) -> RunGroupReport | None
Build one triple's report from its trailing run window, or None
when the triple has no fetchable runs at all.
as_of (a YYYY-MM-DD night) truncates the run history to runs on or
before that night before the trailing window is taken, reproducing the
report that night's runs would have produced — the seam the historical
backfill drives. None judges the full history (the nightly CI case).
Source code in k4bench/regression/report_builder.py
group_report_from_run_dirs ¶
group_report_from_run_dirs(detector: str, platform: str, sample: str, run_dirs: tuple[str, ...], *, predecessor: Callable[[], PredecessorRuns | None] | None = None) -> RunGroupReport | None
Build one triple's report from already-local run directories (ordered oldest → newest; each directory's name is its nightly date).
predecessor loads the platform this one replaced, whose series this group's continue; it never contributes a verdict, a release, a failure or a timing. It is loaded for every replay, because a series' detection state can depend on the nights it continues however long ago they were.
Source code in k4bench/regression/report_builder.py
build_nightly_report ¶
build_nightly_report(data_url: str, cache_dir: str | None = None, *, fetch_window_runs: int = FETCH_WINDOW_RUNS, as_of: str | None = None, night: str | None = None) -> NightlyReport
Build the cross-detector report for the most recent nightly.
The report night is the newest run date seen across all triples. A triple
dated earlier is still reported normally when its CI run says it came from
the report night's own batch, whatever the gap between the two dates (see
:func:_same_batch); for a night whose runs carry no CI run at all, a lag
of up to :data:SAME_BATCH_LAG_DAYS stands in for that. Anything else gets
a missing run job failure (a hard crash uploads nothing, so absence is
itself the failure signal) — unless it is stale by more than
:data:MISSING_RUN_GRACE_DAYS, in which case it is treated as retired and
dropped.
as_of truncates every triple's history to runs on or before that night
(see :func:build_group_report), making the report night the newest run
≤ as_of — the historical-backfill seam. night reports a night newer
than every run found, as one missing run per triple (see
:func:_finalize_report) — the outage seam.
Source code in k4bench/regression/report_builder.py
build_nightly_report_local ¶
build_nightly_report_local(data_dir: str, *, fetch_window_runs: int = FETCH_WINDOW_RUNS, as_of: str | None = None, night: str | None = None) -> NightlyReport
Like :func:build_nightly_report, but over a local directory tree with
the same {detector}/{platform}/{stack}/{sample}/{date} layout as EOS
(used by the integration test and for offline dry-runs; no network).
as_of truncates each sample's runs the same way, and night reports an
outage night the same way.
Source code in k4bench/regression/report_builder.py
report_covers_run ¶
report_covers_run(report: NightlyReport, run_id: str) -> bool
Whether run_id's benchmark jobs produced any of this report's runs.
The question a nightly's report job has to answer before publishing: a fan-out where every job failed uploads nothing, and the newest run on EOS is then still the previous night's — a report built from it would republish and re-mail verdicts that night already sent. One group naming the run is enough, whatever its date: a batch is stamped per job start, so its own runs can straddle midnight.
Compares run ids rather than URLs, so a re-run's /attempts/2 suffix
still matches (see :func:_batch_key).