Skip to content

[bot] Merge master/8168136f into rel/dev - #1860

Merged
yenkins-admin merged 1 commit into
rel/devfrom
snapshot-master-8168136f-to-rel/dev
Oct 8, 2026
Merged

yenkins-admin merged 1 commit into
rel/devfrom
snapshot-master-8168136f-to-rel/dev

Conversation

@yenkins-admin

Copy link
Copy Markdown
Contributor

🚀 Automated PR to perform merge from master into rel/dev with changes up to 8168136 (created by https://github.com/gooddata/gooddata-python-sdk/actions/runs/37749474522).

…ning one (#1816)

* feat(eval): capture the detail of every failing run, not just the winning one

An item that passed 1 of 3 runs reported only the run that won, so the two that
failed -- and the reason they did -- were discarded. Each failing run now keeps its
own detail, conversation id and exit reason, and they reach the JSON report, which
is where a failure is actually read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat(eval): capture failing runs for the agentic kinds too, not just single-shot

The capture landed in core/runner.py, which the agentic kinds never reach: cli/main
routes them to cli/agentic_runner instead, so every agentic result shipped the field
present and empty. Measured on 2026-09-18: 135 agentic_guardrail results, all with
`failed_runs: []`, including 45 items that passed some but not all of their runs.

All twelve K-running kinds now build the records through one shared helper, so a
failing run is described with the same keys as the winning one. Records are kept
when every run went ungraded -- the judge breaking is exactly when the per-run
conversation ids matter most, and that path previously threw them away.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(gooddata-eval): redact inside failed_runs, not just the item's top level

`failed_runs` records each failing run's own `conversation_id`, `response_id`
and `reasoning_steps`, plus its own `detail`. `_redact_item` only ever looked
at the item's top level, so `--redact` dropped the top-level `reasoning` while
the same model reasoning survived verbatim one level down -- in a report whose
whole purpose is being safe to hand a customer. Both builders are affected:
`core/runner._failed_run_record` and `core/agentic/_failed_runs`.

Found by an external check: a consumer searched a rendered redacted report for
every id and model name in its own source document, and ~270 ids came back.

Dropped per entry: conversation_id, response_id, reasoning_steps, and the same
transcript/tool_calls already dropped from the item's `detail`.

Kept, matching what the top level keeps: reasoning_step_count (a count, like
detail.turns), tool_call_count/tool_names (which tool ran, never its arguments
or its result -- where latency_breakdown already draws the line), and the run's
verdict, error and timings, which describe the run rather than our
infrastructure. An unredacted report is unchanged.

The existing test_redact_drops_ids_reasoning_and_model_name also fails without
this fix once its fixture carries failed_runs: the assertion was already right,
it just had no per-run data to catch.

1422 passed, ruff clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* style(gooddata-eval): format the report_skill failed_runs wiring

ruff format on the block this branch added to report_skill.py: the
build_failed_runs call fits one line, and _run_detail needs a blank line after
the assignment above it.

Mine, not master's -- master has no build_failed_runs in this file at all. It
went unnoticed because the integration branch, where I kept seeing it, already
contains this branch, so checking it there proved nothing about its origin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
@yenkins-admin
yenkins-admin merged commit 44fcc55 into rel/dev Oct 8, 2026
1 check passed
@yenkins-admin
yenkins-admin deleted the snapshot-master-8168136f-to-rel/dev branch October 8, 2026 08:23
@coderabbitai

coderabbitai Bot commented Oct 8, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 1e7edcf0-375e-4735-908f-614db6213d4d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants