Repository navigation
[bot] Merge master/8168136f into rel/dev - #1860
Merged
Merged
Conversation
…ning one (#1816) * feat(eval): capture the detail of every failing run, not just the winning one An item that passed 1 of 3 runs reported only the run that won, so the two that failed -- and the reason they did -- were discarded. Each failing run now keeps its own detail, conversation id and exit reason, and they reach the JSON report, which is where a failure is actually read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat(eval): capture failing runs for the agentic kinds too, not just single-shot The capture landed in core/runner.py, which the agentic kinds never reach: cli/main routes them to cli/agentic_runner instead, so every agentic result shipped the field present and empty. Measured on 2026-09-18: 135 agentic_guardrail results, all with `failed_runs: []`, including 45 items that passed some but not all of their runs. All twelve K-running kinds now build the records through one shared helper, so a failing run is described with the same keys as the winning one. Records are kept when every run went ungraded -- the judge breaking is exactly when the per-run conversation ids matter most, and that path previously threw them away. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(gooddata-eval): redact inside failed_runs, not just the item's top level `failed_runs` records each failing run's own `conversation_id`, `response_id` and `reasoning_steps`, plus its own `detail`. `_redact_item` only ever looked at the item's top level, so `--redact` dropped the top-level `reasoning` while the same model reasoning survived verbatim one level down -- in a report whose whole purpose is being safe to hand a customer. Both builders are affected: `core/runner._failed_run_record` and `core/agentic/_failed_runs`. Found by an external check: a consumer searched a rendered redacted report for every id and model name in its own source document, and ~270 ids came back. Dropped per entry: conversation_id, response_id, reasoning_steps, and the same transcript/tool_calls already dropped from the item's `detail`. Kept, matching what the top level keeps: reasoning_step_count (a count, like detail.turns), tool_call_count/tool_names (which tool ran, never its arguments or its result -- where latency_breakdown already draws the line), and the run's verdict, error and timings, which describe the run rather than our infrastructure. An unredacted report is unchanged. The existing test_redact_drops_ids_reasoning_and_model_name also fails without this fix once its fixture carries failed_runs: the assertion was already right, it just had no per-run data to catch. 1422 passed, ruff clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * style(gooddata-eval): format the report_skill failed_runs wiring ruff format on the block this branch added to report_skill.py: the build_failed_runs call fits one line, and _run_detail needs a blank line after the assignment above it. Mine, not master's -- master has no build_failed_runs in this file at all. It went unnoticed because the integration branch, where I kept seeing it, already contains this branch, so checking it there proved nothing about its origin. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configuration
You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🚀 Automated PR to perform merge from master into rel/dev with changes up to 8168136 (created by https://github.com/gooddata/gooddata-python-sdk/actions/runs/37749474522).