Skip to content

🤖 feat: improve plan tool to scale with complexity - #32

Merged
ammario merged 1 commit into
mainfrom
improve-plan-tool-conciseness
Oct 5, 2025
Merged

ammario merged 1 commit into
mainfrom
improve-plan-tool-conciseness

Conversation

@ammario

@ammario ammario commented Oct 5, 2025

Copy link
Copy Markdown
Member

Summary

Updates the plan tool description to encourage proportional planning that scales with task complexity, making plans more concise for simple tasks while still allowing detailed plans for complex work.

Problem

The current plan tool description prescribes a fixed, comprehensive structure regardless of task complexity:

  • Requires: problem/context, chosen approach with rationale, alternatives considered, step-by-step implementation, edge cases and risks, scope/impact, testing strategy
  • Results in over-planning for simple tasks
  • Creates unnecessarily long plans that slow down review
  • Wastes tokens on ceremonial sections

Changes

Tool Description (before):

"Propose a detailed plan for implementation. The plan will be reviewed by the user 
before execution. Use this to outline the approach, steps, and considerations 
before making changes."

Tool Description (after):

"Propose a plan before taking action. The plan should be complete but minimal - 
cover what needs to be decided or understood, nothing more. Use this tool to get 
approval before proceeding with implementation."

Plan Parameter Description (before):

"Implementation plan in markdown (start at h2 level). Cover: problem/context, 
chosen approach with rationale, alternatives considered, step-by-step implementation, 
edge cases and risks, scope/impact on existing code, testing strategy. Break complex 
work into phases."

Plan Parameter Description (after):

"Implementation plan in markdown (start at h2 level). Scale the detail to match the 
task complexity: for straightforward changes, briefly state what and why; for complex 
changes, explain approach, key decisions, risks/tradeoffs; for uncertain changes, 
clarify options and what needs user input. Cover what's necessary to understand and 
approve the approach. Omit obvious details or ceremony."

Key Improvements

  1. Mandatory planning in plan mode - "Propose a plan before taking action" (imperative, not conditional)
  2. Complete but minimal - Guides toward sufficiency without excess
  3. Scales with complexity - Provides examples for simple vs complex vs uncertain tasks
  4. Removes prescriptive requirements - No longer requires "alternatives considered", "testing strategy", etc.
  5. Avoids ceremony - "Omit obvious details or ceremony"

Expected Behavior

Simple task: "Add parameter X"
→ Brief plan (what/why/how)

Complex task: "Redesign auth system"
→ Detailed plan (architecture, phases, risks)

Uncertain task: "Not sure which approach"
→ Decision-focused plan (options with pros/cons)

Benefits

  • ⚡ Faster review cycles - Users can quickly scan simpler plans
  • 💰 Lower token costs - Shorter plans for simple tasks
  • 🎯 Better focus - Agent concentrates on actual decisions, not ceremony
  • 📏 Proportional detail - Complex tasks still get detailed plans when needed

Generated with cmux

Update plan tool description to encourage proportional planning:
- Plans should always be used (mandatory in plan mode)
- Plans should be 'complete but minimal'
- Detail level scales with task complexity
- Removes prescriptive requirements (alternatives, testing strategy, etc.)

Changes:
- Tool description: 'Propose a plan before taking action' (imperative)
- Plan parameter: Provides scaling guidance for simple vs complex tasks
- Emphasizes 'cover what's necessary, omit obvious details'

This allows agents to produce concise plans for simple tasks while
still providing detailed plans for complex work, improving review
speed and reducing token usage without sacrificing completeness.
@ammario
ammario merged commit 38a5904 into main Oct 5, 2025
6 checks passed
@ammario
ammario deleted the improve-plan-tool-conciseness branch October 5, 2025 21:20
ammario pushed a commit that referenced this pull request May 28, 2026
## Summary

`task_await` now accepts an optional `min_completed` integer (default
**1**). It returns as soon as that many awaited tasks have completed —
by default the **first** completion — instead of always blocking until
every awaited task finishes. The parent can act on each result as it
lands (e.g. integrate variant lane #45 while #32/#69 keep running) and
re-await the remainder, rather than idling until the whole batch is
done.

## Background

When the parent spawns a grouped batch (`task` with `n` for best-of-N or
`variants`), there was often useful dependent work available after just
**one** child completed — most obviously for `variants`, where each lane
is independent. But both the foreground `task` path and `task_await`
used `Promise.all` and only returned once the **slowest** task settled,
and the tool/prelude guidance framed the flow as "launch → await as a
batch → synthesize all." Nothing let the parent begin work after the
first child.

## Implementation

- **Schema** (`toolDefinitions.ts`): new `min_completed:
z.number().int().min(1).nullish()` on `TaskAwaitToolArgsSchema`;
absent/`null` ⇒ `1`.
- **Handler** (`task_await.ts`): the per-task await body is now
`awaitOne(taskId, taskSignal)`. Each task gets its own `AbortController`
chained to the tool-call signal. A small coordinator resolves once
`min_completed` tasks have completed **or** every task has otherwise
settled (so an unreachable threshold still returns promptly).
Still-pending "losers" are then **aborted to detach their waiters** —
which only removes the in-memory waiter / interrupts a bash poll; the
child keeps running and its report stays cached and re-awaitable on a
later call. Losers are reported with a live `running`/`queued` status
snapshot rather than an error (a real tool-call interrupt is still
distinguished and surfaced as `error: "Interrupted"`).
- `timeout_secs: 0` stays a non-blocking snapshot of every task
regardless of `min_completed`.
- **Clamped** to the number of awaited tasks (over-large values behave
like "wait for all").
- **Guidance**: `task`/`task_await` descriptions and the `<best-of-n>` /
`<task-variants>` prelude blocks now steer independent lanes toward the
first-completion loop, and tell best-of-N synthesis (which must compare
every candidate) to pass `min_completed` equal to the batch size or use
a foreground grouped spawn. Auto-generated docs (`system-prompt.mdx`,
`hooks/tools.mdx`) and the built-in skill snapshot were regenerated via
`make fmt`.

The foreground `task` (`run_in_background: false`) path is intentionally
**unchanged** — grouped spawns still return all reports, which remains
the natural "give me every candidate" path.

## Validation

- New `task_await` tests: default returns after first (rest `running`);
`min_completed = total` waits for all; `min_completed = k` returns after
the k-th; clamp; re-awaitable loser resolves on a follow-up call;
unreachable threshold returns promptly; `timeout_secs:0` non-blocking
with `min_completed`. The first-completion test also asserts the loser's
per-task signal is aborted while the winner's is not.
- `make static-check` passes (typecheck, eslint, formatting, docs-sync).

## Risks

`min_completed` defaults to `1`, so this is a **behavior change** for
`task_await`: an existing background-spawn-then-await flow that expected
all reports now returns after the first completion. Mitigations: the
all-settled fallback + clamp keep results well-formed, losers remain
re-awaitable (no work lost, children not terminated), and guidance
steers aggregation flows to pass the batch size. Affected area is
limited to sub-agent orchestration.

---

_Generated with `mux` • Model: `anthropic:claude-opus-4-8` • Thinking:
`xhigh` • Cost: `$5.76`_

<!-- mux-attribution: model=anthropic:claude-opus-4-8 thinking=xhigh
costs=5.76 -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant