Repository navigation
Extension capability system: portable commands across agent harnesses #3303
Description
Activity
Thanks for such a thoughtful writeup, @rhuss — and for reading the source before filing. You've correctly identified the real seam: spec-kit solves structural portability (format, paths, placeholders) but not content portability, and the building blocks you pointed at (
post_process_skill_content,process_template,expressions.py) are the right places to be looking. cc-spex is a great stress test for this.I want to give you an honest answer rather than a vague "maybe," so let me split the proposal into the part we'd welcome in core and the part we'd push back on — and then suggest where the bulk of this most naturally lives.
The part that fits in core. Generalizing the skills-only
post_process_skill_content()into apost_process_command_content()available to all formats is reasonable and low-risk. Each integration owns its own transform, it's branch-free, and it's testable per-integration. If you'd like to open a focused PR for just that hook, we'd be glad to review it — it gives extension authors a clean per-agent seam without any shared machinery.The part we'd steer away from — the capability matrix plus conditional template blocks. Not because the pain isn't real (it is), but because that specific design pushes against an invariant that a lot of spec-kit's value quietly rests on.
When you run
specify init, everything is materialized into the project: scripts, templates, memory, workflows, and the rendered command files all land under the project (.specify/…and the agent's own folders). Extensions install into.specify/extensions/<id>/; presets compose the command files in place. After init, the coding agent only ever reads local files — the user never has to clone spec-kit core or any catalog repo, everything works offline, and the on-disk files are the whole truth. That self-containment is what makes the output auditable, forkable, and diffable, and it's what lets non-Python authors (prompt engineers, domain experts writingaide-style workflows) create and review commands with confidence.A capability system with
{{#if subagents}}…{{else}}…{{/if}}would be the first feature to puncture that invariant, and it does so at the level of meaning, not just files:- The rendered command on disk stays concrete (the branch resolves at install), but the semantics of the template — which capability keys exist, what they resolve to for each agent — live in core's Python integration classes, which are never materialized into the project. For the first time, fully understanding or reproducing your own project's commands would require reaching outside the project, into core code on a different release cadence.
- Extension and preset templates live in separate catalog repos, disjoint from core's Python. A
{{#if subagents}}block there references a capability defined in another repo entirely, so those userland templates could no longer be validated on their own — a reviewer would need core's sources they may not have and shouldn't need. - It also turns agnostic, one-file-one-meaning commands into N latent variants, and imports workflow-style logic into files meant to be static docs (the pull toward
expressions.pyis the tell).
And two operational realities make the matrix especially costly in core:
- The data is volatile and externally controlled — tool names, hook names, and even "does this exist at all" change on each agent vendor's schedule, silently. Core would effectively be subscribing to dozens of external changelogs, and staleness would mis-render agnostic commands for everyone.
- The maintainers can't reliably test the matrix — we don't have all ~40 agents installed and observable, so capability claims could be linted for shape but never verified for truth. The people who can verify a given agent's behavior are the folks who actually run it.
It's worth being precise about what the capability system would actually save, because I don't think it's the user's experience. A user doesn't install "N variants" — they pick the agent(s) they actually use, and that maps to the matching extension (if any) plus the matching preset for that agent. Each user only ever consumes the one variant for their own agent. The N-variants exist only as an author-maintenance cost — and that cost sits most appropriately with the person who wants multi-agent reach and can actually test each agent, rather than being centralized into a core matrix no one can verify.
Put together, agent-specific optimization tends to be local (each user only wants their own agent's variant), volatile (tracks that agent's churn), testable only by its owner, and opt-in — and those properties line up with our userland extensibility layers, not with core metadata.
Where we'd suggest putting it instead — and it's genuinely buildable there:
- Presets are built for "take an agnostic command and specialize it for an agent" via
prepend/replace. A community preset likecopilot-sub-agentsalready does your subagent-fanout case for Copilot; aclaude-…counterpart would do it for Claude. A user just installs the one matching the agent they use — each variant is self-contained plain text, owned and tested by someone who runs that agent, and it stays materialized in the project like everything else. - Extensions already install helper scripts to a stable, project-local path (
.specify/extensions/<id>/), and path rewriting preserves extension-local references — so your concern doco(spec-driven): Fix small typo in spec-driven.md #4 (script discovery) is largely handled today without a new__PLUGIN_ROOT__placeholder. Withpost_process_command_content()added, you'd also get the per-agent content-transform seam.
So the capability behavior you're after can be assembled on top of spec-kit — a neutral extension plus a thin per-agent preset — while keeping each piece materialized, self-contained, and auditable: the specialization stays right there in the project as plain text anyone can read, diff, and fork, owned by someone who can actually verify it against the agent it targets. That preserves the guarantee a capability matrix in core would break — that your project is the whole truth, reviewable without reaching into core Python we couldn't reliably test.
If it's helpful, I'm happy to sketch what a "portable subagent" preset would look like using cc-spex's cases as the example, and to review that
post_process_command_content()PR if you'd like to take it on. Really appreciate you pushing on this — it's exactly the right conversation to be having.The materialization invariant argument sold me. I hadn't thought about it that way: a capability matrix would be the first thing that forces you outside the project to understand your own commands. That's not just an engineering tradeoff, it's a design problem.
Here's what I'm planning based on your feedback:
post_process_command_content()PR. I'll open a focused PR. Add the method toIntegrationBasewith a no-op default, call it fromregister_commands()for markdown/TOML/YAML, include tests against the existing skills-format agents to make sure nothing breaks.Presets for agent specialization. Your reframing clicked for me. I was stuck on this as an authoring problem ("one command that works everywhere") when it's really distribution ("each user gets the right variant for their agent"). Presets handle that, and the result stays materialized in the project. I'll prototype a
claude-spexpreset usingprepend/replaceto inject Claude Code tool references into neutral commands.One question on presets. For complex agent-specific sections (say, a ship pipeline that dispatches 4 subagents into isolated worktrees vs. running everything sequentially),
replacemight not be enough. The neutral and specialized versions aren't just different words, they're different control flow. Canreplacehandle multi-line block replacements keyed on a marker? Or is this a case where two genuinely different commands is the cleaner answer?I'll start with the
post_process_command_content()PR since that's useful on its own.Answering your question about markers. No markers are not supported currently.. Would love tho see the PR you said you will start with first
allright, I'm on it :-) Thanks for the quick turnaround.
- added a commit that references this issue
on Jul 2, 2026 Closing this in favor of two focused follow-ups:
post_process_command_content()hook: #3311. This gives extensions a per-integration content transform seam without any shared capability metadata.- Extension lifecycle hook for post-install setup: #3359. Instead of a capability matrix in core, extensions get an
on_installhook to self-adapt based on the detected harness. The result stays fully materialized and auditable, consistent with the invariant @mnriem outlined here.
Thanks for the thorough response on the materialization invariant, it reframed the problem well.
- added a commit that references this issue
on Jul 7, 2026
I'm building an extension (cc-spex, an SDD methodology toolkit on top of spec-kit) and ran into a portability challenge. Spec-kit's extension system handles format conversion and path routing across 28 agent harnesses well. But when my command content needs to reference agent-specific capabilities ("ask the user a question," "spawn a subagent"), I end up hardcoding tool names for one agent.
Before I work around this on my side, I wanted to ask: is there an existing mechanism for making extension command content portable across agents? If not, would it make sense to add one?
What I'm running into
Spec-kit solves the structural layer well (format conversion, path routing, argument placeholders). What I'm struggling with is the content layer. There are several categories of agent-specific behavior that don't seem to have a portable abstraction:
When a command needs to ask "which approach do you prefer?", it needs to know the right prompt tool:
AskUserQuestionon Claude Code,questionon OpenCode, or nothing on Codex. I haven't found a way to express "ask the user" generically.Some workflows fan out to subagents for parallel work (code review across dimensions, research from multiple angles). On Claude Code that's the
Agenttool, on OpenCode it'sTask, and on many agents it simply doesn't exist. Ideally a command could say "if subagents are available, fan out; otherwise, run sequentially."Workflow discipline (e.g., blocking implementation until a spec is reviewed) depends on hooks that vary widely:
PreToolUseon Claude Code,tool.execute.beforeon OpenCode, or nothing at all. An extension currently can't discover what enforcement level the agent supports.Extensions bundle helper scripts that commands invoke at runtime. The install location varies by agent (
~/.claude/plugins/,~/.opencode/plugins/, etc.), and there's no portable way for a command to reference its own script directory.Some workflows benefit from clearing conversation context between phases. Claude Code has
/clearfor this, but most agents don't have an equivalent.These come up in any extension that goes beyond "run a script and report." A code review extension that wants to ask "Fix all / Let me pick / Skip," a CI extension that parallelizes test runs, a workflow extension that enforces stage ordering. In my case, I audited cc-spex and found 46 agent-specific references across its commands, which effectively makes it single-agent despite spec-kit supporting 28.
Ideas (if this doesn't exist yet)
I looked at the source and noticed building blocks that could support this.
Each
IntegrationBasesubclass could declare what the agent supports, following the existing_feature_capabilities()pattern but at the integration level:With capability metadata in place, command templates could use conditional blocks that get resolved at install time, so each agent ends up with optimized instructions rather than generic placeholders:
{{#if interactive_prompts}} Present findings using {{tool:interactive_prompts}} with these options: - "Fix all" - "Let me pick" {{else}} List findings and proceed with fixing all. {{/if}}post_process_skill_content()already does per-agent content transformation, but only for skills-format agents. Could this be extended to all format types via apost_process_command_content()onIntegrationBase?For script discovery, something like
__PLUGIN_ROOT__inprocess_template()would let extensions write"__PLUGIN_ROOT__/scripts/script.sh"instead of hardcoding agent-specific search paths.I also noticed the workflow expressions engine (
expressions.py) already implements a safe Jinja2 subset with conditionals, though it's currently only used by the workflow YAML engine. Could that be reused here?Questions
registrar_config?[edited, my agent was too quick 😬 ]