Skip to content

[v4.0.0] Phase 15: Remediation Policy Safety Gate & Interactive Confirmation #286

Description

@utkarsh232005

Overview

Implement the Remediation Policy Engine and interactive safety gate requiring explicit human confirmation before executing any cluster mutation.

Part of Milestone v4.0.0 — Phase 15 of the Multi-Agent SRE Architecture.


⚠️ IMPORTANT CONTRIBUTOR & PR INSTRUCTIONS

TARGET BRANCH: All Pull Requests implementing this phase MUST target branch v4.0.0 (DO NOT target main).
PR TITLE: feat(security): Phase 15 - Remediation Policy Safety Gate & Interactive Confirmation
PR SCOPE: Security policy and interactive confirmation prompt.


1. Description of What Has to Be Done

Architectural Rule: Diagnosis is automatic; mutation is deliberate and gated.
Under NO circumstances may KDM execute a mutating cluster command without an explicit, interactive affirmative confirmation from the user. Furthermore, destructive commands must be permanently blocked by an immutable policy engine.

Contributors must:

  1. Implement PolicyEngine with deterministic regex rules that block dangerous commands (delete namespace, delete node, rm -rf, --force, --grace-period=0).
  2. Enforce a strict command whitelist: only approved prefixes (kubectl set resources, kubectl scale, kubectl rollout restart, kubectl patch) are allowed.
  3. Implement an interactive terminal prompt in React Ink (src/remediation/safety-prompt.ts) showing the command, risk badge, and a [y/N] prompt.
  4. The default keypress (Enter or n) MUST abort execution. Only an explicit y or Y keypress allows execution to proceed.

2. Desired Outcome & Expected Behavior

Expected Terminal View

┌─ Remediation Proposal ───────────────────────────────────────────────────┐
│ Action: Increase Container Memory Limit                                  │
│ Risk:   LOW · Safe rollback available                                    │
│ Command:                                                                 │
│   $ kubectl set resources deployment checkout-api --limits=memory=512Mi  │
│                                                                          │
│ Diff Preview:                                                            │
│   - limits.memory: 256Mi                                                 │
│   + limits.memory: 512Mi                                                 │
│                                                                          │
│ Apply this remediation to the cluster? [y/N]:                            │
└──────────────────────────────────────────────────────────────────────────┘

3. The Implementation Plan & Architectural Blueprint

Files to Create and Modify

  • [NEW] agents/remediation/policy.py: Command validation policy engine.
  • [NEW] src/remediation/safety-prompt.ts: Ink confirmation prompt.
  • [NEW] agents/tests/test_policy.py: Policy rejection test suite.
  • [NEW] src/__tests__/safety-prompt.test.ts: UI prompt test suite.

4. Detailed Task Breakdown

  • Task 15.1: Implement PolicyEngine (policy.py)

    • Implement regex blacklist: �delete\s+namespace�, �delete\s+node�, �rm\s+-rf�, �--force�.
    • Implement whitelist prefixes: kubectl set resources, kubectl scale, kubectl rollout restart, kubectl patch.
    • Method evaluate(command: str) -> bool: returns False if command matches any blacklisted pattern or lacks whitelisted prefix.
  • Task 15.2: Implement Interactive Safety Prompt (safety-prompt.ts)

    • Render proposed command and diff preview.
    • Capture single keypress. Default Enter -> abort.
  • Task 15.3: Security Tests (test_policy.py)

    • Test blocked commands: kubectl delete namespace kube-system -> Rejected.
    • Test blocked options: kubectl delete pod api --force --grace-period=0 -> Rejected.
    • Test valid commands: kubectl set resources deployment api --limits=memory=512Mi -> Approved.

5. Technical Specifications & Concrete Code Signatures

# agents/remediation/policy.py
import re

BLOCKED_PATTERNS = [
    r"�delete\s+namespace�",
    r"�delete\s+node�",
    r"�rm\s+-rf�",
    r"�--force�",
    r"�--grace-period=0�"
]

ALLOWED_COMMAND_PREFIXES = [
    "kubectl set resources",
    "kubectl scale",
    "kubectl rollout restart",
    "kubectl patch"
]

class PolicyEngine:
    def evaluate(self, command: str) -> bool:
        for pattern in BLOCKED_PATTERNS:
            if re.search(pattern, command, re.IGNORECASE):
                return False
        return any(command.strip().startswith(prefix) for prefix in ALLOWED_COMMAND_PREFIXES)

6. Verification & Acceptance Checklist

  • Run python3 -m unittest agents/tests/test_policy.py — all tests pass.
  • Destructive commands proven to be blocked.
  • Interactive prompt defaults to abort.
  • PR targets branch v4.0.0.

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions