This project compares backend implementation performance of model–harness combinations using Ralph loops. Its goal is practical selection by task, quality and budget. The original Ralph vs plan-then-execute experiments remain historical evidence.
- Unit of comparison: model × harness × reasoning settings × provider/connection environment, including Claude Code and Codex. Combination differences are not isolated model effects.
- Baseline task: RealWorld backend. Planned extensions cover bug fixes, feature additions and database migrations in existing code.
- Primary metrics: budget-constrained pass rate, cost including failures, time and human intervention. Tokens are diagnostic, not billed cost.
- Rules and evidence: quality rules, corrections, roadmap.
- Next pilot: Claude Code + Fable 5.1 vs Codex + Astra. Design.
Model–harness combinations and task types are the priority. Strategy (S), token habits (H) and language (L) remain diagnostic axes in the catalog.
The table and per-experiment summaries below are generated by
scripts/update_readme_results.pyfrom each experiment'sreport.md(translations for the English, Japanese and Chinese READMEs come fromscripts/readme_i18n.json). A pre-commit hook runs it whenever areport.mdis committed (manual run:python3 scripts/update_readme_results.py).
📊 Live dashboard: Ralph loop completion by model — latest experiment included: EXP-038 (2026-10-08)
The external dashboard does not yet include the 2026-09-21 corrections. Use the reports and correction record below for numerical conclusions.
| Experiment | Hypothesis | Verdict | Client |
|---|---|---|---|
| EXP-001 Ralph loop vs Plan-then-execute | S-01: Plan-then-execute uses fewer tokens than a Ralph loop on the same task | Rejected within observed scope | Claude Code 2.1.215 |
| EXP-002 Korean vs English pipeline token comparison | L-01: Running the whole pipeline in English cuts token proxy tokens meaningfully versus Korean | Inconclusive | Claude Code 2.1.216 |
| EXP-003 PTE + skill-style progressive disclosure | S-02: Structuring context per the official skill guidance (documents under 200 lines, load only what is needed via skills) cuts token proxy by 30%+ versus the EXP-001 PTE | Threshold not met (corrected) | Claude Code 2.1.216 |
| EXP-004 Ralph loop + skill structure | S-03: Giving a single-session Ralph loop domain-contract skills reduces token proxy | Inconclusive | Claude Code 2.1.216 |
| EXP-005 Claude Code × Upstage Solar Pro 3 backend | M-01: Swapping the Claude Code backend to Solar Pro 3 lets it finish the same task (RealWorld backend) without intervention, and the total cost on completion is meaningfully lower than Opus. | Pending | Claude Code 2.1.220 + claude-code-router 1.0.73 |
| EXP-006 Claude Code × Upstage Solar Open 2 backend | M-02: Swapping the Claude Code backend to Solar Open 2 lets it finish the same task (RealWorld backend) without intervention, and the total cost on completion is meaningfully lower than Opus. | Pending | Claude Code 2.1.220 (claude-code-router 1.0.73: open2-1 only) |
| EXP-007 Autopsy of the Solar Open 2 non-completion | M-03: The EXP-006 (Solar Open 2) non-completion is not a single convergence-speed bottleneck but an overlap of several failure factors (model behaviour defects · experiment environment contamination · metering distortion). | Verified | N/A (post-hoc analysis of EXP-006 sessions) |
| EXP-008 Solar Open 2 uncontaminated clean run — completion check | M-04: With contamination removed (isolated config), no disturbance and a 30-iteration cap, solar-open2 can finish the RealWorld backend (Hurl 154/154) in a Ralph loop without intervention (cost excluded; completion is the single verdict). | Verified | Claude Code 2.1.220 |
| EXP-009 Opus 5 Ralph loop (EXP-002 en condition, n=3) | M-05: Opus 5 reproduces single-session completion in the EXP-002 en Ralph loop and shows efficiency equal to or better than the Opus 4.x baseline (en 6–7 min · 38–54 API calls). | Partially verified (n=3) | Claude Code 2.1.220 |
| EXP-010 Opus 4.8 vs Opus 5 pure A/B (same day, n=3 each) | M-06: Under fully identical conditions, Opus 5's expanded-output profile (output · commits ↑) reproduces against Opus 4.8 and both models keep single-session completion. | Verified | Claude Code 2.1.220 |
| EXP-011 Codex CLI × gpt-5.6-sol RealWorld backend completion check | M-07: In the Codex CLI (codex exec) harness, gpt-5.6-sol (effort medium) can finish the RealWorld backend (Hurl 13/13 · 154/154) within a 30-iteration cap in an isolated, undisturbed Ralph loop without intervention. |
Verified | Codex CLI 0.144.0 |
| EXP-012 Claude Code × gpt-5.6-sol backend (ccr) RealWorld backend completion check | M-08: Connecting the Claude Code backend to gpt-5.6-sol through ccr (reasoning effort medium) can finish the RealWorld backend (Hurl 13/13 · 154/154) within a 30-iteration cap in an isolated, undisturbed Ralph loop without intervention. | Verified | Claude Code 2.1.220 + claude-code-router 1.0.73 |
| EXP-013 Claude Code × qwen3.8-max direct (ANTHROPIC_BASE_URL) RealWorld backend completion check | M-09: Connecting Claude Code directly to qwen3.8-max through DashScope's Anthropic-compatible endpoint (thinking default) can finish the RealWorld backend (Hurl 13/13 · 154/154) within a 30-iteration cap in an isolated, undisturbed Ralph loop without intervention. | Verified | Claude Code 2.1.220 |
| EXP-014 Claude Code × kimi-k3 direct (ANTHROPIC_BASE_URL) RealWorld backend completion check | M-10: Connecting Claude Code directly to kimi-k3 through Moonshot's Anthropic-compatible endpoint (thinking default) can finish the RealWorld backend (Hurl 13/13 · 154/154) within a 30-iteration cap in an isolated Ralph loop without intervention. | Verified | Claude Code 2.1.220 |
| EXP-015 Claude Code × solar-open2 direct (ANTHROPIC_BASE_URL) RealWorld backend completion check | M-11: Connecting Claude Code directly to solar-open2 through Upstage's Anthropic-compatible endpoint (thinking default) can finish the RealWorld backend (Hurl 13/13 · 154/154) within a 30-iteration cap in an isolated, undisturbed Ralph loop without intervention. | Verified | Claude Code 2.1.220 |
| EXP-016 Replication of the three n=1 completion conditions (n=3 each) | M-12: The unattended completions of the three EXP-011/013/014 conditions (Codex CLI × gpt-5.6-sol, qwen3.8-max direct, kimi-k3 direct) reproduce: two additional runs per condition (n=3 total) all complete within the 30-iteration cap with gate and independent re-check agreeing. | Verified | Claude Code 2.1.220 (kimi, qwen), Codex CLI 0.144.0 (sol) |
| EXP-017 qwen3.8-max · kimi-k3 Korean-condition Ralph loop completion check (n=3 each) | L-02: qwen3.8-max and kimi-k3 (direct, thinking default) can finish the RealWorld backend (Hurl 13/13 · 154/154) within a 30-iteration cap in an isolated, undisturbed Ralph loop without intervention even under the Korean canonical prompt (including the all-deliverables-in-Korean instruction). | Verified | Claude Code 2.1.221 |
| EXP-018 solar-open2 direct replication — impossible after provider endpoint withdrawal | M-13: The EXP-015 solar-open2 direct unattended completion reproduces: two additional runs (n=3 total) both complete within 30 iterations with gate (measure v4) and re-check agreeing. | Pending | Claude Code 2.1.222 |
| EXP-019 Native Opus 4.8 · Opus 5 · Codex×gpt-5.6-sol Korean-condition completion check (n=3 each) | L-03: Native Opus 4.8 and Opus 5 (Claude Code) and gpt-5.6-sol (Codex CLI) can finish the RealWorld backend (Hurl 13/13 · 154/154) within the iteration cap in an isolated, undisturbed Ralph loop without intervention even under the Korean canonical prompt (including the all-deliverables-in-Korean instruction). | Verified | Codex CLI 0.144.0 (solko), Claude Code not recorded (Opus runs) |
| EXP-020 Claude Code × solar-pro4 direct completion check (n=3) | M-14: Connecting Claude Code directly to solar-pro4 through Upstage's Anthropic-compatible endpoint (thinking default) can finish the RealWorld backend (Hurl 13/13 · 154/154) within a 30-iteration cap in an isolated, undisturbed Ralph loop without intervention (n=3, completion-rate verdict, cost excluded). | Verified | Claude Code 2.1.235 |
| EXP-021 Codex CLI × gpt-6-astra Ralph loop completion check (n=3) | M-15: In the Codex CLI (codex exec) harness, gpt-6-astra (effort medium) can finish the RealWorld backend (Hurl 13/13 · 154/154) within a 30-iteration cap in an isolated, undisturbed Ralph loop without intervention (n=3, completion-rate verdict, cost excluded). |
Verified | Codex CLI 0.153.4 |
| EXP-023 Claude Code × Fable 5.1 native Ralph loop completion check (n=3) | M-17: In the native Claude Code harness, Fable 5.1 (claude-fable-5-1, thinking default) can finish the RealWorld backend (Hurl 13/13 · 154/154) within a 10-iteration cap in an EN-canonical Ralph loop without intervention (n=3, completion-rate verdict, cost excluded). |
Verified | Claude Code 2.1.273 |
| EXP-025 Claude Code × DeepSeek V4.1-Flash · V4-Pro direct Ralph loop completion check (EN · KO, n=3 each) | M-18: Connecting Claude Code directly to deepseek-flash (DeepSeek-V4.1-Flash) and deepseek-v4-pro (DeepSeek-V4-Pro-0813) through DeepSeek's Anthropic-compatible endpoint (thinking default) can finish the RealWorld backend (Hurl 13/13 · 154/154) within a 30-iteration · 4-hour cap in an isolated, undisturbed Ralph loop without intervention (EN and KO canon, n=3 each, completion-rate verdict, cost excluded). |
Verified | Claude Code 2.1.278 |
| EXP-026 Claude Code × Opus 5.5 native Ralph loop completion check (EN · KO, n=3 each) | M-19: In the native Claude Code harness, Opus 5.5 (claude-opus-5-5, thinking default) can finish the RealWorld backend (Hurl 13/13 · 154/154) within a 10-iteration · 4-hour cap in a Ralph loop without intervention, with both the EN and KO canon prompts (n=3 each, completion-rate verdict, cost excluded). |
Verified | Claude Code 2.1.280 |
| EXP-027 Codex CLI × gpt-6-sol · gpt-6-luna Ralph loop completion check (n=3 each) | M-20: In the Codex CLI (codex exec) harness, gpt-6-sol and gpt-6-luna (effort medium) can each finish the RealWorld backend (Hurl 13/13 · 154/154) within a 30-iteration · 4-hour cap in an isolated Ralph loop without intervention (n=3 each). |
Verified | Codex CLI 0.155.1 (sol-1, sol-2, luna-1), 0.156.1 (sol-3, luna-2, luna-3) |
| EXP-028 Claude Code × Sonnet 5 native Ralph loop completion check (EN·KO n=3 each) | M-21: In the native Claude Code harness, Sonnet 5 (claude-sonnet-5, default thinking) can finish the RealWorld backend (Hurl 13/13 · 154/154) with the EN and KO canonical Ralph-loop prompts within a 10-iteration · 4-hour cap without intervention (EN·KO n=3 each). |
Verified | Claude Code 2.1.281 |
| EXP-029 pi coding agent × kimi-k3 · qwen3.8-max · deepseek-flash Ralph loop completion check (EN n=3 each) | M-22: With the pi coding agent (pi -p v0.87.1) connected directly to each provider's OpenAI-compatible endpoint, kimi-k3, qwen3.8-max and deepseek-flash (pi default thinking) can finish the RealWorld backend (Hurl 13/13 · 154/154) in an isolated, intervention-free Ralph loop within a 30-iteration · 4-hour cap (EN n=3 each; completion-rate verdict, billing excluded). |
Verified | pi 0.87.1 |
| EXP-030 pi coding agent × Opus 5.5 · gpt-6-sol Ralph loop completion check (EN n=3 each) | M-23: With the pi coding agent (pi -p v0.87.1) running anthropic/claude-opus-5-5 (Anthropic API key, direct) and openai-codex/gpt-6-sol (ChatGPT OAuth) at pi default thinking, both can finish the RealWorld backend (Hurl 13/13 · 154/154) in an isolated, intervention-free Ralph loop within a 30-iteration · 4-hour cap (EN n=3 each; completion-rate verdict, billing excluded). |
Verified | pi 0.87.1 |
| EXP-031 Claude Code × Sonnet 5.5 native Ralph loop completion check (EN · KO n=3 each) | M-24: On the Claude Code native harness, Sonnet 5.5 (claude-sonnet-5-5, default thinking) can finish the RealWorld backend (Hurl 13/13 · 154/154) with the canonical EN and KO Ralph-loop prompts, intervention-free, within a 10-iteration · 4-hour cap (EN · KO n=3 each; completion-rate verdict, billing excluded). |
Verified | Claude Code 2.1.284 |
| EXP-032 pi coding agent × Sonnet 5.5 Ralph loop completion check (EN n=3) | M-25: With the pi coding agent (pi -p v0.87.1) running anthropic/claude-sonnet-5-5 (Anthropic API key, direct; pi default thinking; a custom model entry copied from pi's built-in claude-sonnet-5 definition), it can finish the RealWorld backend (Hurl 13/13 · 154/154) in an isolated, intervention-free Ralph loop within a 30-iteration · 4-hour cap (EN n=3; completion-rate verdict, billing excluded). |
Verified | pi 0.87.1 |
| EXP-033 Codex CLI × gpt-6.1-sol Ralph loop completion check (EN n=3) | M-26: In the Codex CLI (codex exec 0.160.0) harness, gpt-6.1-sol (effort medium) can finish the RealWorld backend (Hurl 13/13 · 154/154) in an isolated, intervention-free Ralph loop within a 30-iteration · 4-hour cap (EN n=3; completion-rate verdict, billing excluded). |
Verified | Codex CLI 0.160.0 |
| EXP-034 pi coding agent × gpt-6.1-sol Ralph loop completion check (EN n=3) | M-27: With the pi coding agent (pi -p 0.87.1) running openai-codex/gpt-6.1-sol (ChatGPT OAuth, pi default thinking, a custom model entry copied from pi's built-in gpt-6-sol definition), it can finish the RealWorld backend (Hurl 13/13 · 154/154) in an isolated, intervention-free Ralph loop within a 30-iteration · 4-hour cap (EN n=3; completion-rate verdict, billing excluded). |
Verified | pi 0.87.1 |
| EXP-035 Antigravity CLI × gemini-3.8-flash Ralph loop completion check (EN n=3) | M-28: With Antigravity CLI (agy -p) running gemini-3.8-flash (effort high, Gemini API key), it can finish the RealWorld backend (Hurl 13/13 · 154/154) in an isolated, intervention-free Ralph loop within a 30-iteration · 4-hour cap (EN n=3; completion-rate verdict, billing excluded). |
Verified | Antigravity CLI (agy) 1.3.1 |
| EXP-036 pi coding agent × gemini-3.8-flash Ralph loop completion check (EN n=3) | M-29: With the pi coding agent (pi -p 0.87.1) running google/gemini-3.8-flash (thinking high, Gemini API key), it can finish the RealWorld backend (Hurl 13/13 · 154/154) in an isolated, intervention-free Ralph loop within a 30-iteration · 4-hour cap (EN n=3; completion-rate verdict, billing excluded). |
Verified | pi 0.87.1 |
| EXP-037 Claude Code × Haiku 5.5 native Ralph loop completion check (EN n=3) | M-30: In the Claude Code native harness, Haiku 5.5 (claude-haiku-5-5, default thinking) can finish the RealWorld backend (Hurl 13/13 · 154/154) with the canonical EN Ralph loop, intervention-free, within a 30-iteration · 4-hour cap (EN n=3; completion-rate verdict, billing excluded). |
Verified | Claude Code 2.1.293 |
| EXP-038 pi coding agent × Haiku 5.5 Ralph loop completion check (EN n=3) | M-31: With the pi coding agent (pi -p 0.87.1) running anthropic/claude-haiku-5-5 (Anthropic API key, default pi thinking, custom model entry copied from the built-in claude-sonnet-5), it can finish the RealWorld backend (Hurl 13/13 · 154/154) in an isolated, intervention-free Ralph loop within a 30-iteration · 4-hour cap (EN n=3; completion-rate verdict, billing excluded). |
Verified | pi 0.87.1 |
EXP-001 — Ralph loop vs Plan-then-execute (Rejected within observed scope)
Deduplicated token proxy: PTE 1,128,420 vs Ralph 136,506 (8.27×). Different grading suites prevent an equal-quality cost conclusion. → report
EXP-002 — Korean vs English pipeline token comparison (Inconclusive)
KO mean 114,605.5 vs EN 112,463.5. The 2,142 difference is below the largest within-condition range (23,033). → report
EXP-003 — PTE + skill-style progressive disclosure (Threshold not met (corrected))
Token proxy fell 22.45% (1,128,420 → 875,083), below the preregistered 30% threshold. The earlier verified verdict is withdrawn. → report
EXP-004 — Ralph loop + skill structure (Inconclusive)
Skills runs: 117,352 / 118,311; mean +2.81% vs KO baseline. Range is 959; claims of no effect and twofold variability are withdrawn. → report
EXP-005 — Claude Code × Upstage Solar Pro 3 backend (Pending)
solar-1 did not finish (0 test runs, 0 commits, stopped early at iteration 6/15): the integration stack was validated, but in the headless autonomous loop the permission-wait and context-overflow failure modes repeated and it never reached a completion trajectory. → report
EXP-006 — Claude Code × Upstage Solar Open 2 backend (Pending)
0/2 completions, but only one run followed the full protocol (open2-1 self-terminated after 1 iteration with a false completion claim): open2-2 used all 15 iterations and reached only 3/13 files (94/154 requests) on independent verification, yet it established the autonomous TDD loop that Solar Pro 3 lacked and converged monotonically — the “behaviour layer” bottleneck moved from autonomy to convergence speed. → report
EXP-007 — Autopsy of the Solar Open 2 non-completion (Verified)
Three layers demonstrated: ① metering distortion (usage overstated 3.07× — actually 483 requests, 23.3M input, an estimated ~$3.8, lower than Opus at $6.41), ② environment contamination (superpowers hooks and the global CLAUDE.md injection ate at least 3 iterations), ③ model behaviour defects (declare-without-execute leaving 0 commits, 25 thinking-only truncations, 2 off-task hallucinations). → report
EXP-008 — Solar Open 2 uncontaminated clean run — completion check (Verified)
Completed at iteration 10/30: .ralph-done created → harness gate passed 13/13 files · 154/154 requests → two independent re-checks by the experimenter agree. About 2 h 53 min wall-clock, no intervention or interruption, 4 git commits (in Korean). Completion came inside EXP-006's cap (15), so the deciding variable was contamination removal and non-disturbance, not a higher cap. → report
EXP-009 — Opus 5 Ralph loop (EXP-002 en condition, n=3) (Partially verified (n=3))
Completion clause verified: all 3/3 runs completed in a single session at iteration 1 (9:03–12:22, gate 13/13 · 154/154 plus two independent re-checks each, 6–7 commits). Efficiency clause confirmed unmet: the time range (8.9–12.2 min) does not overlap 4.x (5.8–6.9 min) — but the cause is not serving speed; it is a behaviour-profile shift with consistently higher output (+62%). → report
EXP-010 — Opus 4.8 vs Opus 5 pure A/B (same day, n=3 each) (Verified)
6/6 completions (every run at iteration 1, gate 13/13 · 154/154). All pre-registered metrics met: non-overlapping output-token ranges (4.8: 29.6–37.1K vs 5: 41.1–47.9K, +37% mean) · non-overlapping git-commit ranges (1–2 vs 4–8), same direction as EXP-009 (5 > 4.8). The generational difference is a real profile, not a timing or metering artefact. → report
EXP-011 — Codex CLI × gpt-5.6-sol RealWorld backend completion check (Verified)
Completed at iteration 1 (gate 13/13 · 154/154 plus two matching independent re-checks, codex exec 5 min 46 s · 1 session · 3 commits, no intervention). → report
EXP-012 — Claude Code × gpt-5.6-sol backend (ccr) RealWorld backend completion check (Verified)
Completed at iteration 11/30 (gate 13/13 · 154/154 plus two matching independent re-checks, 58 min total · 11 commits). One harness-level intervention at iteration 2 for a stream stall (process killed to recover the iteration boundary; model output untouched) — see the protocol issues below. → report
EXP-013 — Claude Code × qwen3.8-max direct (ANTHROPIC_BASE_URL) RealWorld backend completion check (Verified)
Completed at iteration 1 (gate 13/13 · 154/154 plus two matching independent re-checks, 15 min 5 s · 4 commits · 0 interventions, zero direct-connection troubleshooting). → report
EXP-014 — Claude Code × kimi-k3 direct (ANTHROPIC_BASE_URL) RealWorld backend completion check (Verified)
Completed at iteration 1 (gate 13/13 · 154/154 plus two matching independent re-checks, 21 min 18 s · 4 commits · 0 interventions, zero direct-connection troubleshooting). → report
EXP-015 — Claude Code × solar-open2 direct (ANTHROPIC_BASE_URL) RealWorld backend completion check (Verified)
Completed at iteration 2 (two direct re-grades of the output both 13/13 · 154/154, 58 min 15 s · 2 commits. Footnote: one gate false rejection — the measure v3 port-detection flaw rejected a legitimate completion; the EXP-010 48-1 precedent was applied). → report
EXP-016 — Replication of the three n=1 completion conditions (n=3 each) (Verified)
All 6 additional runs completed (gate pass plus two independent re-checks each at 13/13 · 154/154, 0 interventions). Including the original runs, completion is 3/3 per condition, 9/9 combined. → report
EXP-017 — qwen3.8-max · kimi-k3 Korean-condition Ralph loop completion check (n=3 each) (Verified)
All 6/6 runs completed at iteration 1 (gate pass plus two independent re-checks each at 13/13 · 154/154, 0 interventions). Language compliance also holds: every run's commit messages and README are in Korean. → report
EXP-018 — solar-open2 direct replication — impossible after provider endpoint withdrawal (Pending)
External-to-model failure (pre-registered criterion): at the start of the experiment Upstage withdrew the Anthropic-compatible endpoint (/v1/messages) and the solar-open2 hosted API, so the runs themselves were impossible. The hypothesis is unverifiable, not rejected. → report
EXP-019 — Native Opus 4.8 · Opus 5 · Codex×gpt-5.6-sol Korean-condition completion check (n=3 each) (Verified)
All 9/9 runs completed at iteration 1 (gate pass plus two independent re-checks each at 13/13 · 154/154, 0 interventions). Language compliance holds across the board: all 9 runs have Korean commit messages and READMEs (including the English-first model gpt-5.6-sol). → report
EXP-020 — Claude Code × solar-pro4 direct completion check (n=3) (Verified)
3/3 completions (gate pass plus two independent re-checks each at 13/13 · 154/154, 0 model interventions). pro4-1's loop was extended by a gate false rejection (harness fault) — retroactive re-grading fixed its effective completion at iter 3. → report
EXP-021 — Codex CLI × gpt-6-astra Ralph loop completion check (n=3) (Verified)
All 3/3 runs completed at iteration 1 (gate pass plus two independent re-checks each at 13/13 · 154/154, 0 interventions, 7-minute sessions · 3–4 commits). → report
EXP-023 — Claude Code × Fable 5.1 native Ralph loop completion check (n=3) (Verified)
All 3/3 runs completed at iteration 1 (gate pass plus two independent re-checks each at 13/13 · 154/154, 0 interventions, 6–8-minute sessions · 2–4 commits). Secondary metrics moved down, with non-overlapping time and output ranges versus Opus 5 (EXP-009/010): 6.2–7.8 min vs 8.9–17.6 min, 28.7–35.8K vs 41.1–48.4K — Opus 5's expanded-output profile returns to the 4.8 level in Fable 5.1. → report
EXP-025 — Claude Code × DeepSeek V4.1-Flash · V4-Pro direct Ralph loop completion check (EN · KO, n=3 each) (Verified)
All 12/12 runs completed at iteration 1 (3/3 in each of the four conditions, gate pass plus two independent re-checks each at 13/13 · 154/154, 0 interventions, the response model field matched on every call). Flash: EN 4.6–6.0 min, the fastest band of any condition, about $0.1 per run at list price; Pro: 12–16 min, about $0.5. Observed values, not a proof of general success rate or efficiency advantage. → report
EXP-026 — Claude Code × Opus 5.5 native Ralph loop completion check (EN · KO, n=3 each) (Verified)
All 6/6 runs completed at iteration 1 (EN 3/3 · KO 3/3, gate pass plus two independent re-checks each at 13/13 · 154/154, 0 interventions, response model field claude-opus-5-5 on every call). Sessions 4.0–8.4 min (5 of 6 runs at 4.0–4.2 min) and output 18.6–28.2K — the shortest and lowest band among native Claude conditions, below Opus 5 (8.9–17.6 min · 39–49K) and Fable 5.1 (6.2–7.8 min · 28.7–35.8K). Observed values, not a confirmed advantage. → report
EXP-027 — Codex CLI × gpt-6-sol · gpt-6-luna Ralph loop completion check (n=3 each) (Verified)
sol 3/3 · luna 3/3 completed at iteration 1 (gate pass plus two independent re-checks each at 13/13 · 154/154, response model field matched on every run). Sessions: sol 4.8–5.2 min · output 10.4–11.0K; luna 5.4–10.0 min · output 13.4–21.5K. One grader-infrastructure hang (D-1) was handled without touching the agent and did not affect the verdict. First experiment run with the ralph-model-benchmark skill. Observed values, not a confirmed ranking. → report
EXP-028 — Claude Code × Sonnet 5 native Ralph loop completion check (EN·KO n=3 each) (Verified)
EN 3/3 · KO 3/3 completed (gate pass plus two independent re-checks each at 13/13 · 154/154, no intervention, response model field claude-sonnet-5 on every message). Five runs finished at iteration 1; en-1 split the work across three iterations on its own and finished at iteration 3 (no rejected claims). Sessions 11.2–16.8 min · output 57.6–73.7K, above the ranges seen for Opus 5.5 (EXP-026) and Fable 5.1 (EXP-023) in the same harness. Observed values, not a confirmed ranking. → report
EXP-029 — pi coding agent × kimi-k3 · qwen3.8-max · deepseek-flash Ralph loop completion check (EN n=3 each) (Verified)
All three conditions 3/3 completed (gate pass plus two independent re-checks each at 13/13 · 154/154, no intervention, response model field matched on every message). Eight runs finished at iteration 1; qwen-en-1 was recorded at iteration 5 because an external process occupied the grading port (its iteration-1 code also passed on re-grading). Sessions: flash 1.9–4.1 min (after the D-3 re-run), kimi 9.5–9.9 min, qwen 15.6–29.0 min. First run of a third harness (pi) in this repo; comparison with earlier Claude Code direct runs mixes harness and API format. Observed values, not a confirmed ranking. → report
EXP-030 — pi coding agent × Opus 5.5 · gpt-6-sol Ralph loop completion check (EN n=3 each) (Verified)
Both conditions 3/3, all six runs at iteration 1 (gate pass plus two independent re-checks each at 13/13 · 154/154, no intervention, response model field matched on every message). Sessions: Opus 5.5 3.1–4.9 min, gpt-6-sol 4.2–5.6 min, overlapping the native-agent baselines (EXP-026 Opus 5.5 4.0–8.4 min, EXP-027 gpt-6-sol 4.8–5.2 min). Observed values, not a confirmed ranking. → report
EXP-031 — Claude Code × Sonnet 5.5 native Ralph loop completion check (EN · KO n=3 each) (Verified)
EN 3/3 · KO 3/3, all six runs at iteration 1 (gate pass plus two independent re-checks each at 13/13 · 154/154, no intervention, response model field claude-sonnet-5-5 on every message). Sessions 2.5–4.2 min, output 15.2–20.9K, 16–24 API calls, Bash only — shorter and fewer than Opus 5.5 · Fable 5.1 · Sonnet 5 on the same native harness. First run with harness-injected free PORT. Observed values confounded by date and CLI version, not a confirmed ranking. → report
EXP-032 — pi coding agent × Sonnet 5.5 Ralph loop completion check (EN n=3) (Verified)
EN 3/3, all at iteration 1 (gate pass plus two independent re-checks each at 13/13 · 154/154, no intervention, response model field claude-sonnet-5-5 on every message). Sessions 1.8–3.2 min, output 10.9–14.7K and 13–23 requests — shorter and fewer than same-day Claude Code × Sonnet 5.5 (EXP-031 EN 2.9–4.2 min, 15.2–20.0K) and pi × Opus 5.5 (EXP-030 3.1–4.9 min). The agent and API path differ together, so this is a difference of the whole combination, not a confirmed ranking. → report
EXP-033 — Codex CLI × gpt-6.1-sol Ralph loop completion check (EN n=3) (Verified)
EN 3/3, all at iteration 1 (gate pass plus two independent re-checks each at 13/13 · 154/154, response model field gpt-6.1-sol on all 15 records). Sessions 6.2–9.8 min and output 10.0–16.2K — longer than gpt-6-sol on the same harness (EXP-027, 4.8–5.2 min). Codex had no local metadata for the model (fallback) and the CLI version and date differ, so this is not a confirmed model ranking. → report
EXP-034 — pi coding agent × gpt-6.1-sol Ralph loop completion check (EN n=3) (Verified)
EN 3/3, all at iteration 1 (gate pass plus two independent re-checks each at 13/13 · 154/154, response model field gpt-6.1-sol on all 96 records, thinking medium). Sessions 8.3–9.2 min and output 12.5–13.3K — longer than pi × gpt-6-sol (EXP-030, 4.2–5.6 min). Both harnesses moved the same way on the same day, but the date and custom model entry differ, so this is not a confirmed model ranking. → report
EXP-035 — Antigravity CLI × gemini-3.8-flash Ralph loop completion check (EN n=3) (Verified)
EN 3/3 completed (runs 1–2 at iteration 1, run 3 at iteration 2 because the harness grader that the agent itself ran killed agy — D-1). Gate pass plus two independent re-checks each at 13/13 · 154/154; every run authenticated with the Gemini API key and resolved to gemini-3.8-flash-high. Sessions 15.5–22.4 min. → report
EXP-036 — pi coding agent × gemini-3.8-flash Ralph loop completion check (EN n=3) (Verified)
EN 3/3, all at iteration 1 (gate pass plus two independent re-checks each at 13/13 · 154/154, all 333 responses gemini-3.8-flash, thinking high). Sessions 13.1–15.6 min. Run 3 is flagged because the agent read harness files and earlier runs' logs (D-2). → report
EXP-037 — Claude Code × Haiku 5.5 native Ralph loop completion check (EN n=3) (Verified)
EN 3/3 completed (runs 2 and 3 at iteration 1; run 1 at iteration 2 after iteration 1 stopped on its own after scaffolding). Gate pass plus two independent re-checks each at 13/13 · 154/154, every response claude-haiku-5-5. Sessions 5.1–7.9 min, output 49.0–57.6K, estimated $0.06–0.16 per run at list prices, the lowest of the Claude conditions. Run 1 called the exposed context7 MCP. → report
EXP-038 — pi coding agent × Haiku 5.5 Ralph loop completion check (EN n=3) (Verified)
EN 3/3, all at iteration 1 (gate pass plus two independent re-checks each at 13/13 · 154/154, every response claude-haiku-5-5, thinking medium). Sessions 5.4–5.9 min, output 50.0–55.4K, estimated $0.05–0.09 per run at list prices. In run 1 the agent's pgrep output put other processes' secrets into the session log (D-1); the archive is masked. → report
- EXP-001's corrected PTE/Ralph token-proxy ratio is 8.27×; different grading suites prevent an equal-quality cost comparison.
- EXP-003's 22.45% reduction misses the 30% threshold. EXP-004's twofold-variability and no-effect claims are withdrawn.
- First-iteration success supports feasibility under the tested conditions. Loop recovery benefits and maintenance performance need separate evaluation.
- Three successful runs or timings from different dates do not establish general rankings or causal model/harness effects. Unmeasured subscription cost does not mean free.
- DeepSeek V4.1-Flash passed the direct-connection standard unmodified and showed the fastest band and lowest cost profile of any condition (EXP-025). With only the three env values swapped, both Flash and V4-Pro completed 3/3 in EN and KO — 12/12 at iter 1 — with the response model field holding on every call. Flash EN 4.6–6.0 min at about $0.1 per run (cache-hit input $0.006/M) sits below Codex×sol (5.3–10.0 min); Pro took 12–16 min at about $0.5 with 1.5× the output. Timing and harness confounds mean no speed or cost superiority is asserted; this stands as the fourth direct-connection reuse case and a profile record.
- Opus 5.5 passed the native harness with only the model ID swapped and showed the shortest, lowest-output profile among native Claude conditions (EXP-026). EN and KO each 3/3, 6/6 at iteration 1 (two re-checks each matched). Five of six runs took 4.0–4.2 min with about 20K output, below both Fable 5.1 (6.2–7.8 min · 28.7–35.8K) and Opus 5 (8.9–17.6 min · 39–49K), and KO stayed in the same band. This fits item 8 (output expansion is specific to Opus 5). The baselines differ in timing and CLI version and this is an n=3 observation, so a speed advantage is not confirmed until a concurrent cross-run re-measurement. (EXP-031's Sonnet 5.5 later moved this shortest/lowest mark again — see item 15.)
- All three GPT-6 models (astra · sol · luna) completed in the Codex harness with only the model ID swapped (EXP-021 · 027). sol and luna each 3/3 at iter 1 (response model field matched on every run). sol showed a tight spread at 4.8–5.2 min · output ~11K across all 3 runs; luna ran 5.4–10.0 min · output 13.4–21.5K, and despite its "fast" positioning no luna run was shorter than sol on this task. Differences from astra (EXP-021, ~7 min) are confounded by date and CLI version and are not attributed to the model. EXP-027 is the first run of the
ralph-model-benchmarkskill, which packages the new-model benchmark procedure. - Sonnet 5 also completed in both EN and KO in the native harness, but ran slower and produced more output than the higher-tier models in the same harness (EXP-028). EN 3/3 · KO 3/3 completed (response model field
claude-sonnet-5on every message). Five runs finished at iteration 1; in en-1 the agent split scaffolding, test setup and implementation across three iterations and finished at iteration 3 (no rejected claims). Sessions 11.2–16.8 min · output 57.6–73.7K sit above the ranges seen for Opus 5.5 (EXP-026, 4.0–8.4 min · 18.6–28.2K) and Fable 5.1 (EXP-023). Measurement date and Claude Code version differ, so this is not attributed to the model alone. - A third harness, pi, also completed with three open-weight-family models without a translation layer (EXP-029). With the pi coding agent connected directly to each provider's OpenAI-compatible endpoint, kimi-k3, qwen3.8-max and deepseek-flash each completed EN 3/3, 9/9 in total (response model field matched on every message, zero harness troubleshooting). Eight runs finished at iteration 1; qwen-en-1 was recorded at iteration 5 because an external process occupied the grading port (its iteration-1 code also passed on re-grading). Sessions: flash 1.9–4.1 min (after the D-3 re-run), kimi 9.5–9.9 min, qwen 15.6–29.0 min. Compared with earlier Claude Code direct runs of the same models (EXP-013 · 014 · 025), both the harness and the API format (Anthropic-compatible vs OpenAI-compatible) differ, so differences are not attributed to pi. The Ralph loop benchmark now runs under the same procedure on three harnesses: Claude Code, Codex and pi.
- pi also completed with the vendors' reference models (EXP-030). Running Opus 5.5 (Anthropic API key, direct) and gpt-6-sol (ChatGPT OAuth) under pi gave EN 3/3 each, all six runs at iteration 1. Sessions were 3.1–4.9 min for Opus 5.5 and 4.2–5.6 min for gpt-6-sol, overlapping the native-agent baselines (EXP-026 Claude Code 4.0–8.4 min, EXP-027 Codex 4.8–5.2 min), while output was lower under pi (Opus 5.5 16.6–19.8K vs 18.6–28.2K, gpt-6-sol 7.0–8.6K vs 10.4–11.0K). The agent and the API path (system prompt, tools, how thinking is passed, cache TTL) differ together, so these are differences of the whole combination and are not attributed to pi.
- Sonnet 5.5 finished on the native harness via the shortest path of any Claude condition so far (EXP-031). EN 3/3 · KO 3/3, all six runs at iteration 1 (two re-checks each matched, response model field
claude-sonnet-5-5on every message). Sessions 2.5–4.2 min, output 15.2–20.9K and 16–24 API calls sit below Opus 5.5 (4.0–8.4 min, 22–41 calls), Fable 5.1 and Sonnet 5 (11.2–16.8 min, 82–140 calls) on the same native harness, and all six runs wrote files with Bash alone (no Read/Write/Edit). The gap from the previous Sonnet 5 is large, but measurement date, Claude Code version and the harness port injection (a freePORTinjected starting with this experiment) changed together, so it is not attributed to the model alone. - pi × Sonnet 5.5 recorded the shortest completed session in this repo (EXP-032). Running Sonnet 5.5 through pi (Anthropic API key, direct) finished EN 3/3, all at iteration 1 (two re-checks each matched, response model field
claude-sonnet-5-5on every message). Sessions of 1.8–3.2 min (shortest 1 min 49 s) and output of 10.9–14.7K were shorter and smaller than same-day native Claude Code × Sonnet 5.5 (EN 2.9–4.2 min, 15.2–20.0K). pi 0.87.1 does not list this model, so it was called through a custom entry copied from pi's built-in Sonnet 5 definition, and the agent, thinking settings and cache TTL also differ, so the gap is not attributed to pi alone. - Gemini (gemini-3.8-flash) also completed on both Antigravity CLI and pi, but with longer sessions than other Flash-class models (EXP-035 · 036). Using a Gemini API key, Antigravity CLI (agy 1.3.1, effort high) and pi (thinking high) each went EN 3/3, 6/6 in total. On agy, runs 1–2 finished at iteration 1; run 3 finished at iteration 2 because the harness grader, which the agent ran itself, killed the agy process. Sessions were 15.5–22.4 min on agy and 13.1–15.6 min on pi, with 100–192 model responses per run — far more than pi × deepseek-flash (EXP-029, 25–40 responses, 1.9–4.1 min). Both experiments exposed a structural problem: run directories sit inside the harness directory, so agents can reach the grader and earlier runs' records (also seen in three EXP-029 sessions). Separating them is a follow-up.
- Haiku 5.5 completed in the native harness, with the lowest list-price cost estimate among the Claude conditions (EXP-037). EN 3/3 completed (runs 2 and 3 at iteration 1; run 1 at iteration 2 because iteration 1 ended after scaffolding without a completion declaration). Two re-checks each matched, and every response model field was
claude-haiku-5-5. Sessions 5.1–7.9 min, output 49.0–57.6K, and 38–59 API calls were longer and more than Sonnet 5.5 (EXP-031), but the unit price is 1/20 of Sonnet 5.5, so the estimate is $0.06–0.16 per run (median $0.09). This includes the 5x price for requests with prompts over 100K. Because an API key was exported in the shell, the runs used API-key auth instead of the subscription (D-1). They were actually billed and used a 5-minute cache TTL, unlike the other native conditions, which ran on the subscription. Because the run was not isolated, run 1 called the exposed context7 MCP. Date and CLI version are confounded, so this is not a ranking. - pi × Haiku 5.5 also completed 3/3, all at iteration 1 (EXP-038). Output (50.0–55.4K) and API calls (38–46) were in the same range as Claude Code × Haiku 5.5 on the same day (EXP-037), and sessions varied little (5.4–5.9 min). The starting prompt was about 3K versus about 32K on Claude Code, so only 3 requests went over 100K, and the estimate is $0.05–0.09 per run. In run 1 the agent ran
pgrep -fl, which printed other processes' command lines on the machine, and secrets in them ended up in the session log (D-1; the archive is masked). Removing environment variables alone does not stop an agent from reading other processes' information.
- Copy
templates/experiment-readme.mdtoexperiments/NNN-name/README.mdand write the design (hypothesis, conditions, measurement method, success criteria) - Run sessions per condition and store session logs and measurements under
runs/<condition>/ - Write the token-difference analysis and conclusion in
report.md— keep the header lines- 가설: [code](...) — ...and- **판정: ...** — <one-line summary>(the README generator parses these two lines) - Update the status in
hypotheses/catalog.md(untested → in progress → verified/rejected) - Add the new experiment's English, Japanese and Chinese translations (title · hypothesis · verdict · summary) to
scripts/readme_i18n.json— if missing, the Korean original is inserted into that language's README and the script warns - When
report.mdis committed, the pre-commit hook regenerates the results section of all four READMEs (ko/en/ja/zh-CN). On a fresh clone rungit config core.hooksPath hooksonce (manual refresh:python3 scripts/update_readme_results.py)
├── ideation.md # original ideation (kept as-is)
├── ROADMAP.md # phased roadmap
├── hypotheses/catalog.md # hypothesis catalog + experiment status table
├── experiments/ # one directory per experiment (NNN-name/)
│ └── 001-ralph-vs-plan-then-execute/
├── tasks/ # shared task specs (reused across conditions)
│ └── realworld-backend/
├── templates/ # experiment design / report templates
├── scripts/ # measurement / aggregation wrapper scripts
└── docs/specs/ # design documents