#82 added a "Web tasks: rung 3 lookup" section to the always-on ponytail
SKILL.md, about an external `modern-web` CLI most users won't have installed.
It's optional bloat in the always-on ruleset, and it broke CI by leaving the
.openclaw mirror stale.
Reverts the section from skills/ponytail/SKILL.md, the README callout, and
examples/web-platform-lookup.md, then regenerates the .openclaw mirror and
removes the Spanish callout that #174 had mirrored. Suite green (56/56).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
#82 added the "Web tasks: rung 3 lookup" section to skills/ponytail/SKILL.md
but did not run scripts/build-openclaw-skills.js, so the committed
.openclaw/skills/ponytail/SKILL.md mirror drifted from its source. The two
generator-sync tests in tests/openclaw-skills.test.js have failed on main
since that merge. Regenerated the mirror; suite is green again.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
ponytail-mcp reuses the repo's hooks/ via createRequire("../hooks/..."), which
reaches outside the package dir, so it can only run from a checkout (as its
README says), never as a published npm package — a publish tarball wouldn't
include ../hooks/ and would crash. The `bin` field and missing `private` made
it look publishable. Mark it private so an accidental `npm publish` can't ship
a broken package, and drop the dead bin (you point the host at ponytail-mcp/index.js).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The Spanish README merged (#110) carrying stale content: the old flat
"80-94% menos código" single-shot headline (the exact claim #126 corrected),
no CodeWhale section, badge stuck at 13 agents, and missing the Modern Web
Guidance callout (#82) and the Claude Code desktop-install paragraph.
Re-translates the hero + Números section to the corrected agentic numbers
(~54%, up to 94%, 100% safe) with the old figures demoted to the same
<details> block English uses, adds CodeWhale, fixes the badge, and adds a
"community translation, English is the reference" note. Also adds a minimal
Español discoverability link to the English README so readers can find it.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Documents the desktop-app install flow (no /plugin command) and the global command-dir linking needed for /ponytail commands in OpenCode outside a checkout. Covers the recurring questions in #97 and #98.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The hero showed only ~54% (the mean); the rigorous agentic run also reaches 94%
on the over-build tasks (the date picker), so the headline now reads
"~54% (up to 94%)". The sub-line is reworded so 80-94% reads as the per-task
ceiling against a fair baseline, not the old single-shot figure, which would
otherwise contradict the hero.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Grouped bars of LOC, tokens, cost and time as a % of the no-skill baseline
(lower is leaner/cheaper/faster), plus a separate safety strip (baseline,
caveman and ponytail 100%; yagni-oneliner 95%). System-gray palette so it reads
on both GitHub themes. The chart commits landed after #158 had already
squash-merged, so this brings the chart onto main.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Rebuild the benchmark to the standard #126 asked for: real headless Claude Code
sessions (not a bare model) editing a real public repo
(tiangolo/full-stack-fastapi-template @ cd83fc1, MIT), fair arms (baseline,
caveman, ponytail, and the "YAGNI + one-liners" prompt), n=4, Haiku 4.5. LOC is
the git diff; the safety tasks execute the produced code against adversarial
input.
Results: ponytail -54% LOC mean (up to -94% on over-build features like the
date/color picker), -22% tokens, -20% cost, -27% time, and never more than
baseline; 100% safe vs the one-liner prompt's 95% (it dropped a path-traversal
guard once). caveman writes less code but spends more tokens.
Also fixes a baseline-contamination bug (the ponytail plugin's SessionStart hook
fired on every arm; now isolated with --setting-sources project,local + per-arm
--plugin-dir) and a Windows subprocess-timeout hang.
Lead both READMEs with the agentic numbers; demote the single-shot 80-94% to a
labelled "isolated generation" note; supersede the contaminated 2026-06-17
writeup. Dead react-app fixture left untracked.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: correct cost claim to 42-75% from 30-rep re-verification
Re-ran the cost benchmark at 30 reps per cell on Claude (Haiku/Sonnet/Opus):
ponytail is 42-75% cheaper than no-skill, not the previously published 47-77%.
The direction holds, both ends came in a few points lower. Updates the README
headline and body, the benchmark chart subtitle, and the benchmarks/README cost
table, and adds a dated results doc with full method.
Also adds the OpenAI (gpt-4.1-mini/gpt-5.4-mini/gpt-5.5) and Gemini configs. On
OpenAI reasoning models ponytail costs more, not less, so the claim stays
Claude-scoped. Gemini run pending a fresh-quota day.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: scope the body claim to Claude models
"on every model" read as cross-provider, but the 30-rep verification shows
the cost win reverses on OpenAI reasoning models. Match the caption and
benchmarks/README, which already say Claude.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: reframe the pitch as the discipline, not token savings
The cost/code/latency numbers vary by model and on some (terse reasoning
models like GPT-5.5) ponytail costs more, so leading with them as a universal
win was misleading. Adds model-variance to the headline caption and a paragraph
making the stated point the mental model: write only what the task needs,
safety kept, maintainable code. Savings are a model-dependent side effect.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: name the ladder's reasoning cost
The ladder is a deliberation step: on reasoning models the agent spends
thinking tokens working through the rungs before it saves any output, which
together with the always-on ruleset can outweigh the shorter code. Makes the
GPT-5.5 cost increase legible rather than just stating it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: state the single-shot limitation honestly
The benchmark is single-shot (one prompt, one completion); it does not measure
a real multi-turn agent session, where the ruleset re-injects and the ladder
deliberates every turn. Adds that caveat to the README, and corrects the
benchmarks/README note that claimed caching widens the gap "in ponytail's
favor" (unverified, and a measured agentic A/B in #121 found the opposite can
happen). Per-session cost can land either way.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: fix run count in caption (cost is 30 runs, not 10)
Cost was re-verified at 30 reps; code and latency are still the original 10.
The headline caption said "10 runs" across the board, which undersold the cost
verification. Now states the split, matching benchmarks/README.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(examples): replace hand-written examples with real benchmark output
The examples/ before/after blocks were authored by hand, not produced by a
model. Issue #127 correctly noted that nobody hand-rolls quicksort for "sort
this array" - every model just calls .sort(). Regenerate all examples verbatim
from a real benchmark run (Claude Haiku 4.5, no-skill arm vs ponytail arm,
benchmarks/output.json) so the before/after is reproducible, not authored:
email 75->3, debounce 116->10, csv 20->3, countdown 267->9, rate-limit 128->10 LOC
- Delete sorting.md (pure strawman) plus the other hand-written caricatures
(api-endpoint, caching, date-picker)
- Add benchmarks/generate-examples.mjs to regenerate examples from any run
- examples/README.md indexes the set and documents how to reproduce
Closes#127
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: correct cost claim to 42-75% from 30-rep re-verification
Re-ran the cost benchmark at 30 reps per cell on Claude (Haiku/Sonnet/Opus):
ponytail is 42-75% cheaper than no-skill, not the previously published 47-77%.
The direction holds, both ends came in a few points lower. Updates the README
headline and body, the benchmark chart subtitle, and the benchmarks/README cost
table, and adds a dated results doc with full method.
Also adds the OpenAI (gpt-4.1-mini/gpt-5.4-mini/gpt-5.5) and Gemini configs. On
OpenAI reasoning models ponytail costs more, not less, so the claim stays
Claude-scoped. Gemini run pending a fresh-quota day.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: scope the body claim to Claude models
"on every model" read as cross-provider, but the 30-rep verification shows
the cost win reverses on OpenAI reasoning models. Match the caption and
benchmarks/README, which already say Claude.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: reframe the pitch as the discipline, not token savings
The cost/code/latency numbers vary by model and on some (terse reasoning
models like GPT-5.5) ponytail costs more, so leading with them as a universal
win was misleading. Adds model-variance to the headline caption and a paragraph
making the stated point the mental model: write only what the task needs,
safety kept, maintainable code. Savings are a model-dependent side effect.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: name the ladder's reasoning cost
The ladder is a deliberation step: on reasoning models the agent spends
thinking tokens working through the rungs before it saves any output, which
together with the always-on ruleset can outweigh the shorter code. Makes the
GPT-5.5 cost increase legible rather than just stating it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: state the single-shot limitation honestly
The benchmark is single-shot (one prompt, one completion); it does not measure
a real multi-turn agent session, where the ruleset re-injects and the ladder
deliberates every turn. Adds that caveat to the README, and corrects the
benchmarks/README note that claimed caching widens the gap "in ponytail's
favor" (unverified, and a measured agentic A/B in #121 found the opposite can
happen). Per-session cost can land either way.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: fix run count in caption (cost is 30 runs, not 10)
Cost was re-verified at 30 reps; code and latency are still the original 10.
The headline caption said "10 runs" across the board, which undersold the cost
verification. Now states the split, matching benchmarks/README.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Bumps all four plugin manifests to 4.7.0 for the OpenClaw / ClawHub skill package (#102).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds .openclaw/skills/ (ponytail + review/audit/debt/help) generated from the canonical skills/ (verbatim body, no drift), a generator script, and a drift test. Verified live: loads as Ready in OpenClaw 2026.6.6.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(benchmarks): correctness gate scores unfenced code; fix debounce task
The `correct` gate under-reported correctness for terse models, the likely
source of "Ponytail degrades models" reports (issue #65):
- extractBlocks() only matched fenced code blocks, so bare/unfenced code
scored an automatic fail even when correct. Now falls back to the whole
response as one block (and tolerates CRLF). Debounce detection also accepts
unfenced arrow functions.
- The debounce task asked to "add debounce to a search input" but the check
expected a reusable debounce(fn, delay) util, failing correct inline answers.
Task reworded to the deliverable the check verifies.
Adds correctness.test.js (regression guard) and a GPT-mini repro config plus
results writeup: on a clean n=20 run, the reported gpt-4.1-mini drop (10/15)
does not reproduce (100/100). The LOC win (~halved) holds.
README repro fixed: promptfoo needs --env-file ../.env (reads cwd, not root).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(benchmarks): add robustness audit — ponytail vs baseline on edge cases
Answers the real question behind #65: does ponytail's push for the shortest
solution make weak models produce wrong code on edge cases?
robustness-audit.js: 16 self-verifying tasks (12 algorithmic edge-case traps +
4 validators). Each check ships a known-good and known-lazy-wrong reference that
must pass/fail before any model output is scored (--selftest, 16/16).
Findings (gpt-4.1-mini + gpt-5.4-mini, baseline vs ponytail): parity on every
edge-case trap on both models. The one measured soft spot is gpt-5.4-mini email
(~4-5%, reaches for parseaddr). A sharpened SKILL.md validation rule had no
reliable effect in an n=100 A/B (96% vs 95%), so it was not shipped — the
tendency is model-level, not skill-level. Full writeup in results/.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(benchmarks): email slip is provider-specific — 100% on Claude
High-n cross-provider follow-up to the robustness audit. The one ponytail
soft spot (email validation via parseaddr) splits by provider, not model size:
- Claude (haiku/sonnet/opus): 100% under ponytail, n=40 each — and ponytail
beats baseline (unconstrained Sonnet over-engineers into an always-truthy
dict, 0/40; ponytail writes a clean validator).
- OpenAI (gpt-4.1-mini..gpt-5.5): slips at every size under ponytail
(~79-98%), baseline ~100%. The parseaddr reflex lives in OpenAI training.
Not fixable by skill text: 8 distinct SKILL.md edits (incl. an n=100 A/B,
96% vs 95%) all scored <= current, several worse, all bloated LOC. Nothing
shipped. SKILL.md unchanged.
Conclusion: on ponytail's target platform (Claude) email is 100%; the GPT
slip is a documented cross-provider transfer quirk. Adds model-email.js /
claude-email.js to reproduce the tables. Writeup updated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(benchmarks): correct misleading Sonnet baseline 0 percent
The Sonnet baseline 0/40 on email is a return-type artifact, not a logic
failure: unconstrained Sonnet returns a dict {is_valid, message} instead of a
bool, so the bool-contract gate scores every case as accepted. Read dict-aware
via is_valid, its logic is ~75% correct (9/12). Reframed honestly so we are not
presenting 0 vs 100 as a clean win; ponytail still wins (clean 100% bool) but
the point is over-engineered return type, not total failure.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Documents the Antigravity CLI install, the global default-level config, and the OpenCode absolute-path option. Addresses #58, #64, #71.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Bumps all four plugin manifests to 4.6.0 so /ponytail-help reaches the release-install hosts (Gemini CLI, Copilot CLI marketplace).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Fixes the local benchmark LOC counter (counted only fenced code, scored bare output 0), makes summary output ASCII-safe (a Unicode arrow crashed the script on Windows cp1252), gitignores generated artifacts, and refreshes the llama3.2 writeup with n=5 data showing the LOC effect is within the noise floor. Follow-up to #63. Verified live.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Release-prep bump across all four plugin manifests (Claude Code, Codex,
Gemini, Copilot) for v4.5.0. The cross-manifest parity test keeps them aligned.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude Code runs hooks via a non-interactive /bin/sh. On setups where node
isn't on that shell's PATH (Nix/nix-darwin, nvm, fnm), every prompt errored
with "/bin/sh: node: command not found". Guard each hook command so it runs
node only when present and exits 0 otherwise, no more per-prompt noise. The
slash-command skills are unaffected; only the always-on activation needs node.
Document the requirement in the README install section.
Closes#51.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The badge had fallen behind: it stayed at 11 when Antigravity and the VS Code
Codex extension were added, and Copilot CLI is now a full plugin host too. 13
distinct agent rows in docs/agent-portability.md (excluding the generic
fallback).
A `git add -A` during the v4.4.0 bump accidentally tracked the announcement
art (announce/changelog/ponytail-*.gif) in the repo root. Untrack them and
gitignore the pattern; they stay on disk for posting but out of the repo.
The header showed logo.png (black on transparent), which nearly vanishes in
GitHub's dark theme. Wrap it in a <picture> so dark-theme viewers get the
contoured logo-dark.png and light-theme viewers keep the original.
A dark-bg-ready variant of the mark: white face fill plus a die-cut white
contour so it reads on dark backgrounds, where logo.png (black on transparent)
and the social-preview face do not. Ships as SVG (scalable, white + black
layers) and a 1085x1241 PNG.
Contributed by @pixexid in #42.
Co-authored-by: pixexid <54691335+pixexid@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: add a Commands reference to the README
The commands were only mentioned scattered through prose, and two install
blurbs had gone stale (OpenCode omitted /ponytail-debt, Gemini omitted audit
and debt). Add one canonical Commands table (all five commands + what each
does + which hosts support them) and point the install blurbs at it so they
stop drifting as commands are added.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs: add pi to the portability table, de-stale the Gemini row
pi was a supported integration (pi-extension, README install, registers all
the commands) but had no row in the Supported Adapters table. The Gemini row
also enumerated an outdated command list; point it at commands/*.toml
generically so it stops drifting.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Closes the last gap from the field review: deferral creep. /ponytail-debt
greps the repo for `ponytail:` comment markers and prints a ledger
(file:line, what was simplified, ceiling, upgrade trigger), flagging any
marker with no trigger as the rot risk. One-shot, reports only.
Full parity like ponytail-audit: skill + commands/.toml + .opencode/.md + pi
registerCommand (+ test) + agent-portability + README.
Verified: tests 32/32 (pi command list updated), rule check green, the scan
finds the repo's real markers, and a live end-to-end run produced a correct
ledger (2 markers, 1 no-trigger, prose/examples excluded).
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat: refine ruleset from a full-project field review
A reviewer ran ponytail across a 9-phase rewrite (protocol, PC app, simulator,
RPi daemon, ESP32 firmware) and flagged three gaps. All three land in SKILL.md
and propagate to AGENTS.md + the rule copies:
- Promote the one-runnable-check rule to a headline ("Lazy code without its
check is unfinished"), enforced as a check-rule-copies invariant.
- Hardware carve-out in "When NOT to be lazy": a real device is never the spec
ideal (clock drift, sensor offset), leave the calibration knob.
- Clarify the Output rule: explanation the user explicitly asked for is not
debt, only unrequested prose is.
Fallback instructions kept in sync. Rule-copy check + tests green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: add a behavior gate proving the refinements actually fire
The refinements were verified as injected text, but injected != behavioral.
This adds a behavior eval that probes each refined rule on a task that should
trigger it:
- hardware -> does the output leave a calibration knob?
- explanation -> when a write-up is explicitly requested, is it given in full?
- onecheck -> is a runnable check left behind?
benchmarks/behavior.yaml runs the probes (baseline vs ponytail arm); the
grader benchmarks/behavior.js is proven by tests/behavior.test.js (8 cases,
RED/GREEN, no API key, runs in CI). Live-confirmed: the model under the
current ruleset passes all three gates, graded by the same grader.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Release-prep bump across the Claude Code, Codex, and Gemini manifests.
Cutting v4.3.0 also fixes#33: gemini extensions install pulls the latest
GitHub release, and gemini-extension.json was added after v4.2.0, so it
was missing from the release tarball. Shipping it in a release fixes the
plain install command.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ponytail-activate.js and ponytail-runtime.js hardcoded ~/.claude for the
flag file and settings lookup, ignoring CLAUDE_CONFIG_DIR. Add a shared
getClaudeDir() to ponytail-config.js and use it in both. Regression test
added to hooks.test.js.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Both read AGENTS.md, which the repo already ships, so ponytail works
from the repo root with no extra setup. Add agent-portability rows and a
README note. Instruction-tier (no /ponytail levels or hooks).
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Copilot CLI already reads AGENTS.md and .github/copilot-instructions.md
(both shipped), plus a global ~/.copilot/copilot-instructions.md. Add an
agent-portability row and a README note. Instruction-tier only: no
/ponytail levels or hooks.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Use auto-update or marketplace refresh + /reload-plugins (not reinstall), and
note that an unrecognized /plugin means Claude Code itself needs updating.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Swap em dashes for commas/colons/periods in the README, skills, AGENTS.md and
its five rule copies, examples, command files, and benchmark README. Rule
copies stay in sync (same edit applied to all) and the invariant guard passes.
Left untouched on purpose: the vendored caveman SKILL.md (verbatim third-party
text), the dated benchmark writeups in results/ (historical records), and
.js code comments.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The ponytail API example returned the raw ORM model, which exposes every
column. Restore a response_model whitelist (the trust boundary) while still
cutting the repository/service/exception ceremony. Matches the skill's own
'never cut security' rule.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Commit a promptfoo harness (config + arm prompts + LOC metric + vendored
caveman SKILL) so anyone can re-run the comparison: no-skill vs caveman vs
ponytail, across Haiku / Sonnet / Opus, 10 runs per cell, median reported.
Replace the old unreproducible 6-task chart with assets/benchmark-3model.svg
from this run, and reframe the README to the reproducible numbers: ponytail
writes 80-94% less code, costs 47-77% less, and runs 3-6x faster than a
no-skill agent on every model. benchmarks/README.md carries the median tables
and the reproduce command. Drops nothing that is not measured.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Skill preaching minimalism was 2x caveman length. Smaller file cuts
per-read and per-session-injection cost. Benchmark: beats caveman on
all areas now — 135.7k vs 138.4k tokens, 127s vs 136s, 47 vs 117 loc.
v1 lost to caveman on tokens/time despite minimal code: it wrote
essays defending each simplification. v2 caps explanation at three
lines and ships the lazy version instead of stalling on necessity
questions. Benchmark: 136.6k tok vs caveman 138.4k, code 47 vs 117
lines across 5 tasks.