Commit Graph
150 Commits
Author SHA1 Message Date
DietrichGebert 48cdf05a25 Revert "benchmarks: add system prompt to baseline arm so it doesn't ramble (closes #126) (#128)" (#175)
This reverts commit 37f46b8f02.
2026-06-19 01:44:13 +02:00
DietrichGebertandClaude Opus 4.8 a4e5e479d6 docs: re-sync README.es.md to current English (agentic numbers, CodeWhale) (#174)
The Spanish README merged (#110) carrying stale content: the old flat
"80-94% menos código" single-shot headline (the exact claim #126 corrected),
no CodeWhale section, badge stuck at 13 agents, and missing the Modern Web
Guidance callout (#82) and the Claude Code desktop-install paragraph.

Re-translates the hero + Números section to the corrected agentic numbers
(~54%, up to 94%, 100% safe) with the old figures demoted to the same
<details> block English uses, adds CodeWhale, fixes the badge, and adds a
"community translation, English is the reference" note. Also adds a minimal
Español discoverability link to the English README so readers can find it.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 00:44:08 +02:00
Liad YosefandClaude fe963cae99 Add Modern Web Guidance as the rung-3 lookup for web tasks (#82)
Scoped, optional reference to Modern Web Guidance so the agent can look
up native platform features on web work, filling the gap at rung 3 of
the ladder. Three additive changes, no compact-ruleset surgery:

- skills/ponytail/SKILL.md: "Web tasks: rung 3 lookup" section after the
  ladder. Runtime source only, not byte-compared, so no six-file sync.
- README.md: one "Pairs well with" line, matching the Caveman pattern.
- examples/web-platform-lookup.md: a <dialog closedby> vs Radix
  before/after in the date-picker.md style.

Lookup, not license: MWG suggests, the ladder filters. Absent CLI
changes nothing. No new INVARIANT phrase; rule-copy check stays green.

Co-authored-by: Claude <noreply@anthropic.com>
2026-06-19 00:26:36 +02:00
Arthur MorrillandClaude Opus 4.8 b0c5820bb1 Fix scope ambiguity in ponytail-audit and ponytail-review Boundaries (#163)
The Boundaries line opened with "Complexity only, correctness bugs, security
holes, and performance go to a normal review pass." The comma after "Complexity
only" fuses the in-scope item with the out-of-scope list, so a model parsing it
literally can read all four categories as targets of the audit — the opposite of
intent.

Restate the boundary as an explicit scope fence: name what is in scope, then
mark correctness/security/performance as explicitly out of scope. "Out of scope"
is phrasing models reliably honor as a constraint. Also aligns the scope term
with each skill's stated purpose (over-engineering).

Applied to both skills/ and the .openclaw/ mirror so the two trees stay in sync.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 00:26:33 +02:00
Lakshya77089 83f67a261a feat: add Antigravity .agents/rules/ponytail.md rule copy (#119)
Antigravity IDE loads always-on rules from .agents/rules/. The README
already tells Antigravity users to drop the ruleset there, but the file
was missing. Add it as a verbatim copy of the canonical AGENTS.md body
(no frontmatter, like the .windsurf/.clinerules copies) and register it
in check-rule-copies.js so it is validated against AGENTS.md and cannot
drift. Closes #116.
2026-06-19 00:24:49 +02:00
Tianen Cheng a28e5ec123 fix: register skills directory via config hook so opencode discovers ponytail skills (#138)
The ponytail plugin was missing a config hook to register its skills/
directory with opencode's skill discovery system. Without this, the 5
ponytail skills (ponytail, ponytail-audit, ponytail-debt, ponytail-help,
ponytail-review) never appear in the skill tool's available list.

The superpowers plugin already follows this exact pattern — this brings
ponytail in line with the upstream convention.
2026-06-19 00:24:46 +02:00
Sonai Biswas 4dad14fac5 Avoid Gemini loading Claude hook events (#139) 2026-06-19 00:24:43 +02:00
Uchenna 0ac987f995 feat(mcp): add ponytail-mcp, an MCP server for the ruleset (#91)
* feat(mcp): add ponytail-mcp server (prompt + tool)

* test(mcp): cover mode resolution and instruction text

* Report resolved MCP mode
2026-06-19 00:24:40 +02:00
Ben YounesandClaude Opus 4.8 15749f7ffc feat(skills): add /ponytail-gain measured-impact scoreboard (#108)
A one-shot scoreboard showing ponytail's measured benchmark impact
(less code, less cost, more speed) as plain ASCII bars, then points to
/ponytail-debt and /ponytail-audit for this repo's real numbers.

Complements the existing skills rather than duplicating them: debt
harvests the ponytail: ledger, audit finds what's cuttable, gain shows
the measured why-it-matters. No per-repo savings number is ever printed
-- the unbuilt version was never written, so there is no real baseline
to subtract from in a live repo. The bars carry the published benchmark
medians (5 tasks, 3 models); per-repo figures come from debt's count.

Ships every adapter the other commands ship: Claude commands/*.toml,
OpenCode .opencode/command/*.md, OpenClaw skill (generated), Pi command
registration. Help card, command enumeration, portability table, and
README updated in the same change.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 00:24:37 +02:00
Sha256_NulledandFato07 37f46b8f02 benchmarks: add system prompt to baseline arm so it doesn't ramble (closes #126) (#128)
Co-authored-by: Fato07 <fato07@users.noreply.github.com>
2026-06-19 00:24:35 +02:00
Ben YounesandClaude Opus 4.8 e782790b15 feat(benchmarks): add critic-email task — reproduces the critique's own example (#173)
The Scott Logic post ("Ponytail? YAGNI!", see #126) argued a bare one-liner
prompt matches ponytail because both shrink the line count. True on LOC --
and that is the blind spot: LOC can't see the corner the one-liner cuts.

The canonical lazy email validator uses re.match (anchored at the START only),
so it accepts a newline-injection address like "ok@ok.com\n<payload>" -- a real
header/log-injection vector. ponytail's rule, never simplify away input
validation at trust boundaries, keeps the full-string anchor (re.fullmatch).
Same shortness, one keeps the guard.

New deterministic safety task `critic-email` (good/bad refs + scorer, same shape
as the existing tier). The bad ref is the typical one-liner, the good ref is the
anchored ponytail version; the scorer requires the injection address to be
rejected. Verifiable with no API key via `run.py --selftest`.

Refs #126

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 00:24:32 +02:00
Jesus cornelio b345e49385 examples: add 6 new over-engineering survivors + platform-native guide (#109)
New examples (examples/):
- modal-dialog: <dialog> vs Radix/react-modal
- url-params: URLSearchParams vs query-string
- number-formatting: Intl.NumberFormat vs numeral
- infinite-scroll: IntersectionObserver vs react-infinite-scroll-component
- deep-clone: structuredClone vs lodash.cloneDeep / JSON hack
- group-by: Object.groupBy vs lodash.groupBy

New doc (docs/platform-native.md):
Comprehensive reference of platform-native solutions across HTML elements,
CSS, Browser APIs, Node.js stdlib, Python stdlib, and database features.
Covers 60+ cases where the platform already has what developers reach for
a package to do.
2026-06-19 00:24:29 +02:00
Jesus cornelio 7f4dc907fc docs: add Spanish (LATAM) translation of README (#110) 2026-06-19 00:24:26 +02:00
Ben YounesandClaude Opus 4.8 e7e09f8fd4 docs: add CodeWhale support (AGENTS.md native reader, zero setup) (#124)
CodeWhale reads AGENTS.md from project root per its CONFIGURATION.md —
ponytail already works with no adapter file needed. Added dedicated install
section, agent count bump (13→14), and portability table row.

Also adds Zed to the grouped instruction-only adapter list (same
mechanism: reads AGENTS.md natively).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-19 00:24:23 +02:00
Aiden 766c5ca5b1 feat: add argument-hint to ponytail skill (#85) 2026-06-19 00:24:20 +02:00
hashner 25be875fab Fix markdown formatting in ponytail-debt.md (#142)
Fixed "mapping values not allowed in this context error", by adding quotes around the description so parser does not detect a key/value mapping
2026-06-19 00:24:17 +02:00
Colin Eberhardt 70df716a02 Fix path for .env file in README (#104)
The reproduction steps involve running promptfoo from the benchmarks folder. In order for the environment var9ables in `.env` to be discoverable they need to be in this folder, not the project root.
2026-06-19 00:24:14 +02:00
DietrichGebertandClaude Opus 4.8 6d35c10920 docs: desktop install + OpenCode command linking (#105)
Documents the desktop-app install flow (no /plugin command) and the global command-dir linking needed for /ponytail commands in OpenCode outside a checkout. Covers the recurring questions in #97 and #98.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 00:24:11 +02:00
Lakshya77089 795ec0ee36 fix: statusline reads flag from CLAUDE_CONFIG_DIR, not just ~/.claude (#34 follow-up) (#154)
Issue #34 made the hooks honor CLAUDE_CONFIG_DIR when writing the mode flag
($CLAUDE_CONFIG_DIR/.ponytail-active), enforced by tests/hooks.test.js. But
both statusline scripts still hardcoded $HOME/.claude/.ponytail-active, so any
user with CLAUDE_CONFIG_DIR set gets no badge — or a stale mode from a
pre-migration ~/.claude flag that never updates again.

Make both scripts resolve the flag the same way getClaudeDir() does: prefer
CLAUDE_CONFIG_DIR, fall back to ~/.claude. The fallback branch is identical to
the previous behavior, so unset-env users are unaffected. Also corrects the
now-inaccurate path comment in the activation hook header.
2026-06-18 22:50:17 +02:00
Lakshya77089 c30854118e fix: guard final writeHookOutput against stdout EPIPE in ponytail-activate (#149) (#152)
The final writeHookOutput('SessionStart', ...) call was the only operation
in the file outside a try/catch. writeHookOutput ends in a bare
process.stdout.write, so a closed stdout / broken pipe (EPIPE) at hook exit
throws uncaught and crashes the hook with a non-zero exit code. Wrap it to
match the file's existing never-block-session-start posture.
2026-06-18 22:50:13 +02:00
Lakshya77089 a3bc7db722 fix: strip UTF-8 BOM before parsing settings.json in ponytail-activate (#148) (#151)
settings.json written by Notepad or VS Code on Windows can carry a
UTF-8 BOM. JSON.parse then throws SyntaxError, the outer catch swallows
it, hasStatusline stays false, and the statusline setup nudge is never
emitted.

Strip the leading BOM before parsing, matching the existing handling in
ponytail-mode-tracker.js. (#96 added a null guard but not BOM stripping.)
2026-06-18 22:50:10 +02:00
Manvi 55b7cb1925 fix: resolve test failure on Node.js < 20.11.0 by using new URL (#157) 2026-06-18 22:50:06 +02:00
Lakshya77089 53fd1e850e fix: only deactivate on a standalone "stop ponytail" / "normal mode" (#162)
The deactivation check matched the phrase anywhere in the prompt, so an
ordinary request like "add a normal mode toggle" silently turned ponytail
off for the rest of the session. Match the whole message instead (trimmed,
case-insensitive, trailing punctuation ignored) through a shared helper used
by both the Claude/Codex hook and the pi extension.

Fixes #161
2026-06-18 22:50:03 +02:00
Ben YounesandClaude Opus 4.8 955fff537c feat(benchmarks): add completeness judge so LOC wins can't hide under-delivery (#171)
The LOC tier scores the open feature tasks (vibe-*, tmpl-fe-*, open-*) on
git diff alone -- score_vibe only checks "it compiles", score_fixture only
checks "a new file exists". So an arm can win the LOC metric by shipping a
stub: fewer lines because it does less, not because it is less bloated.
That is the most credible attack left on the headline number raised in #126.

complete.py is a second LLM judge (same auditable footing as judge.py: fixed
model, temperature 0, published rubric) that rates how FULLY each submission
implements its task, 0..3. Read alongside the LOC table, a low-LOC arm whose
completeness also drops is caught, not rewarded.

- judge_call gains a `system=` param so the HTTP/key/source plumbing is reused
  instead of duplicated (one rubric is the only delta between the two passes).
- --selftest: the judge must rank a complete reference strictly above a stub.
- --selftest-offline: validates the gate logic with no API call / no key.
- README documents the pass and updates the can/cannot-show limitations.

Fixes #126

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 22:40:09 +02:00
Ben YounesandClaude Opus 4.8 44babb22ef fix(benchmarks): portable plugin-dir resolution for agentic arms (#170)
The ponytail/caveman arms hardcoded one machine's Windows plugin-cache
paths (C:\Users\Dietr\...), so only baseline/yagni/yagni-oneliner were
reproducible off the maintainer's box — undercutting the "fully
reproducible" claim the rebuilt benchmark (#126) was meant to establish.

Resolve per-arm at use-site: env override (PONYTAIL_PLUGIN_DIR /
CAVEMAN_PLUGIN_DIR) -> latest version dir under ~/.claude/plugins/cache
-> clear sys.exit. No pinned version/hash. Selftest extended to cover
env-override and missing-install paths.

Fixes #169

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 22:39:59 +02:00
DietrichGebertandClaude Opus 4.8 91aef5dfef docs(readme): note the up-to-94% peak in the hero line (#165)
The hero showed only ~54% (the mean); the rigorous agentic run also reaches 94%
on the over-build tasks (the date picker), so the headline now reads
"~54% (up to 94%)". The sub-line is reworded so 80-94% reads as the per-task
ceiling against a fair baseline, not the old single-shot figure, which would
otherwise contradict the hero.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 19:25:18 +02:00
DietrichGebertandClaude Opus 4.8 8d5037d9e5 docs(readme): add the agentic benchmark chart (#160)
Grouped bars of LOC, tokens, cost and time as a % of the no-skill baseline
(lower is leaner/cheaper/faster), plus a separate safety strip (baseline,
caveman and ponytail 100%; yagni-oneliner 95%). System-gray palette so it reads
on both GitHub themes. The chart commits landed after #158 had already
squash-merged, so this brings the chart onto main.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 17:09:29 +02:00
DietrichGebertandClaude Opus 4.8 b8d6aa7e9f feat(benchmarks): agentic LOC + safety benchmark answering #126 (#158)
Rebuild the benchmark to the standard #126 asked for: real headless Claude Code
sessions (not a bare model) editing a real public repo
(tiangolo/full-stack-fastapi-template @ cd83fc1, MIT), fair arms (baseline,
caveman, ponytail, and the "YAGNI + one-liners" prompt), n=4, Haiku 4.5. LOC is
the git diff; the safety tasks execute the produced code against adversarial
input.

Results: ponytail -54% LOC mean (up to -94% on over-build features like the
date/color picker), -22% tokens, -20% cost, -27% time, and never more than
baseline; 100% safe vs the one-liner prompt's 95% (it dropped a path-traversal
guard once). caveman writes less code but spends more tokens.

Also fixes a baseline-contamination bug (the ponytail plugin's SessionStart hook
fired on every arm; now isolated with --setting-sources project,local + per-arm
--plugin-dir) and a Windows subprocess-timeout hang.

Lead both READMEs with the agentic numbers; demote the single-shot 80-94% to a
labelled "isolated generation" note; supersede the contaminated 2026-06-17
writeup. Dead react-app fixture left untracked.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 16:42:38 +02:00
salaamdev b4f725f7ce Merge branch 'main' into main 2026-06-18 10:45:26 +03:00
DietrichGebertandClaude Opus 4.8 45f7d2f83f Fix/examples issue 127 (#131)
* docs: correct cost claim to 42-75% from 30-rep re-verification

Re-ran the cost benchmark at 30 reps per cell on Claude (Haiku/Sonnet/Opus):
ponytail is 42-75% cheaper than no-skill, not the previously published 47-77%.
The direction holds, both ends came in a few points lower. Updates the README
headline and body, the benchmark chart subtitle, and the benchmarks/README cost
table, and adds a dated results doc with full method.

Also adds the OpenAI (gpt-4.1-mini/gpt-5.4-mini/gpt-5.5) and Gemini configs. On
OpenAI reasoning models ponytail costs more, not less, so the claim stays
Claude-scoped. Gemini run pending a fresh-quota day.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: scope the body claim to Claude models

"on every model" read as cross-provider, but the 30-rep verification shows
the cost win reverses on OpenAI reasoning models. Match the caption and
benchmarks/README, which already say Claude.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: reframe the pitch as the discipline, not token savings

The cost/code/latency numbers vary by model and on some (terse reasoning
models like GPT-5.5) ponytail costs more, so leading with them as a universal
win was misleading. Adds model-variance to the headline caption and a paragraph
making the stated point the mental model: write only what the task needs,
safety kept, maintainable code. Savings are a model-dependent side effect.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: name the ladder's reasoning cost

The ladder is a deliberation step: on reasoning models the agent spends
thinking tokens working through the rungs before it saves any output, which
together with the always-on ruleset can outweigh the shorter code. Makes the
GPT-5.5 cost increase legible rather than just stating it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: state the single-shot limitation honestly

The benchmark is single-shot (one prompt, one completion); it does not measure
a real multi-turn agent session, where the ruleset re-injects and the ladder
deliberates every turn. Adds that caveat to the README, and corrects the
benchmarks/README note that claimed caching widens the gap "in ponytail's
favor" (unverified, and a measured agentic A/B in #121 found the opposite can
happen). Per-session cost can land either way.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: fix run count in caption (cost is 30 runs, not 10)

Cost was re-verified at 30 reps; code and latency are still the original 10.
The headline caption said "10 runs" across the board, which undersold the cost
verification. Now states the split, matching benchmarks/README.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(examples): replace hand-written examples with real benchmark output

The examples/ before/after blocks were authored by hand, not produced by a
model. Issue #127 correctly noted that nobody hand-rolls quicksort for "sort
this array" - every model just calls .sort(). Regenerate all examples verbatim
from a real benchmark run (Claude Haiku 4.5, no-skill arm vs ponytail arm,
benchmarks/output.json) so the before/after is reproducible, not authored:

  email 75->3, debounce 116->10, csv 20->3, countdown 267->9, rate-limit 128->10 LOC

- Delete sorting.md (pure strawman) plus the other hand-written caricatures
  (api-endpoint, caching, date-picker)
- Add benchmarks/generate-examples.mjs to regenerate examples from any run
- examples/README.md indexes the set and documents how to reproduce

Closes #127

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 04:35:22 +02:00
DietrichGebertandClaude Opus 4.8 1b4914159e docs: correct cost claim to 42-75% from 30-rep re-verification (#129)
* docs: correct cost claim to 42-75% from 30-rep re-verification

Re-ran the cost benchmark at 30 reps per cell on Claude (Haiku/Sonnet/Opus):
ponytail is 42-75% cheaper than no-skill, not the previously published 47-77%.
The direction holds, both ends came in a few points lower. Updates the README
headline and body, the benchmark chart subtitle, and the benchmarks/README cost
table, and adds a dated results doc with full method.

Also adds the OpenAI (gpt-4.1-mini/gpt-5.4-mini/gpt-5.5) and Gemini configs. On
OpenAI reasoning models ponytail costs more, not less, so the claim stays
Claude-scoped. Gemini run pending a fresh-quota day.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: scope the body claim to Claude models

"on every model" read as cross-provider, but the 30-rep verification shows
the cost win reverses on OpenAI reasoning models. Match the caption and
benchmarks/README, which already say Claude.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: reframe the pitch as the discipline, not token savings

The cost/code/latency numbers vary by model and on some (terse reasoning
models like GPT-5.5) ponytail costs more, so leading with them as a universal
win was misleading. Adds model-variance to the headline caption and a paragraph
making the stated point the mental model: write only what the task needs,
safety kept, maintainable code. Savings are a model-dependent side effect.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: name the ladder's reasoning cost

The ladder is a deliberation step: on reasoning models the agent spends
thinking tokens working through the rungs before it saves any output, which
together with the always-on ruleset can outweigh the shorter code. Makes the
GPT-5.5 cost increase legible rather than just stating it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: state the single-shot limitation honestly

The benchmark is single-shot (one prompt, one completion); it does not measure
a real multi-turn agent session, where the ruleset re-injects and the ladder
deliberates every turn. Adds that caveat to the README, and corrects the
benchmarks/README note that claimed caching widens the gap "in ponytail's
favor" (unverified, and a measured agentic A/B in #121 found the opposite can
happen). Per-session cost can land either way.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: fix run count in caption (cost is 30 runs, not 10)

Cost was re-verified at 30 reps; code and latency are still the original 10.
The headline caption said "10 runs" across the board, which undersold the cost
verification. Now states the split, matching benchmarks/README.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 04:23:51 +02:00
Joseph Huang 99139a25d0 fix: pin all four safety carve-outs in the rule-drift canary (#114)
INVARIANTS pinned only 'input validation at trust boundaries'; the other three carve-outs (data-loss, security, accessibility) could drift silently. Adds 'prevents data loss', 'security', 'accessibility' as canaries, each present verbatim in both SKILL.md and AGENTS.md. No rule text changed.
2026-06-16 18:12:22 +02:00
Lakshya77089 c596a2d640 docs: clarify benchmark numbers are per-task, not plan quota (#115)
Adds a caveat under the headline numbers: they are per-task code/latency/cost on the Claude API, not a plan quota promise. Prevents the misread in #111.
2026-06-16 18:03:52 +02:00
DietrichGebertandClaude Opus 4.8 adad50d9b3 chore: bump version to 4.7.0 (#103)
Bumps all four plugin manifests to 4.7.0 for the OpenClaw / ClawHub skill package (#102).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
v4.7.0
2026-06-16 13:56:32 +02:00
DietrichGebertandClaude Opus 4.8 41d6c2f761 feat: ship ponytail to OpenClaw (ClawHub skill package) (#102)
Adds .openclaw/skills/ (ponytail + review/audit/debt/help) generated from the canonical skills/ (verbatim body, no drift), a generator script, and a drift test. Verified live: loads as Ready in OpenClaw 2026.6.6.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 13:35:35 +02:00
DietrichGebertandClaude Opus 4.8 caf138df56 benchmarks: fix correctness gate + robustness audit (#65) (#83)
* fix(benchmarks): correctness gate scores unfenced code; fix debounce task

The `correct` gate under-reported correctness for terse models, the likely
source of "Ponytail degrades models" reports (issue #65):

- extractBlocks() only matched fenced code blocks, so bare/unfenced code
  scored an automatic fail even when correct. Now falls back to the whole
  response as one block (and tolerates CRLF). Debounce detection also accepts
  unfenced arrow functions.
- The debounce task asked to "add debounce to a search input" but the check
  expected a reusable debounce(fn, delay) util, failing correct inline answers.
  Task reworded to the deliverable the check verifies.

Adds correctness.test.js (regression guard) and a GPT-mini repro config plus
results writeup: on a clean n=20 run, the reported gpt-4.1-mini drop (10/15)
does not reproduce (100/100). The LOC win (~halved) holds.

README repro fixed: promptfoo needs --env-file ../.env (reads cwd, not root).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(benchmarks): add robustness audit — ponytail vs baseline on edge cases

Answers the real question behind #65: does ponytail's push for the shortest
solution make weak models produce wrong code on edge cases?

robustness-audit.js: 16 self-verifying tasks (12 algorithmic edge-case traps +
4 validators). Each check ships a known-good and known-lazy-wrong reference that
must pass/fail before any model output is scored (--selftest, 16/16).

Findings (gpt-4.1-mini + gpt-5.4-mini, baseline vs ponytail): parity on every
edge-case trap on both models. The one measured soft spot is gpt-5.4-mini email
(~4-5%, reaches for parseaddr). A sharpened SKILL.md validation rule had no
reliable effect in an n=100 A/B (96% vs 95%), so it was not shipped — the
tendency is model-level, not skill-level. Full writeup in results/.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(benchmarks): email slip is provider-specific — 100% on Claude

High-n cross-provider follow-up to the robustness audit. The one ponytail
soft spot (email validation via parseaddr) splits by provider, not model size:

- Claude (haiku/sonnet/opus): 100% under ponytail, n=40 each — and ponytail
  beats baseline (unconstrained Sonnet over-engineers into an always-truthy
  dict, 0/40; ponytail writes a clean validator).
- OpenAI (gpt-4.1-mini..gpt-5.5): slips at every size under ponytail
  (~79-98%), baseline ~100%. The parseaddr reflex lives in OpenAI training.

Not fixable by skill text: 8 distinct SKILL.md edits (incl. an n=100 A/B,
96% vs 95%) all scored <= current, several worse, all bloated LOC. Nothing
shipped. SKILL.md unchanged.

Conclusion: on ponytail's target platform (Claude) email is 100%; the GPT
slip is a documented cross-provider transfer quirk. Adds model-email.js /
claude-email.js to reproduce the tables. Writeup updated.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(benchmarks): correct misleading Sonnet baseline 0 percent

The Sonnet baseline 0/40 on email is a return-type artifact, not a logic
failure: unconstrained Sonnet returns a dict {is_valid, message} instead of a
bool, so the bool-contract gate scores every case as accepted. Read dict-aware
via is_valid, its logic is ~75% correct (9/12). Reframed honestly so we are not
presenting 0 vs 100 as a clean win; ponytail still wins (clean 100% bool) but
the point is over-engineered return type, not total failure.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 12:17:10 +02:00
salaamdev ce55fd460a Update author in plugin.yaml to Salaamdev 2026-06-15 23:13:09 +03:00
Abdisalam Hassan 4198fc30ac add Hermes plugin for Ponytail with command support and documentation 2026-06-15 19:57:45 +00:00
DietrichGebertandClaude Opus 4.8 687c1b3398 docs: Antigravity CLI install, global default level, OpenCode absolute path (#73)
Documents the Antigravity CLI install, the global default-level config, and the OpenCode absolute-path option. Addresses #58, #64, #71.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 17:04:56 +02:00
DietrichGebertandClaude Opus 4.8 ce153bc95f chore: bump version to 4.6.0 (#72)
Bumps all four plugin manifests to 4.6.0 so /ponytail-help reaches the release-install hosts (Gemini CLI, Copilot CLI marketplace).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
v4.6.0
2026-06-15 16:37:06 +02:00
YIZIHN 084f10fb48 fix: ship the missing /ponytail-help command on Claude Code and OpenCode (#62)
Ships the previously-missing /ponytail-help adapter files (commands/ponytail-help.toml, .opencode/command/ponytail-help.md) and adds tests/commands.test.js, a parity guard asserting every pi-registered command has both adapter files. Thanks @hooni0918.
2026-06-15 16:32:02 +02:00
DietrichGebertandClaude Opus 4.8 2e6a93765a fix(benchmarks): count unfenced code, ASCII-safe output, refresh llama3.2 results (#67)
Fixes the local benchmark LOC counter (counted only fenced code, scored bare output 0), makes summary output ASCII-safe (a Unicode arrow crashed the script on Windows cp1252), gitignores generated artifacts, and refreshes the llama3.2 writeup with n=5 data showing the LOC effect is within the noise floor. Follow-up to #63. Verified live.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 16:22:32 +02:00
Mandavilli Vijay 386f95734a benchmarks: add local model support and Node version note (#63)
Adds benchmarks/benchmark-local.py (Ollama-based local runner), a results writeup, and a Node version note. Thanks @mandavillivijay.
2026-06-15 15:27:11 +02:00
DietrichGebertandClaude Opus 4.8 60a75f8159 chore: bump version to 4.5.0 (#60)
Release-prep bump across all four plugin manifests (Claude Code, Codex,
Gemini, Copilot) for v4.5.0. The cross-manifest parity test keeps them aligned.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
v4.5.0
2026-06-15 11:49:37 +02:00
DietrichGebert b41cb8d3af docs: note the Codex install also covers the desktop app (#59) 2026-06-15 11:25:35 +02:00
DietrichGebertandClaude Opus 4.8 d676635325 fix: hooks degrade gracefully when node is not on PATH (#57)
Claude Code runs hooks via a non-interactive /bin/sh. On setups where node
isn't on that shell's PATH (Nix/nix-darwin, nvm, fnm), every prompt errored
with "/bin/sh: node: command not found". Guard each hook command so it runs
node only when present and exits 0 otherwise, no more per-prompt noise. The
slash-command skills are unaffected; only the always-on activation needs node.
Document the requirement in the README install section.

Closes #51.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 11:23:59 +02:00
Christopher MayfieldandCursor f02f9424a5 fix: use python3 for correctness checks and add CI (#50)
The benchmark harness hardcoded `python`, which is missing on macOS and
many Linux images. Probe python3 first, add npm test, and run checks in
GitHub Actions so regressions are caught on every PR.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-15 11:11:04 +02:00
DietrichGebert 2302fbc843 docs: bump the agents badge to 13 (#56)
The badge had fallen behind: it stayed at 11 when Antigravity and the VS Code
Codex extension were added, and Copilot CLI is now a full plugin host too. 13
distinct agent rows in docs/agent-portability.md (excluding the generic
fallback).
2026-06-15 11:07:04 +02:00
c1c80f3cc8 Adding support for Copilot Marketplace plugin (#47)
* Add GitHub Copilot plugin and marketplace manifests for Ponytail

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Add Copilot hook adapters and plugin data runtime precedence

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Document Copilot plugin install flow and instruction fallback mode

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Fix Copilot hooks for native output context and state-only mode tracking

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* docs: add Copilot CLI namespaced command examples

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Collapse Copilot hooks into shared activate/mode-tracker

The Copilot hook files duplicated ponytail-activate.js and
ponytail-mode-tracker.js, differing only in output shape. Move that
difference into writeHookOutput (isCopilot branch) and point
copilot-hooks.json at the shared hooks. Deletes both forks (-73 lines).

Refs #1

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Align Copilot manifest version to 4.4.0 with cross-manifest parity test

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Make Copilot and Codex host detection exclusive in runtime output routing

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Add Copilot debt command validation with a pull request acceptance checklist

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Removed PR template

* Drop tautological copilot command-form test

The namespaced-form assertion built '/ponytail:ponytail-debt' from two
constants and compared it to itself — it tests string concatenation, not
wiring. The file-exists check above already catches a renamed manifest.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 11:02:59 +02:00
DietrichGebert 1c420ad2f3 chore: untrack one-off social images committed by mistake (#46)
A `git add -A` during the v4.4.0 bump accidentally tracked the announcement
art (announce/changelog/ponytail-*.gif) in the repo root. Untrack them and
gitignore the pattern; they stay on disk for posting but out of the repo.
v4.4.0
2026-06-15 03:03:33 +02:00