* fix(benchmarks): correctness gate scores unfenced code; fix debounce task The `correct` gate under-reported correctness for terse models, the likely source of "Ponytail degrades models" reports (issue #65): - extractBlocks() only matched fenced code blocks, so bare/unfenced code scored an automatic fail even when correct. Now falls back to the whole response as one block (and tolerates CRLF). Debounce detection also accepts unfenced arrow functions. - The debounce task asked to "add debounce to a search input" but the check expected a reusable debounce(fn, delay) util, failing correct inline answers. Task reworded to the deliverable the check verifies. Adds correctness.test.js (regression guard) and a GPT-mini repro config plus results writeup: on a clean n=20 run, the reported gpt-4.1-mini drop (10/15) does not reproduce (100/100). The LOC win (~halved) holds. README repro fixed: promptfoo needs --env-file ../.env (reads cwd, not root). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(benchmarks): add robustness audit — ponytail vs baseline on edge cases Answers the real question behind #65: does ponytail's push for the shortest solution make weak models produce wrong code on edge cases? robustness-audit.js: 16 self-verifying tasks (12 algorithmic edge-case traps + 4 validators). Each check ships a known-good and known-lazy-wrong reference that must pass/fail before any model output is scored (--selftest, 16/16). Findings (gpt-4.1-mini + gpt-5.4-mini, baseline vs ponytail): parity on every edge-case trap on both models. The one measured soft spot is gpt-5.4-mini email (~4-5%, reaches for parseaddr). A sharpened SKILL.md validation rule had no reliable effect in an n=100 A/B (96% vs 95%), so it was not shipped — the tendency is model-level, not skill-level. Full writeup in results/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(benchmarks): email slip is provider-specific — 100% on Claude High-n cross-provider follow-up to the robustness audit. The one ponytail soft spot (email validation via parseaddr) splits by provider, not model size: - Claude (haiku/sonnet/opus): 100% under ponytail, n=40 each — and ponytail beats baseline (unconstrained Sonnet over-engineers into an always-truthy dict, 0/40; ponytail writes a clean validator). - OpenAI (gpt-4.1-mini..gpt-5.5): slips at every size under ponytail (~79-98%), baseline ~100%. The parseaddr reflex lives in OpenAI training. Not fixable by skill text: 8 distinct SKILL.md edits (incl. an n=100 A/B, 96% vs 95%) all scored <= current, several worse, all bloated LOC. Nothing shipped. SKILL.md unchanged. Conclusion: on ponytail's target platform (Claude) email is 100%; the GPT slip is a documented cross-provider transfer quirk. Adds model-email.js / claude-email.js to reproduce the tables. Writeup updated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(benchmarks): correct misleading Sonnet baseline 0 percent The Sonnet baseline 0/40 on email is a return-type artifact, not a logic failure: unconstrained Sonnet returns a dict {is_valid, message} instead of a bool, so the bool-contract gate scores every case as accepted. Read dict-aware via is_valid, its logic is ~75% correct (9/12). Reframed honestly so we are not presenting 0 vs 100 as a clean win; ponytail still wins (clean 100% bool) but the point is over-engineered return type, not total failure. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Benchmark
Three arms (no skill, caveman, ponytail), three models, five everyday tasks, 10 runs per cell, median reported. Code LOC is counted from fenced code blocks; tokens, cost, and latency come straight from the API.
Reproduce
Claude (Haiku / Sonnet / Opus)
Requires an Anthropic API key and Node.js ≥ 22.22.0 (promptfoo's engine constraint —
check with node --version and upgrade if needed):
cp ../.env.example ../.env # add your ANTHROPIC_API_KEY
npx promptfoo@latest eval -c promptfooconfig.yaml --env-file ../.env --repeat 10
npx promptfoo@latest view
--env-file ../.env is required because promptfoo reads .env from the current
directory (benchmarks/), not the repo root where the file lives.
Local models via Ollama
No API key or promptfoo required. Runs against any model served by Ollama:
ollama pull llama3.2 # or any other model
python benchmarks/benchmark-local.py --model llama3.2 --repeat 3
See benchmarks/results/2026-06-15-llama3.2-local.md for what to expect: the skill works
well on instruction-following models (Claude-class) but transfers poorly to small local
models where the multi-step decision ladder isn't reliably followed.
Tasks: email validator, JS debounce, CSV sum, React countdown, FastAPI rate-limit (see promptfooconfig.yaml). Single-shot completions, default temperature.
Median results (10 runs, 2026-06-13)
Code (lines)
| arm | Haiku | Sonnet | Opus |
|---|---|---|---|
| baseline (no skill) | 518 | 693 | 256 |
| caveman | 116 | 120 | 67 |
| ponytail | 39 | 44 | 51 |
Cost (USD, 5 tasks)
| arm | Haiku | Sonnet | Opus |
|---|---|---|---|
| baseline (no skill) | 0.032 | 0.141 | 0.135 |
| caveman | 0.014 | 0.045 | 0.075 |
| ponytail | 0.010 | 0.032 | 0.071 |
Latency (seconds, 5 tasks)
| arm | Haiku | Sonnet | Opus |
|---|---|---|---|
| baseline (no skill) | 37.7 | 124.1 | 58.7 |
| caveman | 14.9 | 34.7 | 23.1 |
| ponytail | 9.9 | 20.1 | 18.0 |
Versus baseline, ponytail writes 80-94% less code, costs 47-77% less, and runs 3-6x faster, on every model.
Metrics
| File | Metric | Behavior |
|---|---|---|
loc.js |
loc |
Measurement - always passes, records line count |
correctness.js |
correct |
Gate - fails if generated code doesn't work |
correctness.js extracts fenced code blocks and runs per-task checks (spawns Python/Node for email, debounce, CSV; structural regex for React and FastAPI). A broken one-liner that scores great on LOC will fail on correctness.
Note: The React countdown and FastAPI rate-limit checks are keyword/structural only (no runtime execution), so they verify plausible structure rather than full correctness. The email, debounce, and CSV checks execute the code.
Prerequisites
Running the benchmark requires Python 3, pandas, and Node.js (18+).
Notes
- Caveman is a prose-compression skill (it leaves code "normal"), so it lands between baseline and ponytail on code size and wins mainly on prose tokens.
- Cost reflects single-shot calls that re-send the skill every time. In real sessions the skill is injected once and prompt-cached, so the cost gap widens further in ponytail's favor.
- These are everyday tasks. For production-grade specs, where an unconstrained agent bloats much harder, see the writeups in
results/.