* fix(benchmarks): correctness gate scores unfenced code; fix debounce task The `correct` gate under-reported correctness for terse models, the likely source of "Ponytail degrades models" reports (issue #65): - extractBlocks() only matched fenced code blocks, so bare/unfenced code scored an automatic fail even when correct. Now falls back to the whole response as one block (and tolerates CRLF). Debounce detection also accepts unfenced arrow functions. - The debounce task asked to "add debounce to a search input" but the check expected a reusable debounce(fn, delay) util, failing correct inline answers. Task reworded to the deliverable the check verifies. Adds correctness.test.js (regression guard) and a GPT-mini repro config plus results writeup: on a clean n=20 run, the reported gpt-4.1-mini drop (10/15) does not reproduce (100/100). The LOC win (~halved) holds. README repro fixed: promptfoo needs --env-file ../.env (reads cwd, not root). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(benchmarks): add robustness audit — ponytail vs baseline on edge cases Answers the real question behind #65: does ponytail's push for the shortest solution make weak models produce wrong code on edge cases? robustness-audit.js: 16 self-verifying tasks (12 algorithmic edge-case traps + 4 validators). Each check ships a known-good and known-lazy-wrong reference that must pass/fail before any model output is scored (--selftest, 16/16). Findings (gpt-4.1-mini + gpt-5.4-mini, baseline vs ponytail): parity on every edge-case trap on both models. The one measured soft spot is gpt-5.4-mini email (~4-5%, reaches for parseaddr). A sharpened SKILL.md validation rule had no reliable effect in an n=100 A/B (96% vs 95%), so it was not shipped — the tendency is model-level, not skill-level. Full writeup in results/. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(benchmarks): email slip is provider-specific — 100% on Claude High-n cross-provider follow-up to the robustness audit. The one ponytail soft spot (email validation via parseaddr) splits by provider, not model size: - Claude (haiku/sonnet/opus): 100% under ponytail, n=40 each — and ponytail beats baseline (unconstrained Sonnet over-engineers into an always-truthy dict, 0/40; ponytail writes a clean validator). - OpenAI (gpt-4.1-mini..gpt-5.5): slips at every size under ponytail (~79-98%), baseline ~100%. The parseaddr reflex lives in OpenAI training. Not fixable by skill text: 8 distinct SKILL.md edits (incl. an n=100 A/B, 96% vs 95%) all scored <= current, several worse, all bloated LOC. Nothing shipped. SKILL.md unchanged. Conclusion: on ponytail's target platform (Claude) email is 100%; the GPT slip is a documented cross-provider transfer quirk. Adds model-email.js / claude-email.js to reproduce the tables. Writeup updated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(benchmarks): correct misleading Sonnet baseline 0 percent The Sonnet baseline 0/40 on email is a return-type artifact, not a logic failure: unconstrained Sonnet returns a dict {is_valid, message} instead of a bool, so the bool-contract gate scores every case as accepted. Read dict-aware via is_valid, its logic is ~75% correct (9/12). Reframed honestly so we are not presenting 0 vs 100 as a clean win; ponytail still wins (clean 100% bool) but the point is over-engineered return type, not total failure. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
84 lines
3.4 KiB
Markdown
84 lines
3.4 KiB
Markdown
# Benchmark
|
|
|
|
Three arms (no skill, [caveman](https://github.com/JuliusBrussee/caveman), ponytail), three models, five everyday tasks, **10 runs per cell, median reported**. Code LOC is counted from fenced code blocks; tokens, cost, and latency come straight from the API.
|
|
|
|
## Reproduce
|
|
|
|
### Claude (Haiku / Sonnet / Opus)
|
|
|
|
Requires an Anthropic API key and **Node.js ≥ 22.22.0** (promptfoo's engine constraint —
|
|
check with `node --version` and upgrade if needed):
|
|
|
|
```bash
|
|
cp ../.env.example ../.env # add your ANTHROPIC_API_KEY
|
|
npx promptfoo@latest eval -c promptfooconfig.yaml --env-file ../.env --repeat 10
|
|
npx promptfoo@latest view
|
|
```
|
|
|
|
`--env-file ../.env` is required because promptfoo reads `.env` from the current
|
|
directory (`benchmarks/`), not the repo root where the file lives.
|
|
|
|
### Local models via Ollama
|
|
|
|
No API key or promptfoo required. Runs against any model served by Ollama:
|
|
|
|
```bash
|
|
ollama pull llama3.2 # or any other model
|
|
python benchmarks/benchmark-local.py --model llama3.2 --repeat 3
|
|
```
|
|
|
|
See `benchmarks/results/2026-06-15-llama3.2-local.md` for what to expect: the skill works
|
|
well on instruction-following models (Claude-class) but transfers poorly to small local
|
|
models where the multi-step decision ladder isn't reliably followed.
|
|
|
|
Tasks: email validator, JS debounce, CSV sum, React countdown, FastAPI rate-limit (see `promptfooconfig.yaml`). Single-shot completions, default temperature.
|
|
|
|
## Median results (10 runs, 2026-06-13)
|
|
|
|
**Code (lines)**
|
|
|
|
| arm | Haiku | Sonnet | Opus |
|
|
|---|--:|--:|--:|
|
|
| baseline (no skill) | 518 | 693 | 256 |
|
|
| caveman | 116 | 120 | 67 |
|
|
| **ponytail** | **39** | **44** | **51** |
|
|
|
|
**Cost (USD, 5 tasks)**
|
|
|
|
| arm | Haiku | Sonnet | Opus |
|
|
|---|--:|--:|--:|
|
|
| baseline (no skill) | 0.032 | 0.141 | 0.135 |
|
|
| caveman | 0.014 | 0.045 | 0.075 |
|
|
| **ponytail** | **0.010** | **0.032** | **0.071** |
|
|
|
|
**Latency (seconds, 5 tasks)**
|
|
|
|
| arm | Haiku | Sonnet | Opus |
|
|
|---|--:|--:|--:|
|
|
| baseline (no skill) | 37.7 | 124.1 | 58.7 |
|
|
| caveman | 14.9 | 34.7 | 23.1 |
|
|
| **ponytail** | **9.9** | **20.1** | **18.0** |
|
|
|
|
Versus baseline, ponytail writes **80-94% less code**, costs **47-77% less**, and runs **3-6x faster**, on every model.
|
|
|
|
## Metrics
|
|
|
|
| File | Metric | Behavior |
|
|
|------|--------|----------|
|
|
| `loc.js` | `loc` | Measurement - always passes, records line count |
|
|
| `correctness.js` | `correct` | Gate - fails if generated code doesn't work |
|
|
|
|
`correctness.js` extracts fenced code blocks and runs per-task checks (spawns Python/Node for email, debounce, CSV; structural regex for React and FastAPI). A broken one-liner that scores great on LOC will fail on correctness.
|
|
|
|
> **Note:** The React countdown and FastAPI rate-limit checks are keyword/structural only (no runtime execution), so they verify plausible structure rather than full correctness. The email, debounce, and CSV checks execute the code.
|
|
|
|
### Prerequisites
|
|
|
|
Running the benchmark requires **Python 3**, **pandas**, and **Node.js** (18+).
|
|
|
|
## Notes
|
|
|
|
- Caveman is a prose-compression skill (it leaves code "normal"), so it lands between baseline and ponytail on code size and wins mainly on prose tokens.
|
|
- Cost reflects single-shot calls that re-send the skill every time. In real sessions the skill is injected once and prompt-cached, so the cost gap widens further in ponytail's favor.
|
|
- These are everyday tasks. For production-grade specs, where an unconstrained agent bloats much harder, see the writeups in `results/`.
|