* fix(benchmarks): correctness gate scores unfenced code; fix debounce task
The `correct` gate under-reported correctness for terse models, the likely
source of "Ponytail degrades models" reports (issue #65):
- extractBlocks() only matched fenced code blocks, so bare/unfenced code
scored an automatic fail even when correct. Now falls back to the whole
response as one block (and tolerates CRLF). Debounce detection also accepts
unfenced arrow functions.
- The debounce task asked to "add debounce to a search input" but the check
expected a reusable debounce(fn, delay) util, failing correct inline answers.
Task reworded to the deliverable the check verifies.
Adds correctness.test.js (regression guard) and a GPT-mini repro config plus
results writeup: on a clean n=20 run, the reported gpt-4.1-mini drop (10/15)
does not reproduce (100/100). The LOC win (~halved) holds.
README repro fixed: promptfoo needs --env-file ../.env (reads cwd, not root).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(benchmarks): add robustness audit — ponytail vs baseline on edge cases
Answers the real question behind #65: does ponytail's push for the shortest
solution make weak models produce wrong code on edge cases?
robustness-audit.js: 16 self-verifying tasks (12 algorithmic edge-case traps +
4 validators). Each check ships a known-good and known-lazy-wrong reference that
must pass/fail before any model output is scored (--selftest, 16/16).
Findings (gpt-4.1-mini + gpt-5.4-mini, baseline vs ponytail): parity on every
edge-case trap on both models. The one measured soft spot is gpt-5.4-mini email
(~4-5%, reaches for parseaddr). A sharpened SKILL.md validation rule had no
reliable effect in an n=100 A/B (96% vs 95%), so it was not shipped — the
tendency is model-level, not skill-level. Full writeup in results/.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(benchmarks): email slip is provider-specific — 100% on Claude
High-n cross-provider follow-up to the robustness audit. The one ponytail
soft spot (email validation via parseaddr) splits by provider, not model size:
- Claude (haiku/sonnet/opus): 100% under ponytail, n=40 each — and ponytail
beats baseline (unconstrained Sonnet over-engineers into an always-truthy
dict, 0/40; ponytail writes a clean validator).
- OpenAI (gpt-4.1-mini..gpt-5.5): slips at every size under ponytail
(~79-98%), baseline ~100%. The parseaddr reflex lives in OpenAI training.
Not fixable by skill text: 8 distinct SKILL.md edits (incl. an n=100 A/B,
96% vs 95%) all scored <= current, several worse, all bloated LOC. Nothing
shipped. SKILL.md unchanged.
Conclusion: on ponytail's target platform (Claude) email is 100%; the GPT
slip is a documented cross-provider transfer quirk. Adds model-email.js /
claude-email.js to reproduce the tables. Writeup updated.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(benchmarks): correct misleading Sonnet baseline 0 percent
The Sonnet baseline 0/40 on email is a return-type artifact, not a logic
failure: unconstrained Sonnet returns a dict {is_valid, message} instead of a
bool, so the bool-contract gate scores every case as accepted. Read dict-aware
via is_valid, its logic is ~75% correct (9/12). Reframed honestly so we are not
presenting 0 vs 100 as a clean win; ponytail still wins (clean 100% bool) but
the point is over-engineered return type, not total failure.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The benchmark harness hardcoded `python`, which is missing on macOS and
many Linux images. Probe python3 first, add npm test, and run checks in
GitHub Actions so regressions are caught on every PR.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(benchmarks): add correctness assertion - proves less code is not broken code
The existing benchmark measures lines-of-code (loc.js) but never checks
whether the generated code actually works. This adds a functional
correctness gate (correctness.js) that extracts code from fenced blocks
and runs per-task checks:
- email validator: spawns Python, asserts accept/reject on 5 inputs
- debounce: spawns Node, asserts delayed execution + reset on re-call
- csv sum: spawns Python with a test CSV, asserts correct total (351)
- countdown (React): structural check (useState + useEffect + decrement)
- rate limiter (FastAPI): structural check (limit logic + framework usage)
12 unit tests (node:test) cover good/bad outputs for every task plus the
unknown-task edge case. Existing tests and rule-copy checks unaffected.
* fix: address review feedback
- csv check: use regex lookaround instead of substring match to prevent
false positives (e.g. 13510 containing '351')
- ratelimit: fix operator precedence in block finder by adding parens
around the || inside the !b.lang guard
- README: note that React/FastAPI checks are structural only, add
prerequisites section (Python 3, pandas, Node.js 18+)
- test: add regression test for csv substring false positive