Files
ponytail/benchmarks/results/2026-06-16-robustness-audit.md
DietrichGebertandClaude Opus 4.8 caf138df56 benchmarks: fix correctness gate + robustness audit (#65) (#83)
* fix(benchmarks): correctness gate scores unfenced code; fix debounce task

The `correct` gate under-reported correctness for terse models, the likely
source of "Ponytail degrades models" reports (issue #65):

- extractBlocks() only matched fenced code blocks, so bare/unfenced code
  scored an automatic fail even when correct. Now falls back to the whole
  response as one block (and tolerates CRLF). Debounce detection also accepts
  unfenced arrow functions.
- The debounce task asked to "add debounce to a search input" but the check
  expected a reusable debounce(fn, delay) util, failing correct inline answers.
  Task reworded to the deliverable the check verifies.

Adds correctness.test.js (regression guard) and a GPT-mini repro config plus
results writeup: on a clean n=20 run, the reported gpt-4.1-mini drop (10/15)
does not reproduce (100/100). The LOC win (~halved) holds.

README repro fixed: promptfoo needs --env-file ../.env (reads cwd, not root).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(benchmarks): add robustness audit — ponytail vs baseline on edge cases

Answers the real question behind #65: does ponytail's push for the shortest
solution make weak models produce wrong code on edge cases?

robustness-audit.js: 16 self-verifying tasks (12 algorithmic edge-case traps +
4 validators). Each check ships a known-good and known-lazy-wrong reference that
must pass/fail before any model output is scored (--selftest, 16/16).

Findings (gpt-4.1-mini + gpt-5.4-mini, baseline vs ponytail): parity on every
edge-case trap on both models. The one measured soft spot is gpt-5.4-mini email
(~4-5%, reaches for parseaddr). A sharpened SKILL.md validation rule had no
reliable effect in an n=100 A/B (96% vs 95%), so it was not shipped — the
tendency is model-level, not skill-level. Full writeup in results/.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(benchmarks): email slip is provider-specific — 100% on Claude

High-n cross-provider follow-up to the robustness audit. The one ponytail
soft spot (email validation via parseaddr) splits by provider, not model size:

- Claude (haiku/sonnet/opus): 100% under ponytail, n=40 each — and ponytail
  beats baseline (unconstrained Sonnet over-engineers into an always-truthy
  dict, 0/40; ponytail writes a clean validator).
- OpenAI (gpt-4.1-mini..gpt-5.5): slips at every size under ponytail
  (~79-98%), baseline ~100%. The parseaddr reflex lives in OpenAI training.

Not fixable by skill text: 8 distinct SKILL.md edits (incl. an n=100 A/B,
96% vs 95%) all scored <= current, several worse, all bloated LOC. Nothing
shipped. SKILL.md unchanged.

Conclusion: on ponytail's target platform (Claude) email is 100%; the GPT
slip is a documented cross-provider transfer quirk. Adds model-email.js /
claude-email.js to reproduce the tables. Writeup updated.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(benchmarks): correct misleading Sonnet baseline 0 percent

The Sonnet baseline 0/40 on email is a return-type artifact, not a logic
failure: unconstrained Sonnet returns a dict {is_valid, message} instead of a
bool, so the bool-contract gate scores every case as accepted. Read dict-aware
via is_valid, its logic is ~75% correct (9/12). Reframed honestly so we are not
presenting 0 vs 100 as a clean win; ponytail still wins (clean 100% bool) but
the point is over-engineered return type, not total failure.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 12:17:10 +02:00

6.4 KiB
Raw Permalink Blame History

Robustness audit: does ponytail degrade weak models? (2026-06-16)

Follow-up to issue #65. After fixing the correctness-gate bugs, the open question was the real one: does Ponytail's push toward the shortest solution make weak models produce wrong code on edge cases? This audit answers it directly, with a deliberately hostile test set and high sample counts.

TL;DR

  • Across 12 classic edge-case traps (off-by-one, n=0, leap-century, subtractive Roman, deep nesting, …) on two weak models (gpt-4.1-mini, gpt-5.4-mini), Ponytail holds baseline parity — it does not produce more wrong answers than the unconstrained model.
  • The one measured soft spot is email validation, and it is provider-specific. OpenAI models, at every size, sometimes reach for email.utils.parseaddr (a parser, not a validator) under "stdlib-first" pressure and accept "@missing-local.com". On Claude, ponytail's target platform, email is 100% (haiku/sonnet/opus, n=40 each).
  • The slip is not fixable by skill text: 8 distinct SKILL.md edits (including an n=100 A/B, 96% → 95%) all scored ≤ the current skill, several worse, all bloating LOC. Counter- instructions make small models overthink and fail more. Nothing was shipped — adding skill text that doesn't move the number is exactly the cargo-cult Ponytail exists to avoid.

Method

baseline (no skill) vs ponytail (full SKILL.md), single-shot, default params, gpt-4.1-mini and gpt-5.4-mini. Each task runs generated code against edge-case assertions. Every check is self-verified: a known-correct and a known-lazy-wrong reference must pass/fail respectively before any model output is scored (node robustness-audit.js --selftest, 16/16). Runs were serial to avoid quota 429s shrinking denominators.

Edge-case traps (n=20/cell)

All 12 algorithmic tasks: baseline 20/20 == ponytail 20/20 on both models. Examples of the traps (the lazy version passes the common case, fails the edge):

task the trap a lazy impl misses
is_prime n = 0, 1, negatives
factorial / fibonacci n = 0
binary_search empty list, target at the last index (off-by-one)
is_leap_year / days_in_month 1900 not leap, 2000 leap (century rule)
int_to_roman subtractive forms (4=IV, 9=IX, 40=XL)
flatten nesting deeper than one level
clamp value already in range
chunk trailing remainder

The only sub-20 cell in the first run was gpt-5.4-mini flatten at 19/20 — a single stochastic miss that did not reproduce: 50/50 at n=50. (clamp showed 19/19, i.e. one API error, not a wrong answer.)

Validators: the email slip is provider-specific

The one place ponytail measurably affects correctness is email validation, via the parse ≠ validate trap: under "stdlib-first" pressure a model reaches for email.utils.parseaddr — a parser that accepts malformed input like @missing-local.com — instead of writing an explicit check. The split is by provider, not model size.

OpenAI (email, baseline vs ponytail, n=50100):

model baseline ponytail
gpt-4.1-mini 100% 98%
gpt-4.1 100% 79%
gpt-5.4-mini ~100% ~92%
gpt-5.4 100% 98%
gpt-5.5 98% 94%

Claude (email, baseline vs ponytail, n=40):

model baseline ponytail
claude-haiku-4-5 35/40 40/40
claude-sonnet-4-6 0/40 * 40/40
claude-opus-4-8 39/40 40/40

Every OpenAI model slips regardless of size (gpt-4.1 full is the worst). Every Claude model is 100% under ponytail.

* The Sonnet baseline 0/40 is a return-type artifact, not a logic failure, and should not be read as "Sonnet cannot validate email." Unconstrained Sonnet over-engineers the validator into a dict ({is_valid, message}) instead of a bool. The test calls the function as a bool, and a non-empty dict is always truthy, so it "accepts" every address and scores 0. Read dict-aware (via is_valid), its logic is about 75% correct (9/12). The honest point is narrow: ponytail writes the plain correct bool the task implies, while the unconstrained model over-builds the interface and trips a naive if validate(x) caller. url, creditcard, and ipv4 hold at ~100% under ponytail on both providers, because their lazy stdlib choice (ipaddress, Luhn, scheme checks) is already strict. Only email's obvious stdlib tool is a parser.

The fix that wasn't

SKILL.md already says "never simplify away input validation" and "pick the stdlib option correct on edge cases." We tried hard to push the OpenAI rate to 100% by editing the skill — 8 distinct edits across counter-pressure wording, a check-mandate, explicit-over-delegate, a few-shot example, combinations, and three placements. Every one scored ≤ the current skill; several were far worse (one cratered to 78%); all bloated median LOC. The definitive n=100 A/B of the most promising edit:

OLD skill: 96/100 (96.0%)
NEW skill: 95/100 (95.0%)   -> within noise, no reliable effect

Counter-instructions backfire: piling validation rules onto the skill makes models overthink and produce more broken validators, not fewer. The reflex to reach for parseaddr lives in the OpenAI models' training, and no skill wording reliably overrides it — so nothing was shipped. Adding skill text that doesn't work is the cargo-cult Ponytail exists to prevent.

Conclusion

"Ponytail degrades model performance" is not supported. Across 12 edge-case traps, ponytail holds baseline parity. On validation it is 100% on every Claude model, which is its target platform. The only blemish is an email-validator slip on OpenAI models (a cross-provider parseaddr reflex, present at every size), documented here and not fixable by skill text. The LOC win (about half the code) comes with no correctness tax on Claude.

Reproduce

cd benchmarks
node robustness-audit.js --selftest        # verify all 16 instruments (no API)
node robustness-audit.js                    # 16-task audit, gpt-5.4-mini, n=20
AUDIT_MODEL=gpt-4.1-mini node robustness-audit.js

# email cross-provider (the slip)
ME_MODELS="gpt-4.1,gpt-5.4,gpt-5.5" ME_N=50 node model-email.js   # OpenAI  (OPENAI_API_KEY)
node claude-email.js                                              # Claude  (ANTHROPIC_API_KEY)

OPENAI_API_KEY / ANTHROPIC_API_KEY read from ../.env.