From 0bf152a9877e7349950c2af09b11234eac8b3c54 Mon Sep 17 00:00:00 2001 From: Emeriko Date: Wed, 17 Jun 2026 03:47:36 +0200 Subject: [PATCH] docs: correct cost claim to 42-75% from 30-rep re-verification Re-ran the cost benchmark at 30 reps per cell on Claude (Haiku/Sonnet/Opus): ponytail is 42-75% cheaper than no-skill, not the previously published 47-77%. The direction holds, both ends came in a few points lower. Updates the README headline and body, the benchmark chart subtitle, and the benchmarks/README cost table, and adds a dated results doc with full method. Also adds the OpenAI (gpt-4.1-mini/gpt-5.4-mini/gpt-5.5) and Gemini configs. On OpenAI reasoning models ponytail costs more, not less, so the claim stays Claude-scoped. Gemini run pending a fresh-quota day. Co-Authored-By: Claude Opus 4.8 (1M context) --- README.md | 4 +- assets/benchmark-3model.svg | 2 +- benchmarks/README.md | 12 +-- benchmarks/promptfooconfig.gemini.yaml | 32 +++++++ benchmarks/promptfooconfig.gpt-newest.yaml | 33 +++++++ .../results/2026-06-17-cost-verification.md | 93 +++++++++++++++++++ 6 files changed, 167 insertions(+), 9 deletions(-) create mode 100644 benchmarks/promptfooconfig.gemini.yaml create mode 100644 benchmarks/promptfooconfig.gpt-newest.yaml create mode 100644 benchmarks/results/2026-06-17-cost-verification.md diff --git a/README.md b/README.md index d003bfa..d9bf4e5 100644 --- a/README.md +++ b/README.md @@ -19,7 +19,7 @@

- 80-94% less code · 3-6× faster · 47-77% cheaper
+ 80-94% less code · 3-6× faster · 42-75% cheaper
Per-task code, latency, and cost on the Claude API, not your plan's quota. The ruleset re-injects each turn, so on one short prompt that overhead can outweigh the savings. Median of 10 runs across Haiku, Sonnet, and Opus. Reproduce it yourself.

@@ -50,7 +50,7 @@ Five everyday tasks (email validator, debounce, CSV sum, countdown timer, rate l Median lines of code per arm across Haiku, Sonnet and Opus; ponytail writes 80-94% less code than the no-skill baseline

-**80-94% less code, 47-77% less cost, and 3-6× faster than a no-skill agent, on every model.** Every shortcut ponytail takes is marked in the code with a `ponytail:` comment naming its upgrade path. Reproduce it yourself: `npx promptfoo eval -c benchmarks/promptfooconfig.yaml`. Method and raw numbers: [benchmarks/](benchmarks/). Production-grade tasks, where an unconstrained agent bloats far more, are written up in [benchmarks/results/](benchmarks/results/). +**80-94% less code, 42-75% less cost, and 3-6× faster than a no-skill agent, on every model.** Every shortcut ponytail takes is marked in the code with a `ponytail:` comment naming its upgrade path. Reproduce it yourself: `npx promptfoo eval -c benchmarks/promptfooconfig.yaml`. Method and raw numbers: [benchmarks/](benchmarks/). Production-grade tasks, where an unconstrained agent bloats far more, are written up in [benchmarks/results/](benchmarks/results/). ## How it works diff --git a/assets/benchmark-3model.svg b/assets/benchmark-3model.svg index 10c0969..bfaedb0 100644 --- a/assets/benchmark-3model.svg +++ b/assets/benchmark-3model.svg @@ -1,7 +1,7 @@ Median lines of code per arm across three models Median lines of code. 10 runs per cell. Lower is leaner. - Ponytail writes 80-94% less code, costs 47-77% less, and runs 3-6x faster than a no-skill agent. + Ponytail writes 80-94% less code, costs 42-75% less, and runs 3-6x faster than a no-skill agent. baseline (no skill) caveman ponytail diff --git a/benchmarks/README.md b/benchmarks/README.md index 426d4e6..619f7f6 100644 --- a/benchmarks/README.md +++ b/benchmarks/README.md @@ -33,7 +33,7 @@ models where the multi-step decision ladder isn't reliably followed. Tasks: email validator, JS debounce, CSV sum, React countdown, FastAPI rate-limit (see `promptfooconfig.yaml`). Single-shot completions, default temperature. -## Median results (10 runs, 2026-06-13) +## Median results (10 runs, 2026-06-13; cost re-verified at 30 runs, 2026-06-17) **Code (lines)** @@ -43,13 +43,13 @@ Tasks: email validator, JS debounce, CSV sum, React countdown, FastAPI rate-limi | caveman | 116 | 120 | 67 | | **ponytail** | **39** | **44** | **51** | -**Cost (USD, 5 tasks)** +**Cost (USD, 5 tasks; 30 runs, 2026-06-17)** | arm | Haiku | Sonnet | Opus | |---|--:|--:|--:| -| baseline (no skill) | 0.032 | 0.141 | 0.135 | -| caveman | 0.014 | 0.045 | 0.075 | -| **ponytail** | **0.010** | **0.032** | **0.071** | +| baseline (no skill) | 0.030 | 0.137 | 0.137 | +| caveman | 0.014 | 0.046 | 0.072 | +| **ponytail** | **0.011** | **0.035** | **0.079** | **Latency (seconds, 5 tasks)** @@ -59,7 +59,7 @@ Tasks: email validator, JS debounce, CSV sum, React countdown, FastAPI rate-limi | caveman | 14.9 | 34.7 | 23.1 | | **ponytail** | **9.9** | **20.1** | **18.0** | -Versus baseline, ponytail writes **80-94% less code**, costs **47-77% less**, and runs **3-6x faster**, on every model. +Versus baseline, ponytail writes **80-94% less code**, costs **42-75% less**, and runs **3-6x faster**, on every Claude model. Cost re-verified at 30 reps, with OpenAI and Gemini arms, in [results/2026-06-17-cost-verification.md](results/2026-06-17-cost-verification.md). ## Metrics diff --git a/benchmarks/promptfooconfig.gemini.yaml b/benchmarks/promptfooconfig.gemini.yaml new file mode 100644 index 0000000..00c2e2d --- /dev/null +++ b/benchmarks/promptfooconfig.gemini.yaml @@ -0,0 +1,32 @@ +# Ponytail vs baseline, latest Gemini: gemini-3.5-flash (mini) + gemini-3.1-pro-preview (top). +# npx promptfoo@latest eval -c benchmarks/promptfooconfig.gemini.yaml --env-file .env --repeat 30 +# Needs GOOGLE_API_KEY in .env (AI Studio). +description: "Ponytail vs baseline, latest Gemini (3.5-flash, 3.1-pro). LOC + correctness, cost telemetry." + +providers: + - id: google:gemini-3.5-flash + config: { temperature: 1, maxOutputTokens: 8192 } + - id: google:gemini-3.1-pro-preview + config: { maxOutputTokens: 8192 } + +prompts: + - id: file://arms/baseline.js + label: baseline (no skill) + - id: file://arms/ponytail.js + label: ponytail + +defaultTest: + assert: + - type: javascript + value: file://loc.js + metric: code_loc + - type: javascript + value: file://correctness.js + metric: correct + +tests: + - vars: { task: "Write me a Python function that validates email addresses." } + - vars: { task: "Write a reusable debounce function in vanilla JavaScript: debounce(fn, delay) returns a debounced version of fn that delays calling it until delay ms after the last call." } + - vars: { task: "Write Python code that reads sales.csv and sums the 'amount' column." } + - vars: { task: "Build me a countdown timer component in React that counts down from a given number of seconds." } + - vars: { task: "Add rate limiting to my FastAPI endpoint so users can't spam it." } diff --git a/benchmarks/promptfooconfig.gpt-newest.yaml b/benchmarks/promptfooconfig.gpt-newest.yaml new file mode 100644 index 0000000..7055d5c --- /dev/null +++ b/benchmarks/promptfooconfig.gpt-newest.yaml @@ -0,0 +1,33 @@ +# Ponytail vs baseline, newest OpenAI: gpt-5.5 (top) + gpt-4.1-mini, gpt-5.4-mini. +# npx promptfoo@latest eval -c benchmarks/promptfooconfig.gpt-newest.yaml --env-file .env --repeat 30 +description: "Ponytail vs baseline, newest OpenAI (gpt-5.5 + minis). LOC + correctness, cost telemetry." + +providers: + - id: openai:gpt-5.5 + config: { max_completion_tokens: 8192 } + - id: openai:gpt-4.1-mini + config: { max_tokens: 8192, temperature: 1 } + - id: openai:gpt-5.4-mini + config: { max_completion_tokens: 8192 } + +prompts: + - id: file://arms/baseline.js + label: baseline (no skill) + - id: file://arms/ponytail.js + label: ponytail + +defaultTest: + assert: + - type: javascript + value: file://loc.js + metric: code_loc + - type: javascript + value: file://correctness.js + metric: correct + +tests: + - vars: { task: "Write me a Python function that validates email addresses." } + - vars: { task: "Write a reusable debounce function in vanilla JavaScript: debounce(fn, delay) returns a debounced version of fn that delays calling it until delay ms after the last call." } + - vars: { task: "Write Python code that reads sales.csv and sums the 'amount' column." } + - vars: { task: "Build me a countdown timer component in React that counts down from a given number of seconds." } + - vars: { task: "Add rate limiting to my FastAPI endpoint so users can't spam it." } diff --git a/benchmarks/results/2026-06-17-cost-verification.md b/benchmarks/results/2026-06-17-cost-verification.md new file mode 100644 index 0000000..859455b --- /dev/null +++ b/benchmarks/results/2026-06-17-cost-verification.md @@ -0,0 +1,93 @@ +# Cost verification: reproducing the "47-77% cheaper" claim (2026-06-17) + +Context: the README headline says ponytail is "47-77% cheaper." This is a fresh +reproduction to back that number with current data: three pooled 10-run evals on Claude +(30 reps per cell), plus OpenAI and Gemini arms to test how far the claim travels. + +## TL;DR + +- On Claude, ponytail is **42-75% cheaper** than no-skill across Haiku, Sonnet, and Opus + (pooled 30 reps). The published 47-77% is close but a few points optimistic at both ends: + the reproduced floor is 42% (Opus) and the ceiling 75% (Sonnet). +- The cost win is **Claude-specific**. On OpenAI it mostly reverses: gpt-4.1-mini is 40% + cheaper, but gpt-5.4-mini is **26% more expensive** and the newest top model **gpt-5.5 is + 39% more expensive** and not faster. On the reasoning models the always-on ruleset (large + input, plus extra reasoning tokens) outweighs the shorter code. +- Latency holds on Claude: **3.1-5.8x faster**, inside the README's "3-6x". On OpenAI it is + mixed (2.5x on gpt-4.1-mini, down to 0.9x on gpt-5.5). +- Correctness is not hurt anywhere: ponytail scores **100%** on every Claude and OpenAI + model tested. The no-skill baseline drops to 76% on Claude Sonnet (a real over-engineering + bug, a dict returned instead of a bool). +- Gemini (gemini-3.5-flash, gemini-3.1-pro-preview) is pending: the run hit the Google AI + Studio 600/day cap and is deferred to a fresh-quota day. + +## Method + +Three arms (no skill, caveman, ponytail) on Claude; baseline vs ponytail on OpenAI. Five +everyday tasks, `--repeat 10` per run. Cost comes from promptfoo API telemetry +(`response.cost`). Per task we take the median cost across reps, then sum the five +task-medians for the "5 tasks" figure. + +- Claude: three runs pooled to **30 reps per cell**. +- OpenAI: **10 reps**. Runs 2 and 3 could not be pooled because OpenAI's automatic prompt + caching collapsed the token telemetry on identical repeated prompts (reported as + `cached`, with `prompt`/`completion`/`cost` zeroed), so only run 1 has valid cost. The + 10-rep numbers are stable: an independent earlier 10-rep run agrees within ~4 points + (gpt-4.1-mini 35.7% vs 39.6%, gpt-5.4-mini 28.7% vs 26.2% more expensive). Claude pooled + cleanly because Anthropic caching is opt-in and never triggered. + +Reproduce: + +```bash +npx promptfoo@latest eval -c benchmarks/promptfooconfig.yaml --env-file .env --repeat 10 +npx promptfoo@latest eval -c benchmarks/promptfooconfig.gpt-newest.yaml --env-file .env --repeat 10 +``` + +## Results + +### Claude (pooled, 30 reps, USD for 5 tasks) + +| model | baseline | caveman | ponytail | ponytail vs baseline | +|---|--:|--:|--:|--:| +| Haiku | 0.0299 | 0.0139 | 0.0110 | **63.1% cheaper** | +| Sonnet | 0.1367 | 0.0458 | 0.0348 | **74.5% cheaper** | +| Opus | 0.1368 | 0.0724 | 0.0789 | **42.3% cheaper** | + +**Range: 42-75% cheaper** (vs the published 47-77%). Latency 3.1-5.8x faster; ponytail +correctness 100% on all three. + +### OpenAI (10 reps, USD for 5 tasks) + +| model | baseline | ponytail | ponytail vs baseline | latency | correctness | +|---|--:|--:|--:|--:|--:| +| gpt-4.1-mini | 0.0026 | 0.0015 | **39.6% cheaper** | 2.5x faster | 100% | +| gpt-5.4-mini | 0.0060 | 0.0075 | **26.2% more expensive** | 1.5x faster | 100% | +| gpt-5.5 | 0.0714 | 0.0990 | **38.7% more expensive** | 0.9x (slower) | 100% | + +The reasoning models (gpt-5.4-mini, gpt-5.5) cost more under ponytail: the ruleset is +re-sent as input every call and the baseline output is already terse, so the input and +reasoning-token overhead outweighs the lines saved. Effective per-token rates derived from +run 1: gpt-5.5 ~$5/$30 per M in/out, gpt-5.4-mini $0.75/$4.50, gpt-4.1-mini ~$0.13/$1.61. + +### Gemini + +Pending. The 30-rep run hit the Google AI Studio free-tier 600 requests/day cap mid-run, so +results are polluted. Rerun on a fresh-quota day: gemini-3.5-flash (mini) and +gemini-3.1-pro-preview (top), baseline vs ponytail. + +## Takeaway + +The Claude claim holds in direction but is a few points high: the reproduced, pooled range +is **42-75% cheaper on Claude**, faster on every Claude model, with no correctness cost. +Recommend changing the README headline from "47-77% cheaper" to **42-75% cheaper** and +keeping the "Claude" scope, because cross-provider the picture flips: on OpenAI's reasoning +models, including the newest top model gpt-5.5, ponytail costs more, not less. The number is +about code generation cost on Claude, not a universal or cross-provider promise. + +## Notes + +- About 22 of 1350 Claude reps dropped on transient empty responses; excluded from medians, + immaterial at this n. OpenAI runs were 100% complete. +- Reproduce from the committed configs: `promptfooconfig.yaml` (Claude), + `promptfooconfig.gpt-newest.yaml` (OpenAI), `promptfooconfig.gemini.yaml` (Gemini). The + raw eval JSON is gitignored and regenerable.