The examples/ before/after blocks were authored by hand, not produced by a model. Issue #127 correctly noted that nobody hand-rolls quicksort for "sort this array" - every model just calls .sort(). Regenerate all examples verbatim from a real benchmark run (Claude Haiku 4.5, no-skill arm vs ponytail arm, benchmarks/output.json) so the before/after is reproducible, not authored: email 75->3, debounce 116->10, csv 20->3, countdown 267->9, rate-limit 128->10 LOC - Delete sorting.md (pure strawman) plus the other hand-written caricatures (api-endpoint, caching, date-picker) - Add benchmarks/generate-examples.mjs to regenerate examples from any run - examples/README.md indexes the set and documents how to reproduce Closes #127 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
18 lines
774 B
Markdown
18 lines
774 B
Markdown
# Examples
|
|
|
|
Real model output, verbatim from benchmark runs — the same task answered by the same model
|
|
with no skill (`## Without Ponytail`) and with ponytail (`## With Ponytail`), so you can
|
|
compare side by side. Model: Claude Haiku 4.5, temperature 1, source `benchmarks/output.json`.
|
|
|
|
These are not hand-written. Reproduce them yourself:
|
|
`npx promptfoo@latest eval -c benchmarks/promptfooconfig.yaml`. Method, all three models, and
|
|
median-of-10 numbers: [../benchmarks/](../benchmarks/).
|
|
|
|
| Example | Without (LOC) | With (LOC) |
|
|
|---|--:|--:|
|
|
| [Email Validation](email-validation.md) | 75 | 3 |
|
|
| [Debounce](debounce.md) | 116 | 10 |
|
|
| [CSV Sum](csv-sum.md) | 20 | 3 |
|
|
| [Countdown Timer](react-countdown.md) | 267 | 9 |
|
|
| [Rate Limiting](rate-limit.md) | 128 | 10 |
|