Commit Graph
3 Commits
Author SHA1 Message Date
DietrichGebert 48cdf05a25 Revert "benchmarks: add system prompt to baseline arm so it doesn't ramble (closes #126) (#128)" (#175)
This reverts commit 37f46b8f02.
2026-06-19 01:44:13 +02:00
Sha256_NulledandFato07 37f46b8f02 benchmarks: add system prompt to baseline arm so it doesn't ramble (closes #126) (#128)
Co-authored-by: Fato07 <fato07@users.noreply.github.com>
2026-06-19 00:24:35 +02:00
EmerikoandClaude Opus 4.8 321a59c82f feat: reproducible promptfoo benchmark + 3-model results
Commit a promptfoo harness (config + arm prompts + LOC metric + vendored
caveman SKILL) so anyone can re-run the comparison: no-skill vs caveman vs
ponytail, across Haiku / Sonnet / Opus, 10 runs per cell, median reported.

Replace the old unreproducible 6-task chart with assets/benchmark-3model.svg
from this run, and reframe the README to the reproducible numbers: ponytail
writes 80-94% less code, costs 47-77% less, and runs 3-6x faster than a
no-skill agent on every model. benchmarks/README.md carries the median tables
and the reproduce command. Drops nothing that is not measured.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-13 05:08:55 +02:00