Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
cf97ccc509 | ||
|
|
2d91f6a957 | ||
|
|
9c99843725 | ||
|
|
b609e019a0 | ||
|
|
983255e2a1 | ||
|
|
cbb8859f39 |
@@ -1,6 +1,7 @@
|
||||
{
|
||||
"name": "ponytail",
|
||||
"description": "Lazy senior dev mode. Forces the simplest, shortest solution that actually works — YAGNI, stdlib first, no unrequested abstractions.",
|
||||
"version": "4.0.0",
|
||||
"description": "Lazy senior dev mode. Forces the simplest, shortest solution that actually works: YAGNI, stdlib first, no unrequested abstractions.",
|
||||
"author": {
|
||||
"name": "Dietrich Gebert",
|
||||
"url": "https://github.com/DietrichGebert"
|
||||
|
||||
@@ -18,6 +18,7 @@ Rules:
|
||||
- No boilerplate nobody asked for.
|
||||
- Deletion over addition. Boring over clever. Fewest files possible.
|
||||
- Question complex requests: "Do you actually need X, or does Y cover it?"
|
||||
- Mark intentional simplifications with a `ponytail:` comment.
|
||||
- Pick the edge-case-correct option when two stdlib approaches are the same size — lazy means less code, not the flimsier algorithm.
|
||||
- Mark intentional simplifications with a `ponytail:` comment. If the shortcut has a known ceiling (global lock, O(n²) scan, naive heuristic), the comment names the ceiling and the upgrade path.
|
||||
|
||||
Not lazy about: input validation at trust boundaries, error handling that prevents data loss, security, accessibility, anything explicitly requested.
|
||||
Not lazy about: input validation at trust boundaries, error handling that prevents data loss, security, accessibility, anything explicitly requested. Non-trivial logic leaves ONE runnable check behind — the smallest thing that fails if the logic breaks (an assert-based demo/self-check or one small test file; no frameworks, no fixtures). Trivial one-liners need no test.
|
||||
|
||||
@@ -24,6 +24,7 @@ Rules:
|
||||
- No boilerplate nobody asked for.
|
||||
- Deletion over addition. Boring over clever. Fewest files possible.
|
||||
- Question complex requests: "Do you actually need X, or does Y cover it?"
|
||||
- Mark intentional simplifications with a `ponytail:` comment.
|
||||
- Pick the edge-case-correct option when two stdlib approaches are the same size — lazy means less code, not the flimsier algorithm.
|
||||
- Mark intentional simplifications with a `ponytail:` comment. If the shortcut has a known ceiling (global lock, O(n²) scan, naive heuristic), the comment names the ceiling and the upgrade path.
|
||||
|
||||
Not lazy about: input validation at trust boundaries, error handling that prevents data loss, security, accessibility, anything explicitly requested.
|
||||
Not lazy about: input validation at trust boundaries, error handling that prevents data loss, security, accessibility, anything explicitly requested. Non-trivial logic leaves ONE runnable check behind — the smallest thing that fails if the logic breaks (an assert-based demo/self-check or one small test file; no frameworks, no fixtures). Trivial one-liners need no test.
|
||||
|
||||
@@ -18,6 +18,7 @@ Rules:
|
||||
- No boilerplate nobody asked for.
|
||||
- Deletion over addition. Boring over clever. Fewest files possible.
|
||||
- Question complex requests: "Do you actually need X, or does Y cover it?"
|
||||
- Mark intentional simplifications with a `ponytail:` comment.
|
||||
- Pick the edge-case-correct option when two stdlib approaches are the same size — lazy means less code, not the flimsier algorithm.
|
||||
- Mark intentional simplifications with a `ponytail:` comment. If the shortcut has a known ceiling (global lock, O(n²) scan, naive heuristic), the comment names the ceiling and the upgrade path.
|
||||
|
||||
Not lazy about: input validation at trust boundaries, error handling that prevents data loss, security, accessibility, anything explicitly requested.
|
||||
Not lazy about: input validation at trust boundaries, error handling that prevents data loss, security, accessibility, anything explicitly requested. Non-trivial logic leaves ONE runnable check behind — the smallest thing that fails if the logic breaks (an assert-based demo/self-check or one small test file; no frameworks, no fixtures). Trivial one-liners need no test.
|
||||
|
||||
@@ -18,6 +18,7 @@ Rules:
|
||||
- No boilerplate nobody asked for.
|
||||
- Deletion over addition. Boring over clever. Fewest files possible.
|
||||
- Question complex requests: "Do you actually need X, or does Y cover it?"
|
||||
- Mark intentional simplifications with a `ponytail:` comment.
|
||||
- Pick the edge-case-correct option when two stdlib approaches are the same size — lazy means less code, not the flimsier algorithm.
|
||||
- Mark intentional simplifications with a `ponytail:` comment. If the shortcut has a known ceiling (global lock, O(n²) scan, naive heuristic), the comment names the ceiling and the upgrade path.
|
||||
|
||||
Not lazy about: input validation at trust boundaries, error handling that prevents data loss, security, accessibility, anything explicitly requested.
|
||||
Not lazy about: input validation at trust boundaries, error handling that prevents data loss, security, accessibility, anything explicitly requested. Non-trivial logic leaves ONE runnable check behind — the smallest thing that fails if the logic breaks (an assert-based demo/self-check or one small test file; no frameworks, no fixtures). Trivial one-liners need no test.
|
||||
|
||||
@@ -18,8 +18,9 @@ Rules:
|
||||
- No boilerplate nobody asked for.
|
||||
- Deletion over addition. Boring over clever. Fewest files possible.
|
||||
- Question complex requests: "Do you actually need X, or does Y cover it?"
|
||||
- Mark intentional simplifications with a `ponytail:` comment.
|
||||
- Pick the edge-case-correct option when two stdlib approaches are the same size — lazy means less code, not the flimsier algorithm.
|
||||
- Mark intentional simplifications with a `ponytail:` comment. If the shortcut has a known ceiling (global lock, O(n²) scan, naive heuristic), the comment names the ceiling and the upgrade path.
|
||||
|
||||
Not lazy about: input validation at trust boundaries, error handling that prevents data loss, security, accessibility, anything explicitly requested.
|
||||
Not lazy about: input validation at trust boundaries, error handling that prevents data loss, security, accessibility, anything explicitly requested. Non-trivial logic leaves ONE runnable check behind — the smallest thing that fails if the logic breaks (an assert-based demo/self-check or one small test file; no frameworks, no fixtures). Trivial one-liners need no test.
|
||||
|
||||
(Yes, this file also applies to agents working on the ponytail repo itself. Especially to them.)
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
<p align="center">
|
||||
<img src="assets/logo.png" width="220" alt="Ponytail — the lazy senior dev">
|
||||
<img src="assets/logo.png" width="220" alt="Ponytail, the lazy senior dev">
|
||||
</p>
|
||||
|
||||
<h1 align="center">Ponytail</h1>
|
||||
@@ -27,6 +27,16 @@ With ponytail:
|
||||
|
||||
More survivors in [examples/](examples/).
|
||||
|
||||
## Numbers
|
||||
|
||||
Six tasks: streaming log parser, atomic file sync, notification dispatcher, validation engine, auth module, concurrent money ledger. One spec each, one fresh agent per arm, same model. Three arms: no skill, the [caveman](https://github.com/JuliusBrussee/caveman) skill, and ponytail. Every arm passes the same adversarial security and concurrency probes. Then the agreement ends:
|
||||
|
||||
<p align="center">
|
||||
<img src="assets/benchmark-loc.svg" width="860" alt="Lines of code per task: ponytail 490 total vs caveman 1,440 vs no-skill control 3,629, all passing the same adversarial probes">
|
||||
</p>
|
||||
|
||||
**47% fewer tokens than the no-skill agent. 3× faster. A seventh of the code.** The 3,139 lines nobody wrote have never caused an incident. When a surprise feature request hit two of the tasks, ponytail extended in 96 changed lines; caveman needed 413, the no-skill agent 1,115. Every shortcut ponytail took is marked in the code with a `ponytail:` comment naming its upgrade path. Data: [benchmarks/](benchmarks/).
|
||||
|
||||
## How it works
|
||||
|
||||
Before writing code, the agent stops at the first rung that holds:
|
||||
@@ -63,7 +73,7 @@ Cursor, Windsurf, Cline, Copilot, Aider: copy the matching rules file from this
|
||||
No.
|
||||
|
||||
**What if I really need the 120-line cache class?**
|
||||
You don't. Insist anyway and he'll build it — slowly, correctly, while looking at you.
|
||||
You don't. Insist anyway and he'll build it. Slowly. Correctly. While looking at you.
|
||||
|
||||
**Does it scale?**
|
||||
The code you never wrote scales infinitely. Zero bugs, zero CVEs, 100% uptime since forever.
|
||||
@@ -71,10 +81,6 @@ The code you never wrote scales infinitely. Zero bugs, zero CVEs, 100% uptime si
|
||||
**Why "ponytail"?**
|
||||
You know exactly why.
|
||||
|
||||
## Numbers
|
||||
|
||||
5 coding tasks, same agent with and without ponytail: **−16% tokens, ~4× faster, 293 → 47 lines.** The 246 lines nobody wrote have never caused an incident. Data: [benchmarks/](benchmarks/).
|
||||
|
||||
## License
|
||||
|
||||
[MIT](LICENSE). The shortest license that works.
|
||||
|
||||
@@ -0,0 +1,69 @@
|
||||
<svg viewBox="0 0 860 470" xmlns="http://www.w3.org/2000/svg" font-family="-apple-system, 'Segoe UI', Helvetica, Arial, sans-serif">
|
||||
<title>Non-blank lines of code per task: control vs caveman vs ponytail</title>
|
||||
|
||||
<text x="20" y="26" font-size="15" font-weight="600" fill="#8b949e">Six tasks. Adversarial probes: everyone passes. Lines of code: not everyone.</text>
|
||||
|
||||
<!-- legend -->
|
||||
<rect x="20" y="42" width="12" height="12" rx="2" fill="#8b949e"/>
|
||||
<text x="38" y="53" font-size="13" fill="#8b949e">Control (no skill) · 3,629 total</text>
|
||||
<rect x="250" y="42" width="12" height="12" rx="2" fill="#d9822b"/>
|
||||
<text x="268" y="53" font-size="13" fill="#8b949e">Caveman · 1,440</text>
|
||||
<rect x="420" y="42" width="12" height="12" rx="2" fill="#2da44e"/>
|
||||
<text x="438" y="53" font-size="13" fill="#8b949e">Ponytail · 490</text>
|
||||
|
||||
<!-- Task A -->
|
||||
<text x="112" y="117" font-size="12" fill="#8b949e" text-anchor="end">log-analysis CLI</text>
|
||||
<rect x="120" y="90" width="662" height="13" rx="2" fill="#8b949e"/>
|
||||
<text x="788" y="101" font-size="11" fill="#8b949e">946</text>
|
||||
<rect x="120" y="106" width="198" height="13" rx="2" fill="#d9822b"/>
|
||||
<text x="324" y="117" font-size="11" fill="#d9822b">283</text>
|
||||
<rect x="120" y="122" width="102" height="13" rx="2" fill="#2da44e"/>
|
||||
<text x="228" y="133" font-size="11" fill="#2da44e" font-weight="600">145</text>
|
||||
|
||||
<!-- Task B -->
|
||||
<text x="112" y="177" font-size="12" fill="#8b949e" text-anchor="end">file sync</text>
|
||||
<rect x="120" y="150" width="459" height="13" rx="2" fill="#8b949e"/>
|
||||
<text x="585" y="161" font-size="11" fill="#8b949e">656</text>
|
||||
<rect x="120" y="166" width="160" height="13" rx="2" fill="#d9822b"/>
|
||||
<text x="286" y="177" font-size="11" fill="#d9822b">228</text>
|
||||
<rect x="120" y="182" width="69" height="13" rx="2" fill="#2da44e"/>
|
||||
<text x="195" y="193" font-size="11" fill="#2da44e" font-weight="600">99</text>
|
||||
|
||||
<!-- Task C -->
|
||||
<text x="112" y="237" font-size="12" fill="#8b949e" text-anchor="end">notification dispatcher</text>
|
||||
<rect x="120" y="210" width="566" height="13" rx="2" fill="#8b949e"/>
|
||||
<text x="692" y="221" font-size="11" fill="#8b949e">808</text>
|
||||
<rect x="120" y="226" width="277" height="13" rx="2" fill="#d9822b"/>
|
||||
<text x="403" y="237" font-size="11" fill="#d9822b">396</text>
|
||||
<rect x="120" y="242" width="51" height="13" rx="2" fill="#2da44e"/>
|
||||
<text x="177" y="253" font-size="11" fill="#2da44e" font-weight="600">73</text>
|
||||
|
||||
<!-- Task D -->
|
||||
<text x="112" y="297" font-size="12" fill="#8b949e" text-anchor="end">validation engine</text>
|
||||
<rect x="120" y="270" width="474" height="13" rx="2" fill="#8b949e"/>
|
||||
<text x="600" y="281" font-size="11" fill="#8b949e">677</text>
|
||||
<rect x="120" y="286" width="153" height="13" rx="2" fill="#d9822b"/>
|
||||
<text x="279" y="297" font-size="11" fill="#d9822b">218</text>
|
||||
<rect x="120" y="302" width="49" height="13" rx="2" fill="#2da44e"/>
|
||||
<text x="175" y="313" font-size="11" fill="#2da44e" font-weight="600">70</text>
|
||||
|
||||
<!-- Task E -->
|
||||
<text x="112" y="357" font-size="12" fill="#8b949e" text-anchor="end">auth module</text>
|
||||
<rect x="120" y="330" width="182" height="13" rx="2" fill="#8b949e"/>
|
||||
<text x="308" y="341" font-size="11" fill="#8b949e">260</text>
|
||||
<rect x="120" y="346" width="104" height="13" rx="2" fill="#d9822b"/>
|
||||
<text x="230" y="357" font-size="11" fill="#d9822b">148</text>
|
||||
<rect x="120" y="362" width="34" height="13" rx="2" fill="#2da44e"/>
|
||||
<text x="160" y="373" font-size="11" fill="#2da44e" font-weight="600">49</text>
|
||||
|
||||
<!-- Task F -->
|
||||
<text x="112" y="417" font-size="12" fill="#8b949e" text-anchor="end">money ledger</text>
|
||||
<rect x="120" y="390" width="197" height="13" rx="2" fill="#8b949e"/>
|
||||
<text x="323" y="401" font-size="11" fill="#8b949e">282</text>
|
||||
<rect x="120" y="406" width="117" height="13" rx="2" fill="#d9822b"/>
|
||||
<text x="243" y="417" font-size="11" fill="#d9822b">167</text>
|
||||
<rect x="120" y="422" width="38" height="13" rx="2" fill="#2da44e"/>
|
||||
<text x="164" y="433" font-size="11" fill="#2da44e" font-weight="600">54</text>
|
||||
|
||||
<text x="120" y="458" font-size="11" fill="#8b949e" opacity="0.8">Non-blank LOC, AST-counted. Same model, same specs, one fresh agent per arm. Every arm passes the same security (8/8) and concurrency (6/6) probes. 2026-06-12.</text>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 4.2 KiB |
@@ -0,0 +1,125 @@
|
||||
# Ponytail v4 hardening — A–F benchmark vs Caveman (2026-06-12)
|
||||
|
||||
Response to the hardening brief in `C:\dev\ponytail-bench\PONYTAIL-BENCHMARK-WRITEUP.md`.
|
||||
Harness reused as-is: same 6 tasks (specs reconstructed in `ponytail-bench\specs.md` —
|
||||
the originals were not preserved; both new arms got identical text), same scorer
|
||||
(`score.py`, arms now auto-discovered), same adversarial probes (`probe_e.py`,
|
||||
`probe_f.py`), same extension protocol (phase-1 git commit, cost = `git diff
|
||||
--numstat` insertions + new-file LOC). Caveman = `JuliusBrussee/caveman` SKILL.md
|
||||
verbatim (full level), saved at `ponytail-bench\caveman-SKILL.md`. One fresh
|
||||
subagent per task × arm, same model for all 16 runs. Caveat: this model/harness
|
||||
differs from the original Cursor runs, so comparisons to the old treatment
|
||||
numbers are directional; the ponytail4-vs-caveman head-to-head is same-model.
|
||||
|
||||
## v4 changes (the hardening, ~10 lines of prompt total)
|
||||
|
||||
1. **Test reflex** (brief 5.1): non-trivial logic leaves ONE runnable check —
|
||||
assert-based `demo()`/`__main__` self-check or one small `test_*.py`. No
|
||||
frameworks. One-liners need no test.
|
||||
2. **Ceiling comments** (5.2): a `ponytail:` shortcut with a known ceiling must
|
||||
name the ceiling and the upgrade path in the comment.
|
||||
3. **Robust variant rule** (5.3): between two same-size stdlib options, take the
|
||||
edge-case-correct one.
|
||||
|
||||
Applied to SKILL.md, all five cross-agent rule copies, the hook fallback, and a
|
||||
guard line in ponytail-review (never flag the minimal check as bloat).
|
||||
|
||||
## Build phase — non-blank LOC / .py files (scorer-verified)
|
||||
|
||||
| Task | Control (orig) | Treatment v3 (orig) | **Ponytail v4** | **Caveman** |
|
||||
|---|--:|--:|--:|--:|
|
||||
| A log CLI | 970 / 13 | 150 / 1 | **145 / 2** | 283 / 1 |
|
||||
| B file sync | 587 / 9 | 175 / 2 | **99 / 1** | 228 / 2 |
|
||||
| C dispatcher | 726 / 13 | 85 / 1 | **73 / 1** | 396 / 10 |
|
||||
| D validation | 343 / 8 | 93 / 1 | **70 / 1** | 218 / 3 |
|
||||
| E auth | 155 / 1 | 74 / 1 | **49 / 1** | 148 / 1 |
|
||||
| F ledger | 162 / 1 | 86 / 1 | **54 / 1** | 167 / 1 |
|
||||
| **Total** | 2943 | 663 | **490** | 1440 |
|
||||
|
||||
v4 is at or below v3 on every task (−3% to −43%) **despite now shipping a
|
||||
runnable check in all six arms** — the test reflex did not cause bloat creep.
|
||||
v4 is 34% of Caveman's size. Task A is the one place Caveman has fewer .py
|
||||
files (1 vs 2): v4's second file is the 24-line regression check Caveman
|
||||
doesn't ship — deleting it to win file count would sacrifice the safety clause
|
||||
to win on size, which the brief forbids.
|
||||
|
||||
## Extension phase (tasks C, D — surprise requests, git-measured)
|
||||
|
||||
| Metric | C v4 | C caveman | D v4 | D caveman |
|
||||
|---|--:|--:|--:|--:|
|
||||
| Lines changed (insertions + new-file LOC) | **41** | 156 | **55** | 257 |
|
||||
| Files touched | 1 | 7 | 1 | 3 |
|
||||
| Still correct after | yes | yes | yes | yes |
|
||||
|
||||
v4 honored the requested seams (duck-typed registry in C, `@rule` registry in
|
||||
D) and extended 74–79% cheaper. Both arms' extended demos re-run exit 0.
|
||||
|
||||
## Safety — adversarial probes (independently executed)
|
||||
|
||||
| Probe | v4 | caveman |
|
||||
|---|--:|--:|
|
||||
| Security, task E (8 checks) | **8/8** | 8/8 |
|
||||
| Concurrency, task F (6 checks) | **6/6** | 6/6 |
|
||||
|
||||
No regression from the added rules. v4's E chose PBKDF2-HMAC-SHA256 (600k
|
||||
iters) + 16-byte `secrets` salt + `hmac.compare_digest` + `token_urlsafe(32)`;
|
||||
F kept integer cents + a global lock with the ceiling comment naming the
|
||||
per-account-lock upgrade (5.2 working as designed; Caveman built per-account
|
||||
locks at 3× the LOC).
|
||||
|
||||
## Correctness
|
||||
|
||||
19/19 independent re-runs exit 0 (14 build demos/tests + 5 post-extension).
|
||||
|
||||
## Acceptance criteria (brief §5.6)
|
||||
|
||||
1. Probes 100% — **pass** (8/8 + 6/6).
|
||||
2. Every treatment arm ships a runnable check — **pass** (A: `test_loganalyze.py`;
|
||||
B–F: assert-based `__main__` checks; all executed). This was the #1 gap (was 1/4).
|
||||
3. LOC within ~20% of v3 treatment numbers — **pass on intent**: every arm at or
|
||||
below v3 (A −3%, C −14%; B/D/E/F 25–43% *below* — leaner, not bloated).
|
||||
4. Ceiling-bearing `ponytail:` comments name upgrade paths — **pass**, verified
|
||||
per arm: global lock→per-account locks (F), no token TTL→add TTL (E),
|
||||
sequential sends→async/threaded + hardcoded route→routing table (C),
|
||||
special-cased `unique`→DATASET_RULES registry (D), observed-hours stats→
|
||||
impute full range (A), no empty-dir handling→dir pass (B).
|
||||
5. Head-to-head vs Caveman — **pass**: ≥ on every axis, strictly better on three.
|
||||
- Safety: tie at 100% (≥, never regressed to win on size).
|
||||
- Size: LOC strictly better 6/6; files ≤ on 5/6 (A caveat above).
|
||||
- Extension cost: strictly better on both tasks.
|
||||
- Reviewability: strictly better — every v4 simplification is `ponytail:`-marked
|
||||
with its ceiling; Caveman's code marks only spec-allowed simulated transports,
|
||||
and its design trade-offs live in the chat report, invisible to a later reviewer.
|
||||
|
||||
## Addendum: same-model control arm (control2, added same day)
|
||||
|
||||
The control numbers above were inherited from the original Cursor harness,
|
||||
which could not expose token counts. Six fresh `task*-control2` arms were run
|
||||
through this harness (no skill, "build production-normal", same model, same
|
||||
specs), making all three arms same-model. Control2 passes both probes (8/8,
|
||||
6/6) and all 10 demo/test runs exit 0; extensions on C and D re-verified.
|
||||
|
||||
| Whole benchmark (6 builds + C/D extensions) | Control2 | Caveman | Ponytail v4 |
|
||||
|---|--:|--:|--:|
|
||||
| Build LOC | 3,629 | 1,440 | **490** |
|
||||
| Build LOC per task (A-F) | 946/656/808/677/260/282 | 283/228/396/218/148/167 | **145/99/73/70/49/54** |
|
||||
| Extension lines changed (C, D) | 378, 737 | 156, 257 | **41, 55** |
|
||||
| Agent tokens, total | 430,697 | 290,546 | **229,370 (-47% vs control2)** |
|
||||
| Agent wall time, total | 2,749s | 1,596s | **821s (3.3x)** |
|
||||
| Probes | 8/8 + 6/6 | 8/8 + 6/6 | 8/8 + 6/6 |
|
||||
|
||||
Wall times carry parallel-scheduling noise (arms ran concurrently, n=1 per
|
||||
cell); token counts are exact from agent telemetry. The README "Numbers"
|
||||
section now cites this same-model dataset and retires the older 5-task v3
|
||||
figures (still recorded in `2026-06-12-caveman-vs-ponytail.md`).
|
||||
|
||||
## Residual (honest notes)
|
||||
|
||||
- A's spike stats still use observed-hours-only mean+3σ rather than a
|
||||
leave-one-out/imputed baseline (Caveman zero-filled the hour range). The 5.3
|
||||
rule softened but did not eliminate the naive-algorithm tendency; the choice
|
||||
is now at least documented with its upgrade path (5.2). Candidate for a
|
||||
future eval if it bites in practice.
|
||||
- Caveman is a prose-compression skill that explicitly writes code "normal" —
|
||||
it loses on code size by design. The meaningful result is that adding the
|
||||
test reflex did not erode ponytail's size advantage or its 100% probe record.
|
||||
@@ -110,13 +110,15 @@ if (skillContent) {
|
||||
'No abstractions that were not requested. No avoidable dependencies. No boilerplate nobody asked for. ' +
|
||||
'Deletion over addition. Boring over clever. Fewest files possible. ' +
|
||||
'Ship the lazy version and question the complex request in the same response — never stall. ' +
|
||||
'Mark intentional simplifications with a `ponytail:` comment.\n\n' +
|
||||
'Between two same-size stdlib options, pick the one correct on edge cases. ' +
|
||||
'Mark intentional simplifications with a `ponytail:` comment — a shortcut with a known ceiling names the ceiling and the upgrade path in the comment.\n\n' +
|
||||
'## Output\n\n' +
|
||||
'Code first. Then at most three short lines: what was skipped, when to add it. ' +
|
||||
'If the explanation is longer than the code, delete the explanation.\n\n' +
|
||||
'## When NOT to be lazy\n\n' +
|
||||
'Never simplify away: input validation at trust boundaries, error handling that prevents data loss, ' +
|
||||
'security measures, accessibility basics, anything the user explicitly asked to keep.\n\n' +
|
||||
'security measures, accessibility basics, anything the user explicitly asked to keep. ' +
|
||||
'Non-trivial logic leaves ONE runnable check behind (assert-based demo/self-check or one small test file; no frameworks). Trivial one-liners need no test.\n\n' +
|
||||
'## Boundaries\n\n' +
|
||||
'Ponytail governs what you build, not how you talk. "stop ponytail" or "normal mode": revert. Level persists until changed or session end.';
|
||||
}
|
||||
|
||||
@@ -50,5 +50,7 @@ If there is nothing to cut, say `Lean already. Ship.` and stop.
|
||||
## Boundaries
|
||||
|
||||
Complexity only — correctness bugs, security holes, and performance go to a
|
||||
normal review pass, not this one. Does not apply the fixes, only lists them.
|
||||
normal review pass, not this one. A single smoke test or `assert`-based
|
||||
self-check is the ponytail minimum, not bloat — never flag it for deletion.
|
||||
Does not apply the fixes, only lists them.
|
||||
"stop ponytail-review" or "normal mode": revert to verbose review style.
|
||||
|
||||
@@ -46,7 +46,8 @@ higher one and move on. The first lazy solution that works is the right one.
|
||||
- Deletion over addition. Boring over clever — clever is what someone decodes at 3am.
|
||||
- Fewest files possible. Shortest working diff wins.
|
||||
- Complex request? Ship the lazy version and question it in the same response — "Did X; Y covers it. Need full X? Say so." Never stall on an answer you can default.
|
||||
- Mark deliberate simplifications with a `ponytail:` comment (`// ponytail: this exists`) — simple reads as intent, not ignorance.
|
||||
- Two stdlib options, same size? Take the one that's correct on edge cases. Lazy means writing less code, not picking the flimsier algorithm.
|
||||
- Mark deliberate simplifications with a `ponytail:` comment (`// ponytail: this exists`) — simple reads as intent, not ignorance. Shortcut with a known ceiling (global lock, O(n²) scan, naive heuristic)? The comment names the ceiling and the upgrade path: `# ponytail: global lock — per-account locks if throughput matters`.
|
||||
|
||||
## Output
|
||||
|
||||
@@ -77,6 +78,12 @@ that prevents data loss, security measures, accessibility basics, anything
|
||||
explicitly requested. User insists on the full version → build it, no
|
||||
re-arguing.
|
||||
|
||||
Non-trivial logic (a branch, a loop, a parser, a money/security path) leaves
|
||||
ONE runnable check behind — the smallest thing that fails if the logic
|
||||
breaks: an `assert`-based `demo()`/`__main__` self-check or one small
|
||||
`test_*.py`. No frameworks, no fixtures, no per-function suites unless
|
||||
asked. Trivial one-liners need no test — YAGNI applies to tests too.
|
||||
|
||||
## Boundaries
|
||||
|
||||
Ponytail governs what you build, not how you talk (pair with Caveman for
|
||||
|
||||
Reference in New Issue
Block a user