feat: refine ruleset from a full-project field review (#39)
* feat: refine ruleset from a full-project field review
A reviewer ran ponytail across a 9-phase rewrite (protocol, PC app, simulator,
RPi daemon, ESP32 firmware) and flagged three gaps. All three land in SKILL.md
and propagate to AGENTS.md + the rule copies:
- Promote the one-runnable-check rule to a headline ("Lazy code without its
check is unfinished"), enforced as a check-rule-copies invariant.
- Hardware carve-out in "When NOT to be lazy": a real device is never the spec
ideal (clock drift, sensor offset), leave the calibration knob.
- Clarify the Output rule: explanation the user explicitly asked for is not
debt, only unrequested prose is.
Fallback instructions kept in sync. Rule-copy check + tests green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test: add a behavior gate proving the refinements actually fire
The refinements were verified as injected text, but injected != behavioral.
This adds a behavior eval that probes each refined rule on a task that should
trigger it:
- hardware -> does the output leave a calibration knob?
- explanation -> when a write-up is explicitly requested, is it given in full?
- onecheck -> is a runnable check left behind?
benchmarks/behavior.yaml runs the probes (baseline vs ponytail arm); the
grader benchmarks/behavior.js is proven by tests/behavior.test.js (8 cases,
RED/GREEN, no API key, runs in CI). Live-confirmed: the model under the
current ruleset passes all three gates, graded by the same grader.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
b545f1536a
commit
f3da910b4f
@@ -54,7 +54,9 @@ higher one and move on. The first lazy solution that works is the right one.
|
||||
Code first. Then at most three short lines: what was skipped, when to add it.
|
||||
No essays, no feature tours, no design notes. If the explanation is longer
|
||||
than the code, delete the explanation, every paragraph defending a
|
||||
simplification is complexity smuggled back in as prose.
|
||||
simplification is complexity smuggled back in as prose. Explanation the user
|
||||
explicitly asked for (a report, a walkthrough, per-phase notes) is not debt,
|
||||
give it in full, the rule is only against unrequested prose.
|
||||
|
||||
Pattern: `[code] → skipped: [X], add when [Y].`
|
||||
|
||||
@@ -78,11 +80,16 @@ that prevents data loss, security measures, accessibility basics, anything
|
||||
explicitly requested. User insists on the full version → build it, no
|
||||
re-arguing.
|
||||
|
||||
Non-trivial logic (a branch, a loop, a parser, a money/security path) leaves
|
||||
ONE runnable check behind, the smallest thing that fails if the logic
|
||||
breaks: an `assert`-based `demo()`/`__main__` self-check or one small
|
||||
`test_*.py`. No frameworks, no fixtures, no per-function suites unless
|
||||
asked. Trivial one-liners need no test, YAGNI applies to tests too.
|
||||
Hardware is never the ideal on paper: a real clock drifts, a real sensor
|
||||
reads off, a PCA9685 runs a few percent fast. Leave the calibration knob, not
|
||||
just less code, the physical world needs tuning a minimal model can't see.
|
||||
|
||||
Lazy code without its check is unfinished. Non-trivial logic (a branch, a
|
||||
loop, a parser, a money/security path) leaves ONE runnable check behind, the
|
||||
smallest thing that fails if the logic breaks: an `assert`-based
|
||||
`demo()`/`__main__` self-check or one small `test_*.py`. No frameworks, no
|
||||
fixtures, no per-function suites unless asked. Trivial one-liners need no
|
||||
test, YAGNI applies to tests too.
|
||||
|
||||
## Boundaries
|
||||
|
||||
|
||||
Reference in New Issue
Block a user