Files
ponytail/examples/rate-limit.md
T
DietrichGebertandClaude Opus 4.8 45f7d2f83f Fix/examples issue 127 (#131)
* docs: correct cost claim to 42-75% from 30-rep re-verification

Re-ran the cost benchmark at 30 reps per cell on Claude (Haiku/Sonnet/Opus):
ponytail is 42-75% cheaper than no-skill, not the previously published 47-77%.
The direction holds, both ends came in a few points lower. Updates the README
headline and body, the benchmark chart subtitle, and the benchmarks/README cost
table, and adds a dated results doc with full method.

Also adds the OpenAI (gpt-4.1-mini/gpt-5.4-mini/gpt-5.5) and Gemini configs. On
OpenAI reasoning models ponytail costs more, not less, so the claim stays
Claude-scoped. Gemini run pending a fresh-quota day.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: scope the body claim to Claude models

"on every model" read as cross-provider, but the 30-rep verification shows
the cost win reverses on OpenAI reasoning models. Match the caption and
benchmarks/README, which already say Claude.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: reframe the pitch as the discipline, not token savings

The cost/code/latency numbers vary by model and on some (terse reasoning
models like GPT-5.5) ponytail costs more, so leading with them as a universal
win was misleading. Adds model-variance to the headline caption and a paragraph
making the stated point the mental model: write only what the task needs,
safety kept, maintainable code. Savings are a model-dependent side effect.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: name the ladder's reasoning cost

The ladder is a deliberation step: on reasoning models the agent spends
thinking tokens working through the rungs before it saves any output, which
together with the always-on ruleset can outweigh the shorter code. Makes the
GPT-5.5 cost increase legible rather than just stating it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: state the single-shot limitation honestly

The benchmark is single-shot (one prompt, one completion); it does not measure
a real multi-turn agent session, where the ruleset re-injects and the ladder
deliberates every turn. Adds that caveat to the README, and corrects the
benchmarks/README note that claimed caching widens the gap "in ponytail's
favor" (unverified, and a measured agentic A/B in #121 found the opposite can
happen). Per-session cost can land either way.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: fix run count in caption (cost is 30 runs, not 10)

Cost was re-verified at 30 reps; code and latency are still the original 10.
The headline caption said "10 runs" across the board, which undersold the cost
verification. Now states the split, matching benchmarks/README.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(examples): replace hand-written examples with real benchmark output

The examples/ before/after blocks were authored by hand, not produced by a
model. Issue #127 correctly noted that nobody hand-rolls quicksort for "sort
this array" - every model just calls .sort(). Regenerate all examples verbatim
from a real benchmark run (Claude Haiku 4.5, no-skill arm vs ponytail arm,
benchmarks/output.json) so the before/after is reproducible, not authored:

  email 75->3, debounce 116->10, csv 20->3, countdown 267->9, rate-limit 128->10 LOC

- Delete sorting.md (pure strawman) plus the other hand-written caricatures
  (api-endpoint, caching, date-picker)
- Add benchmarks/generate-examples.mjs to regenerate examples from any run
- examples/README.md indexes the set and documents how to reproduce

Closes #127

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 04:35:22 +02:00

6.9 KiB

Rate Limiting

Task: "Add rate limiting to my FastAPI endpoint so users can't spam it."

Verbatim model output from a benchmark run — Claude Haiku 4.5, no-skill arm vs ponytail arm, temperature 1, source benchmarks/output.json. Reproduce: npx promptfoo@latest eval -c benchmarks/promptfooconfig.yaml.

Without Ponytail — 128 lines of code

Rate Limiting in FastAPI

Here are several approaches, from simplest to most robust:

pip install slowapi
from fastapi import FastAPI, Request
from slowapi import Limiter
from slowapi.util import get_remote_address
from slowapi.errors import RateLimitExceeded
from fastapi.responses import JSONResponse

app = FastAPI()

limiter = Limiter(key_func=get_remote_address)
app.state.limiter = limiter
app.add_exception_handler(RateLimitExceeded, lambda request, exc: JSONResponse(
    status_code=429,
    content={"detail": "Rate limit exceeded"}
))

@app.get("/api/data")
@limiter.limit("5/minute")
async def get_data(request: Request):
    return {"message": "Success"}

Rate limit formats:

  • "5/minute" - 5 requests per minute
  • "100/hour" - 100 requests per hour
  • "1000/day" - 1000 requests per day

2. Using limits Library (More Control)

pip install limits
from fastapi import FastAPI, HTTPException, Request
from limits import parse
from limits.storage import MemoryStorage
from limits.strategies import MovingWindowRateLimiter

app = FastAPI()

storage = MemoryStorage()
limiter = MovingWindowRateLimiter(storage)
rate_limit = parse("5/minute")

@app.get("/api/data")
async def get_data(request: Request):
    client_id = request.client.host
    
    if not limiter.hit(rate_limit, client_id):
        raise HTTPException(status_code=429, detail="Rate limit exceeded")
    
    return {"message": "Success"}

3. Custom Middleware (Full Control)

from fastapi import FastAPI, HTTPException, Request
from datetime import datetime, timedelta
from collections import defaultdict
import asyncio

app = FastAPI()

# Store request timestamps per IP
request_history = defaultdict(list)
RATE_LIMIT = 5  # requests
TIME_WINDOW = 60  # seconds

@app.middleware("http")
async def rate_limit_middleware(request: Request, call_next):
    client_ip = request.client.host
    now = datetime.now()
    
    # Clean old requests outside time window
    request_history[client_ip] = [
        req_time for req_time in request_history[client_ip]
        if now - req_time < timedelta(seconds=TIME_WINDOW)
    ]
    
    # Check if limit exceeded
    if len(request_history[client_ip]) >= RATE_LIMIT:
        return JSONResponse(
            status_code=429,
            content={"detail": "Rate limit exceeded"}
        )
    
    # Record this request
    request_history[client_ip].append(now)
    
    return await call_next(request)

@app.get("/api/data")
async def get_data():
    return {"message": "Success"}

4. Per-User Rate Limiting (With Authentication)

from fastapi import FastAPI, Depends, HTTPException
from slowapi import Limiter
from slowapi.util import get_remote_address

app = FastAPI()
limiter = Limiter(key_func=get_remote_address)

def get_user_id(token: str = Header(None)) -> str:
    # Your auth logic here
    return token or "anonymous"

@app.get("/api/data")
@limiter.limit("10/minute")
async def get_data(request: Request, user_id: str = Depends(get_user_id)):
    return {"message": f"Success for {user_id}"}

5. Redis-Based Rate Limiting (Production)

pip install slowapi redis
from fastapi import FastAPI, Request
from slowapi import Limiter
from slowapi.util import get_remote_address
from slowapi.errors import RateLimitExceeded
from slowapi.storage import RedisStorage
from redis import Redis
from fastapi.responses import JSONResponse

redis_client = Redis.from_url("redis://localhost:6379")
storage = RedisStorage(redis_client)
limiter = Limiter(key_func=get_remote_address, storage=storage)

app = FastAPI()
app.state.limiter = limiter

app.add_exception_handler(
    RateLimitExceeded,
    lambda request, exc: JSONResponse(
        status_code=429,
        content={"detail": "Rate limit exceeded"}
    )
)

@app.get("/api/data")
@limiter.limit("5/minute")
async def get_data(request: Request):
    return {"message": "Success"}

6. Complete Example with Multiple Endpoints

from fastapi import FastAPI, Request
from slowapi import Limiter
from slowapi.util import get_remote_address
from slowapi.errors import RateLimitExceeded
from fastapi.responses import JSONResponse

app = FastAPI()
limiter = Limiter(key_func=get_remote_address)
app.state.limiter = limiter

@app.exception_handler(RateLimitExceeded)
async def rate_limit_handler(request: Request, exc: RateLimitExceeded):
    return JSONResponse(
        status_code=429,
        content={
            "detail": "Too many requests",
            "retry_after": exc.detail
        }
    )

# Strict limit
@app.post("/api/login")
@limiter.limit("3/minute")
async def login(request: Request):
    return {"token": "xxx"}

# Moderate limit
@app.get("/api/data")
@limiter.limit("30/minute")
async def get_data(request: Request):
    return {"data": []}

# Generous limit
@app.get("/api/public")
@limiter.limit("100/minute")
async def get_public(request: Request):
    return {"public": "data"}

Comparison Table

Method Pros Cons Best For
slowapi Easy, decorator-based In-memory only Development, small apps
limits Flexible, multiple backends More setup Custom logic
Middleware Full control, no dependencies Manual cleanup Learning, simple cases
Redis Distributed, scalable Requires Redis Production, multi-server

Testing Rate Limits

import httpx

async def test_rate_limit():
    async with httpx.AsyncClient(app=app, base_url="http://test") as client:
        for i in range(6):
            response = await client.get("/api/data")
            print(f"Request {i+1}: {response.status_code}")
            # First 5 should be 200, 6th should be 429

Recommendation: Use slowapi for most cases—it's the most popular and easiest to implement! 🚀

With Ponytail — 10 lines of code

from fastapi import FastAPI, HTTPException
from slowapi import Limiter
from slowapi.util import get_remote_address

app = FastAPI()
limiter = Limiter(key_func=get_remote_address)
app.state.limiter = limiter

@app.get("/api/endpoint")
@limiter.limit("10/minute")
async def my_endpoint(request):
    return {"status": "ok"}

Skipped: custom rate limit logic, Redis, sliding windows — slowapi handles it. Add when: you need distributed rate limiting across multiple servers (swap Limiter for Redis backend) or per-user limits (add key_func=lambda r: r.headers.get("authorization")).

128 → 10 lines of code — same model, same prompt.