Benchmark LLMs Like a Pro. In 5 Lines of Code.
Stop writing boilerplate to test LLMs. Start getting results.
A dead-simple Python library for benchmarking LLM providers. Write tests once, run them across any model, get beautiful reports.
benchmark = Benchmark(provider=client, name="my_test")
benchmark.add_test(TestCase(
prompt="What is 2+2?",
model="gpt-3.5-turbo",
validator=Contains("4")
))
report = await benchmark.run_async()That's it. No setup. No config files. Just results.
Before promptum:
# Custom API client for each provider
openai_client = OpenAI(api_key=...)
anthropic_client = Anthropic(api_key=...)
# Manual validation logic
if "correct answer" not in response:
failed_tests.append(...)
# Track metrics yourself
latency = end_time - start_time
tokens = response.usage.total_tokens
# Write your own retry logic
for attempt in range(max_retries):
try:
response = client.chat.completions.create(...)
break
except Exception:
sleep(2 ** attempt)After promptum:
report = await benchmark.run_async()
summary = report.get_summary() # Metrics captured automaticallypip install promptum # (or: uv pip install promptum)
export OPENROUTER_API_KEY="your-key"import asyncio
from promptum import Benchmark, TestCase, OpenRouterClient, Contains
async def main():
async with OpenRouterClient(api_key="your-key") as client:
benchmark = Benchmark(provider=client, name="quick_test")
benchmark.add_test(TestCase(
name="basic_math",
prompt="What is 15 * 7? Reply with just the number.",
model="openai/gpt-3.5-turbo",
validator=Contains("105")
))
report = await benchmark.run_async()
summary = report.get_summary()
print(f"✓ {summary['passed']}/{summary['total']} tests passed")
print(f"⚡ {summary['avg_latency_ms']:.0f}ms average")
print(f"💰 ${summary['total_cost_usd']:.6f} total cost")
asyncio.run(main())Run it:
python your_script.py- One API for 100+ Models - OpenRouter support out of the box (OpenAI, Anthropic, Google, etc.)
- Smart Validation - ExactMatch, Contains, Regex, JsonSchema, or write your own
- Automatic Retries - Exponential/linear backoff with configurable attempts
- Metrics Tracking - Latency, tokens, cost - automatically captured
- Async by Default - Run 100 tests in parallel without breaking a sweat
- Type Safe - Full type hints, catches errors before runtime
- Zero Config - No YAML files, no setup scripts, just Python
Compare GPT-4 vs Claude on your tasks:
from promptum import Benchmark, TestCase, ExactMatch, Contains, Regex
tests = [
TestCase(
name="json_output",
prompt='Output JSON: {"status": "ok"}',
model="openai/gpt-4",
validator=Regex(r'\{"status":\s*"ok"\}')
),
TestCase(
name="json_output",
prompt='Output JSON: {"status": "ok"}',
model="anthropic/claude-3-5-sonnet",
validator=Regex(r'\{"status":\s*"ok"\}')
),
TestCase(
name="creative_writing",
prompt="Write a haiku about Python",
model="openai/gpt-4",
validator=Contains("Python", case_sensitive=False)
),
]
benchmark.add_tests(tests)
report = await benchmark.run_async()
# Side-by-side model comparison
for model, summary in report.compare_models().items():
print(f"{model}: {summary['pass_rate']:.0%} pass rate, {summary['avg_latency_ms']:.0f}ms avg")🔬 Model Evaluation - Compare GPT-4, Claude, Gemini on your specific tasks 🎯 Prompt Engineering - Test 100 prompt variations, find what works ⚡ Latency Testing - Measure real-world response times across providers 💰 Cost Analysis - Track spending per model/task before production 🔄 Regression Testing - Ensure model updates don't break your prompts 📊 A/B Testing - Data-driven model selection for your product
- Python 3.13+
- An OpenRouter API key (or implement your own provider)
That's it. No Docker, no complex setup.
Most libraries force inheritance:
class MyProvider(BaseProvider): # Tightly coupled
def generate(self): ...We use protocols (structural typing):
class MyProvider: # No inheritance needed
async def generate(self) -> tuple[str, Metrics]:
# Your implementation
return response, metrics
# It just works
benchmark = Benchmark(provider=MyProvider())Cleaner. More flexible. More Pythonic.
Found a bug? Want a feature? PRs welcome!
# Development setup
git clone https://github.com/deyna256/promptum.git
cd promptum
just sync # Install dependencies
just test # Run tests
# Development commands
just lint # Check code style
just format # Format code
just typecheck # Type checkingMIT - do whatever you want with it.
⭐ Star on GitHub | 🐛 Report Bug | 💡 Request Feature
Made for developers who value their time.