You're choosing between GPT-4, Claude, and Gemini for your product. You need to know which one actually handles your prompts better — not some generic benchmark, but your real tasks, your edge cases, your expected formats.
promptum turns that into a few lines of Python. One client for 100+ models, automatic retries, latency/cost/token tracking, structured validation — all async, all typed, zero config files.
session = Session(provider=client, name="my_test")
session.add_test(Prompt(
name="basic_math",
prompt="What is 2+2?",
model="openai/gpt-3.5-turbo",
validator=Contains("4"),
))
report = await session.run()
summary = report.get_summary()No YAML. No config files. Just Python you can read in 30 seconds.
pip install promptum # (or: uv pip install promptum)
export OPENROUTER_API_KEY="your-key"import asyncio
from promptum import Session, Prompt, OpenRouterClient, Contains
async def main():
async with OpenRouterClient(api_key="your-key") as client:
session = Session(provider=client, name="quick_test")
session.add_test(Prompt(
name="basic_math",
prompt="What is 15 * 7? Reply with just the number.",
model="openai/gpt-3.5-turbo",
validator=Contains("105"),
))
report = await session.run()
summary = report.get_summary()
print(f"Passed: {summary.passed}/{summary.total}")
print(f"Avg latency: {summary.avg_latency_ms:.0f}ms")
print(f"Total cost: ${summary.total_cost_usd:.6f}")
asyncio.run(main())Most LLM testing is ad-hoc scripts that grow into unmaintainable messes. You end up with separate API clients per provider, hand-rolled retry logic, manual latency tracking, and validation scattered across files.
promptum replaces all of that with a single coherent API:
- 100+ Models via OpenRouter — one client for OpenAI, Anthropic, Google, and more
- Smart Validation — ExactMatch, Contains, Regex, JsonSchema, or write your own
- Automatic Retries — exponential/fixed-delay backoff with configurable attempts
- Metrics Tracking — latency, tokens, cost — automatically captured
- Async by Default — run tests in parallel with concurrency control
- Extensible — implement
LLMProviderorValidatorto plug in any model or validation logic - Type Safe — full type hints, catches errors before runtime
- Session & Testing — Session, Prompt, Report, Summary, TestResult
- Providers — LLMProvider, OpenRouterClient, Metrics, Retry, Exceptions
- Validation — Validator, ExactMatch, Contains, Regex, JsonSchema
- Python 3.10+
- An OpenRouter API key (or implement your own provider)
Found a bug? Want a feature? PRs welcome!
git clone https://github.com/deyna256/promptum.git
cd promptum
just sync # Install dependencies
just test # Run tests
just style # Check code style
just format # Format code
just type # Type checkingMIT - do whatever you want with it.
Star on GitHub | Report Bug | Request Feature
Made for developers who value their time.