Threshold testing is blind to agent regression. agent-eval runs your agent 50x on version A and B and gives you a p-value on whether behavior actually shifted, not just whether one run looked different.
statistics ci-cd python3 p-value regression-testing ai-evaluation langchain llm-agent llm-evaluation crewai langgraph llm-observability llm-testing agent-ops openai-agents-sdk agent-testing agent-eval mann-whitney-u bootstrap-confidence-interval promptfoo-alternative
-
Updated
Aug 3, 2026 - Python