Pinned Loading
-
trailhead-travel-agent-eval
trailhead-travel-agent-eval PublicLLM-agent quality suite for a fictional travel-support bot (single-turn Q&A, multi-turn chatbot, RAG) using DeepEval — correctness, safety, and hallucination metrics scored by a local judge model.
Python
-
trailhead-travel-rag-eval
trailhead-travel-rag-eval PublicRAG-pipeline evaluation for a fictional travel-support agent using Ragas — faithfulness, context precision/recall, and answer relevancy, with noise-aware CI thresholds.
Python
-
trailhead-travel-observability
trailhead-travel-observability PublicObservability and CI regression gating for an LLM agent using LangSmith — tracing, versioned eval datasets, a GitHub Actions prompt-regression gate, and production drift monitoring.
Python
-
trailhead-travel-red-team
trailhead-travel-red-team PublicAutomated red-teaming and safety layering for a RAG agent — Promptfoo attack generation, Guardrails AI validation, and Llama Guard classification, with before/after break-rate comparison.
Python
If the problem persists, check the GitHub status page or contact support.