Production-ready automated evaluation pipeline for LLM prompts, prompt regressions, RAG systems, and autonomous agent workflows.
-
Updated
Sep 4, 2026 - Python
Production-ready automated evaluation pipeline for LLM prompts, prompt regressions, RAG systems, and autonomous agent workflows.
Production-grade evaluation framework for LLMs, RAG pipelines, and autonomous AI agents with deterministic checks, LLM-as-a-Judge, RAG Triad, and an interactive observatory dashboard.
GenAI Chatbot Testing Framework using Python & Pytest | Validates intent, context, prompt variations, hallucination risks, safety, and response quality scoring
GenAITesting — training and certification in GenAI, LLM and AI agent application testing, with hands-on RAG and multi-agent projects, plus a Python and DSA track.
To associate your repository with the genai-testing topic, visit your repo's landing page and select "manage topics."