ai-evaluation
Open-source LLM and agent evaluation framework with 50+ metrics, LLM-as-Judge augmentation, and guardrail scanners (jailbreak, PII, prompt-injection). Useful for scoring RAG outputs, agent trajectories, and function-calling behavior in data-science workflows.
- from
- Data Science
- added
- 2026-10-10
- likes
- 0
similar
-
OpenAI Evals github.com
An open-source framework and registry for evaluating language models and systems.
-
Opik github.com
An open-source platform for tracing, evaluating, and monitoring LLM applications. #opensource
-
crilio github.com
An open-source Python CLI that uses LLM-as-a-Judge to automate semantic regression testing for LLM prompts in CI/CD, blocking GitHub PRs that cause hallucinations or break formatting rules. Supports OpenAI, Anthropic, and local Ollama models.
-
Where the LLM Stops: Deterministic Scoring in an AI-Assisted VAPT Pipeline aayushyadav.hashnode.dev
-
NLG-eval github.com
Evaluation code for various unsupervised automated metrics for Natural Language Generation.
-
CALM: Curiosity-Driven Auditing for Large Language Models arxiv.org
(AAAI) Auditing as a black-box optimization problem where the goal is to automatically uncover input-output pairs of the target LLMs that exhibit illegal, immoral, or unsafe behaviors.
Data Science › Agents > Tools: “Open-source LLM and agent evaluation framework with 50+ metrics, LLM-as-Judge augmentation, and guardrail scanners (jailbreak, PII, prompt-injection). Useful for scoring RAG outputs, agent trajectories, and function-calling behavior in data-science workflows.”