dogear

enter for all results · esc to close

ai-evaluation

github.comtool

Open-source LLM and agent evaluation framework with 50+ metrics, LLM-as-Judge augmentation, and guardrail scanners (jailbreak, PII, prompt-injection). Useful for scoring RAG outputs, agent trajectories, and function-calling behavior in data-science workflows.

from
Data Science
added
2026-10-10
likes
0

similar

  1. OpenAI Evals github.com

    An open-source framework and registry for evaluating language models and systems.

  2. Opik github.com

    An open-source platform for tracing, evaluating, and monitoring LLM applications. #opensource

  3. crilio github.com

    An open-source Python CLI that uses LLM-as-a-Judge to automate semantic regression testing for LLM prompts in CI/CD, blocking GitHub PRs that cause hallucinations or break formatting rules. Supports OpenAI, Anthropic, and local Ollama models.

  4. Where the LLM Stops: Deterministic Scoring in an AI-Assisted VAPT Pipeline aayushyadav.hashnode.dev
  5. NLG-eval github.com

    Evaluation code for various unsupervised automated metrics for Natural Language Generation.

  6. CALM: Curiosity-Driven Auditing for Large Language Models arxiv.org

    (AAAI) Auditing as a black-box optimization problem where the goal is to automatically uncover input-output pairs of the target LLMs that exhibit illegal, immoral, or unsafe behaviors.

Data Science › Agents > Tools: “Open-source LLM and agent evaluation framework with 50+ metrics, LLM-as-Judge augmentation, and guardrail scanners (jailbreak, PII, prompt-injection). Useful for scoring RAG outputs, agent trajectories, and function-calling behavior in data-science workflows.”