Skip to content
#

llm-evaluation-toolkit

Here are 27 public repositories matching this topic...

An easy python package to run quick basic QA evaluations. This package includes standardized QA evaluation metrics and semantic evaluation metrics: Black-box and Open-Source large language model prompting and evaluation, exact match, F1 Score, PEDANT semantic match, transformer match. Our package also supports prompting OPENAI and Anthropic API.

  • Updated Jul 18, 2025
  • Python

Core engine behind Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations

  • Updated Aug 12, 2026
  • JavaScript

Frontend for Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations

  • Updated Aug 13, 2026
  • TypeScript

An epistemic AI lab for scientifically evaluating how language models track belief, truth, knowledge, inference, defeaters, and luck.

  • Updated Aug 5, 2026
  • Python

Improve this page

Add a description, image, and links to the llm-evaluation-toolkit topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the llm-evaluation-toolkit topic, visit your repo's landing page and select "manage topics."

Learn more