EVAL-FIRST TOOLKIT FOR LLM APPLICATIONS

Measure the intelligence. Improve the outcome.

EvalCore turns "the model seems to work" into evidence: structured metrics, test datasets, and quality signals your team can act on — built into the engineering workflow, not bolted on at the end.

$ pip install evalcore
Read the docs
Apache-2.0 license Python toolkit 2 commits in, actively shaping up
● eval run — summary_accuracy live
faithfulness
0.94
context relevance
0.88
answer accuracy
0.97
01 — What it is

A developer-focused evaluation toolkit

EvalCore is built for the moments where a demo looks convincing but you still can't say, with evidence, whether it's good. It gives every project the same short list of questions to answer before shipping a change.

  • Is the response accurate?
  • Is the retrieved context relevant?
  • Does the answer stay grounded in the available information?
  • How does a new model or prompt compare with the previous one?
  • Where does the application need improvement?

Evaluation isn't a final inspection. It's part of the engineering workflow — build the application, measure its behavior, improve what matters.

02 — The loop

Uncertainty, turned into evidence

LLM applications can produce convincing answers without consistently producing correct ones. EvalCore is built around a closed feedback loop that catches the difference.

LLM application Test dataset Evaluation Quality signals Engineering decisions prompts · retrieval · model choice continuous improvement
03 — Capabilities

Five things EvalCore is built to do

Each capability targets one part of the loop above — from generating test data to feeding results back into the next iteration.

01 Objective metrics Evaluate application outputs against measurable quality signals, not gut feel.
02 Test generation Build evaluation datasets that cover real scenarios and edge cases.
03 Quality analysis Surface weak points across retrieval, generation, and end-to-end responses.
04 Framework integrations Plug evaluation into the tools already used to build your LLM application.
05 Feedback loops Route evaluation results back into prompts, retrieval, and model decisions.
04 — Quick start

Define a metric, score a response

The example below defines a custom accuracy check and scores a single application output — the smallest complete evaluation loop.

quickstart.py python
import asyncio

from openai import AsyncOpenAI
from ragas.metrics import DiscreteMetric
from ragas.llms import llm_factory

# Configure the evaluation model
client = AsyncOpenAI()
llm = llm_factory("gpt-4o", client=client)

# Define what you want to measure
metric = DiscreteMetric(
    name="summary_accuracy",
    allowed_values=["accurate", "inaccurate"],
    prompt="Evaluate if the summary is accurate."
)

# Evaluate an application output
async def main():
    score = await metric.ascore(
        llm=llm,
        response="The summary of the text is..."
    )
    print(f"Score: {score.value}")
    print(f"Reason: {score.reason}")

asyncio.run(main())
05 — Workflow

A typical EvalCore project

Five stages, repeated. The goal is never a single score — it's a system that keeps getting better.

Define the quality signal

Decide what matters — accuracy, relevance, faithfulness, completeness, or a custom domain-specific criterion.

Prepare evaluation data

Pull representative examples from your application, including the cases where quality matters most.

Run evaluations

Apply metrics to your application's outputs and collect structured, comparable results.

Inspect the results

Look for patterns, recurring failure cases, and shifts in performance over time.

Improve and repeat

Update prompts, retrieval strategy, models, or application logic — then evaluate again.

06 — Roadmap

What's being built

EvalCore is early — two commits in. This is the shape of what's planned next.

Core evaluation metrics
RAG evaluation workflows
Test-data generation
Custom evaluation criteria
Experiment comparison
Production feedback workflows
Agent evaluation support
LLM benchmarking
Prompt evaluation
Workflow evaluation