EVAL-FIRST TOOLKIT FOR LLM APPLICATIONS
Measure the intelligence. Improve the outcome.
EvalCore turns "the model seems to work" into evidence: structured metrics, test datasets, and quality signals your team can act on — built into the engineering workflow, not bolted on at the end.
A developer-focused evaluation toolkit
EvalCore is built for the moments where a demo looks convincing but you still can't say, with evidence, whether it's good. It gives every project the same short list of questions to answer before shipping a change.
- Is the response accurate?
- Is the retrieved context relevant?
- Does the answer stay grounded in the available information?
- How does a new model or prompt compare with the previous one?
- Where does the application need improvement?
Evaluation isn't a final inspection. It's part of the engineering workflow — build the application, measure its behavior, improve what matters.
Uncertainty, turned into evidence
LLM applications can produce convincing answers without consistently producing correct ones. EvalCore is built around a closed feedback loop that catches the difference.
Five things EvalCore is built to do
Each capability targets one part of the loop above — from generating test data to feeding results back into the next iteration.
Define a metric, score a response
The example below defines a custom accuracy check and scores a single application output — the smallest complete evaluation loop.
import asyncio
from openai import AsyncOpenAI
from ragas.metrics import DiscreteMetric
from ragas.llms import llm_factory
# Configure the evaluation model
client = AsyncOpenAI()
llm = llm_factory("gpt-4o", client=client)
# Define what you want to measure
metric = DiscreteMetric(
name="summary_accuracy",
allowed_values=["accurate", "inaccurate"],
prompt="Evaluate if the summary is accurate."
)
# Evaluate an application output
async def main():
score = await metric.ascore(
llm=llm,
response="The summary of the text is..."
)
print(f"Score: {score.value}")
print(f"Reason: {score.reason}")
asyncio.run(main())
A typical EvalCore project
Five stages, repeated. The goal is never a single score — it's a system that keeps getting better.
Define the quality signal
Decide what matters — accuracy, relevance, faithfulness, completeness, or a custom domain-specific criterion.
Prepare evaluation data
Pull representative examples from your application, including the cases where quality matters most.
Run evaluations
Apply metrics to your application's outputs and collect structured, comparable results.
Inspect the results
Look for patterns, recurring failure cases, and shifts in performance over time.
Improve and repeat
Update prompts, retrieval strategy, models, or application logic — then evaluate again.
What's being built
EvalCore is early — two commits in. This is the shape of what's planned next.