The LLM Evaluation Framework
Data-Driven Evaluation for LLM-Powered Applications
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
Open-source infrastructure that turns scattered SKILL.md files into curated, retrieval-ready agent-skill corpora—with retrieval and evaluation tooling included.