Eval & LLMOps

Evaluation harnesses, tracing, observability, guardrails and prompt management.

Tools in this category

Eval & LLMOps ★ 34.7k

Open Code Review

Hybrid code review that puts deterministic rules first and the LLM second — the only ordering that survives an enterprise security review.

Teams that need AI code review to pass a security review, not just look impressive in a demo Read more →

Guardrails AI

Validates model output against a schema instead of hoping — and re-asks when validation fails.

Structured output that must parse reliably, and PII or policy checks before content reaches a user Read more →

OpenLLMetry

LLM traces as standard OpenTelemetry spans — so they land in the observability stack you already pay for.

Organisations with an existing OpenTelemetry stack who do not want a second observability silo Read more →
Eval & LLMOps ★ 15.0k

Langfuse

Self-hostable tracing, evals and prompt management — the observability layer most teams forget to build.

Tracing, prompt versioning and production evals Read more →
Eval & LLMOps ★ 11.0k

Ragas

Metric suite for RAG: faithfulness, context precision, recall — the numbers you need before a go-live review.

RAG regression testing and go-live metrics Read more →
Eval & LLMOps ★ 9.0k

Arize Phoenix

Notebook-first observability with clustering that surfaces failure modes you did not think to test.

Local debugging and failure-mode discovery Read more →
Eval & LLMOps ★ 7.0k

Helicone

Drop-in proxy that logs every request and shows cost per user, per feature, per day.

Fast cost attribution and request logging Read more →

Notes in this category