📊

Best Observability Top 20

Top 20 most popular open-source Observability projects, ranked by GitHub Stars.

1

Netdata

80.2k Stars

Netdata is an open-source real-time observability platform with second-resolution metrics, AI health checks and zero-config deployment, widely used for monitoring AI agent services.

monitoringmetricsreal-timeagent
2

Kong

44.0k Stars

The cloud-native API and AI Gateway providing LLM request routing, rate limiting, load balancing and observability for AI agent applications.

observabilityapiagentlua
3

PostHog

37.7k Stars

PostHog is a full-stack observability platform for AI agents and self-driving products, combining analytics, session replay, feature flags and error tracking.

analyticssession-replayfeature-flagsai-observability
4

Langfuse

33.2k Stars

Open-source LLM engineering platform providing tracing, evaluations, prompt management, and dataset management with integrations for LangChain, OpenAI, Anthropic, and more.

observabilitytracingllm-evaluationprompt-management
5

Prompt Optimizer

33.1k Stars

An AI prompt optimizer that helps users write better prompts and achieve improved AI results.

prompt-engineeringevaluationllmtypescript
6

SigNoz

31.9k Stars

SigNoz is an open-source OpenTelemetry-native observability platform combining APM, logs, metrics and alerts.

opentelemetrytracinglogsmetrics
7

MLflow LLMOps

27.5k Stars

MLflow LLM / GenAI extensions. One-stop LLMOps: prompt versioning, evaluation, model registry, tracing.

mlopsevaluationmodel-registry
8

MLflow

27.5k Stars

MLflow is the open-source AI engineering platform for debugging, evaluating, monitoring, and optimizing AI agents and LLM applications, with model and data access management.

mlflowllmopsevaluationobservability
9

12 Factor Agents

25.3k Stars

What are the principles we can use to build LLM-powered software that is actually good enough to put in the hands of production customers?

agentframeworkevaluationobservability
10

Promptfoo

24.3k Stars

Test and evaluate LLM prompts, agents, and RAG pipelines. Built-in red teaming and security evaluation for reliable AI applications.

testingevaluationred-teamingprompt-testing
11

Opik

21.4k Stars

Opik is an open-source LLM observability platform providing agent tracing, evaluation testing, and prompt experiment management to help developers monitor and optimize AI agent systems.

observabilityllm-evaluationtracingprompt-testing
12

Agents Towards Production

21.3k Stars

End-to-end, code-first tutorials for building production-grade GenAI agents. From prototype to enterprise deployment.

agentframeworkevaluationobservability
13

openobserve

21.1k Stars

OpenObserve is a high-performance observability platform for logs, metrics, and traces, well suited for monitoring AI agent runtimes and tool calls.

observabilitylogsmetricstracing
14

OpenAI Evals

19.2k Stars

OpenAI's framework for evaluating LLMs and LLM systems, providing an open-source registry of benchmarks and tools for systematic model assessment.

llm-evaluationbenchmarkevalsred-teaming
15

ccusage

18.0k Stars

Analyze coding (agent) CLI token usage and costs from local data.

token-usagecost-analysisclirust
16

DeepEval

17.6k Stars

DeepEval is an open-source evaluation framework for LLM applications. It provides rich evaluation metrics and tools, supporting unit testing and integration testing to help developers build reliable LLM applications.

llmevaluationtestingrag
17

RagaAI Catalyst

16.1k Stars

RagaAI Catalyst is an observability, monitoring, and evaluation framework for Agent AI, supporting agent/LLM/tool tracing, multi-agent debugging, and self-hosted dashboard analytics.

observabilitytracingevaluationagent-monitoring
18

Ragas

15.3k Stars

Ragas is a framework for evaluating RAG (Retrieval Augmented Generation) systems. It provides various evaluation metrics including faithfulness, answer relevance, context precision, helping developers optimize RAG application performance.

ragevaluationllmtesting
19

OpenMetadata

14.9k Stars

OpenMetadata is a unified metadata platform for data and AI, providing data asset discovery, lineage, governance, and agent context retrieval capabilities.

observabilitymetadatadata-governancelineage
20

LM Evaluation Harness

13.7k Stars

A framework for few-shot evaluation of language models by EleutherAI, providing standardized evaluation pipelines supporting hundreds of benchmark tasks and widely adopted as a core LLM evaluation tool in the community.

llm-evaluationbenchmarkevaluation-frameworklanguage-model

Related Articles