MLflow LLMOps

Active
GitHub Python Apache-2.0

Description

MLflow LLM / GenAI extensions. One-stop LLMOps: prompt versioning, evaluation, model registry, tracing.

Strengths & Limitations

Strengths

  • Actively maintained, recent updates
  • High community interest (27.5k stars)
  • Permissive open-source license (Apache-2.0)

⚠️ Limitations

  • High issue backlog (2.0k open issues)

Related Projects

Agenta

4.5k · TypeScript
Active B

Agenta is an open-source LLMOps platform providing prompt playground, prompt management, LLM evaluation, and LLM observability all in one place.

observabilityllmopsprompt-management +2
  • · Interactive LLM Playground for side-by-side prompt comparison with 50+ model support
  • · Integrated prompt management with version control, branching, and environment management
  • · Systematic LLM evaluation with 20+ pre-built evaluators, LLM-as-judge, and human feedback

Giskard

5.8k · Python
Active A+

An open-source evaluation and testing library for LLM agents providing automated model scanning, bias detection, performance benchmarking, and compliance checks.

evaluationtestingllm-safety +3
  • · Scenario API for creating evaluations that test non-deterministic LLM outputs
  • · Built-in checks including Groundedness, Conformity, and LLM-as-judge assessments
  • · Red-teaming vulnerability scanner generating adversarial test suites across OWASP LLM Top-10 categories

Agents Towards Production

21.3k · Jupyter Notebook
Active A+

End-to-end, code-first tutorials for building production-grade GenAI agents. From prototype to enterprise deployment.

agentframeworkevaluation +2
  • · 28 production-grade tutorials — Covering stateful workflows, vector memory, web search APIs and more
  • · End-to-end coverage — From Docker deployment, FastAPI endpoints to security guardrails and GPU scaling
  • · Multi-agent coordination — Teaching multi-Agent collaboration architecture design and implementation

SwanLab

4.2k · Python
Active A+

An open-source, modern-design AI training tracking and visualization tool. Supports PyTorch, Transformers and more. Monitor and evaluate AI agent training processes.

pythonobservabilityevaluation +2
  • · Seamless integration with 50+ mainstream frameworks: native support for PyTorch, Transformers, HuggingFace Accelerate, PaddleNLP, NVIDIA NeMo RL and more, with two lines of code to connect training pipelines
  • · Rich visualization system: supports line charts, scalar plots, PR curves, ROC curves, confusion matrices, 3D point clouds, molecular structures, ECharts custom charts and 20+ chart types
  • · Multi-dimensional hardware monitoring: real-time monitoring of GPU (NVIDIA/AMD ROCm/Hygon DCU/Cambricon MLU/Moore Threads/Muxi/Iluvatar/Kunlun), disk utilization, network traffic and other hardware metrics