11 packages found
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, autom
MCP for the open-genes
Evidence-first multi-agent debate skill: get a second opinion by pitting Codex × Claude Code (or GLM/DeepSeek/Qwen) to i
MCP for Scorable Evaluation Platform
A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, tes
A curated list of Generative AI tools, works, models, and references
Autonomous agent for Kubernetes incident detection, diagnosis, and mitigation using LLMs and modular workflows. Integrat
The Mind Palace for AI Agents - HIPAA-hardened Cognitive Architecture with on-device LLM (prism-coder:7b), Hebbian learn
🪢 Open source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Int
AssetOpsBench - Industry 4.0: A unified benchmark and framework for building, orchestrating, and evaluating domain-speci