3 packages found
Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for Claude Code,
Diff your AI agent's behavior between two runs. See exactly which tool calls, args, costs and outputs changed when you s
MCP server and typed SDK for the VerifyAX agent-evaluation platform.