4 packages found
Benchmarking the gap between AI agent hype and architecture. Three agent archetypes, 73-point performance spread, stress
A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, tes
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewA
Lightweight, auditable Python code agent (~1500 LOC) — ReAct + Planner + Reflexion + Hybrid RAG, with SWE-bench Lite e