2 packages found
Benchmarking the gap between AI agent hype and architecture. Three agent archetypes, 73-point performance spread, stress
Lightweight, auditable Python code agent (~1500 LOC) — ReAct + Planner + Reflexion + Hybrid RAG, with SWE-bench Lite e