57 packages found
Drawdown-first portfolio tool with a read-only MCP addon for Claude — a deterministic core computes every number; the AI
Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents.
Benchmarking the gap between AI agent hype and architecture. Three agent archetypes, 73-point performance spread, stress
Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for Claude Code,
15 AI-powered GTM skills for Claude Code. Campaign tested frameworks for cold email, ICP research, signal scoring, campa
Screen-reader navigation cost analyzer — models the real effort to discover, reach, and operate interactive web content
Local context engine for AI coding agents. Routes tasks to relevant files, tests, rules, and skills, supports prompt cac
A native desktop application for developing, testing, and debugging Model Context Protocol servers.
📄 Production-ready MCP server for PDF processing - 5-10x faster with parallel processing and 94%+ test coverage
One MCP server, 51 tools for AI agents on mobile + canvas UIs: iOS & Android automation, Maestro E2E, evidenced assertio
Autonomous LLM build orchestrator: plans a goal into tasks, edits code with anchored SEARCH/REPLACE, and verifies every
Task-induced context normalization for coding agents — a native Claude Code plugin. The task induces a projection; opera
✨✨Latest Advances on Neuro-Symbolic Learning in the era of Large Language Models
Reproducible evaluation for AI coding agents. Multi-turn scenarios against Claude Code, Codex, Copilot, Cursor, Gemini C
CI for ink stories — compile checks, exhaustive branch playtesting, dead-content detection. MCP server + CLI.
Local-first repo maps for coding agents—ranked files, test routes, risks, CLI/MCP/GitHub Action, and public GitHub URLs.
A curated system of production-ready Claude Code skills with quantitative evaluation reports, golden test fixtures, and
Convert AI coding sessions between Claude Code, Codex, OpenCode, Zed, Aider, Gemini CLI, Cursor, Cline & more — lossless
Autonomous multi-agent SDLC harness: describe a feature in plain English and AI agents scope, code, test, review, and op
The open-source web-browsing backend for AI agents & workflow engines. Ships a 42-tool MCP server for Claude Code/Cursor
Claude 2026: The Developer’s Smartest AI for Long-Context Code & Writing
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
Opt-in Codex Skill for practical coding and externally anchored delivery with elastic agent teams, task DAGs, isolated c
Claude Code skills for medical research — literature search, reporting guidelines, statistical analysis, publication fig
A Claude Code skill that burns tokens on demand. Stress test, inflate metrics, or just set money on fire.
Build Claude Code–style deep agents in Python: tool-calling, sandboxed execution, multi-agent teams, skills, checkpoints
ConcoLLMic: the first language- and theory-agonistic concolic execution engine via LLM agents
Self-installing personal AI orchestrator. Hand the latest blueprint file to a Claude Code session and it builds a full m
Agents make claims. Reelier writes receipts — record an agent's tool-call workflow once, replay it deterministically at
Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ mod
63 deterministic quant computation tools for autonomous financial agents. Options, derivatives, risk, portfolio, statist
AI-native framework for building trading systems with polyglot bindings.
CLI-first API testing tool. YAML-defined tests, structured JSON output, built for AI-assisted workflows.
Synthetic monitoring and CI smoke tests for LLM inference endpoints.
A curated list of developer tools, SDKs, libraries, and testing utilities for Model Context Protocol (MCP) server develo
Combining a five-level AI framework with git-native memory overcomes session amnesia, enabling anticipation of problems
Regression testing for AI agents. Snapshot behavior,diff tool calls,catch regressions in CI. Works with LangGraph, CrewA
Permanent guardrails for AI-coded apps
A Claude Code skill that encodes battle-tested editorial principles, section-specific rhetorical moves, and a structured
A Claude Skill to give your agent the ability to use a web browser
Pre-execution governance for AI agents. Sub-millisecond tool call validation, drift detection, circuit breakers, human-i
AI skill for Claude Code and Codex that helps agents write correct R for Six Sigma and SPC work, including control chart
CLI, MCP server, and npm library that turns any website into an API — no docs, no SDK, no browser.
Zero-dependency browser automation CLI. 70+ commands, 10 test assertions, smart commands (click/fill by text — no LLM ne
A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, tes
Tired of paying for frontier AI models? Break free with this guide to 100% local, private LLM code generation on Apple S
High-performance Rust hooks for Claude Code skill auto-activation. ~2ms startup, zero dependencies, production-tested pa
Open-source rules engine for Cursor IDE — 110 rules, agents, skills, commands that give AI permanent memory of your stan
Adversarial multi-model reasoning verification MCP server for AI agents. Claude, Grok, and DeepSeek challenge each decis
Autonomous orchestration framework for Claude Code with MemPalace-inspired memory (4-layer stack, 818-token wake-up), pa