Are you the author? Sign in to claim
LLM-as-judge evaluation framework for assessing AI agent output quality
LLM-as-judge evaluation framework for assessing AI agent output quality
Browser automation using accessibility snapshots instead of screenshots
Official GitHub integration for repos, issues, PRs, and CI/CD workflows
Run Claude Code as an MCP server so any agent can delegate coding tasks to it
Google's universal MCP server supporting PostgreSQL, MySQL, MongoDB, Redis, and 10+ databases