827 packages found
A repo lists papers related to LLM based agent
It is a comprehensive resource hub compiling all LLM papers accepted at the International Conference on Learning Represe
A curated list of Generative AI tools, works, models, and references
Bilig WorkPaper: headless spreadsheet formula engine and MCP server for Node agents: complete spreadsheet work autonomou
Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool
All-in-one Web Agent framework for post-training. Start building with a few clicks!
[NeurIPS 2024 D&B] GTA: A Benchmark for General Tool Agents & [arXiv 2026] GTA-2
ML-Dev-Bench is a benchmark for evaluating AI agents against various ML development tasks.
Unreal Engine Vibe Coding tool
4-stage evaluation framework for testing Claude Code plugin component triggering. Validates skills, agents, and commands
历年ICLR论文和开源项目合集,包含ICLR2021、ICLR2022、ICLR2023、ICLR2024、ICLR2025.
Awesome papers involving LLMs in Social Science.
CLI & MCP server for Tuning Engines — fine-tune LLMs on code repositories
Translation quality evaluation MCP Server powered by xCOMET (eXplainable COMET).
A Node.js package and GitHub Action for evaluating MCP (Model Context Protocol) tool implementations using LLM-based sco
A Claude Code skill that encodes battle-tested editorial principles, section-specific rhetorical moves, and a structured
🚀 MassGen is an open-source multi-agent scaling system that runs in your terminal, autonomously orchestrating frontier
Agent-native image-editing SDK for Claude Code. 21 MCP tools + /decompose skill — semantic layer splits, L1–L5 cultural
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your cl
AI-powered job search system built on Claude Code. 14 skill modes, Go dashboard, PDF generation, batch processing.
Multi-agent orchestration system for Claude Code with parallel execution, automated quality gates, Board of Directors, a
Evaluate agent skill quality. Find the weakest link. Fix it. Prove it worked.
mcp-occams-razor
HealthFlow: Automating electronic health record analysis via a strategically self-evolving multi-agent framework
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
AI-powered code review CLI with multiple providers (Gemini, Claude, OpenAI). Features 95%+ token reduction via semantic
TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and e
Advanced cognitive reasoning MCP server — DAG thought graph, 10 strategies, metacognition, self-critique, knowledge inte
A curated list of tools, papers, and datasets for applying AI to cybersecurity tasks. This list primarily focuses on mod
B2B software vendor evaluation skill for Claude Code — domain-expert questions, vendor AI agent conversations, evidence-
Intelligent prompt improver hook for Claude Code. Type vibes, ship precision.
A curated system of production-ready Claude Code skills with quantitative evaluation reports, golden test fixtures, and
Open source GitLab MCP server for AI assistants: 2-tool dynamic find/execute over 870+ GitLab actions (1,000+ Enterprise
18 mental models and critical thinking frameworks for Claude Code - First Principles, Bayesian, Systems Thinking, OODA,
Academic Research Skills for Claude Code: research → write → review → revise → finalize
Code and Data for "MIRAI: Evaluating LLM Agents for Event Forecasting"
Framework and toolkits for building and evaluating collaborative agents that can work together with humans.
AssetOpsBench - Industry 4.0: A unified benchmark and framework for building, orchestrating, and evaluating domain-speci
jetbrains-debugger-mcp-plugin
Timestamped web intelligence for AI agents. MCP server with guaranteed freshness envelopes.
A powerful RAG (Retrieval-Augmented Generation) system built with LangChain, designed as an MCP server for Cursor, VS Co
SPEC-First Agentic Development Kit for Claude Code — 24 AI agents + 52 skills with TDD/DDD quality gates, 16-language pr
A Systematic Survey of Deep Research
A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, a
Open source implementation and extension of Google Research’s PaperBanana for automated academic figures, diagrams, and
Production-ready Claude Code skills for building AI agents with Pydantic AI. Includes dependency injection, tools, valid
This repo contains detailed implementation information about Anthropic's paired prompts approach for evaluating politica
The most comprehensive toolkit for Claude Code -- 135 agents, 35 curated skills, 42 commands, 176+ plugins, 20 hooks, 15