137 packages found
Monocle is a framework for tracing GenAI app code. This repo contains implementation of Monocle for GenAI apps written i
Lightweight, auditable Python code agent (~1500 LOC) — ReAct + Planner + Reflexion + Hybrid RAG, with SWE-bench Lite e
AI agent platform for building multi-agent systems with orchestration, memory, RAG, workflows, and enterprise observabi
Desktop app for multi-workspace Claude Code management
An AI-powered QA tool for Android. Claude Code controls Android via ADB, executes test scenarios, and records every inte
Ask the oracle when you're stuck. Invoke GPT-5 Pro with a custom context and files.
Hivemind turns your traces into reusable skills across agents
A simple and well-tailored LLM application framework that enables you to seamlessly integrate LLM capabilities in the mo
A LangGraph-powered multi-agent deep research system featuring task planning, human-in-the-loop review, multi-source ret
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
GoClaw - GoClaw is OpenClaw rebuilt in Go — with multi-tenant isolation, 5-layer security, and native concurrency. Deplo
Most AI agents forget you the moment the tab closes. Constellation Engine gives them a hippocampus — a living star map w
Build, run and scale AI agents like API and microservices - observable,auditable and identity-aware from day one.
Run Claude Desktop’s Cowork mode natively on Linux — no macOS or VM required
A collection of production-ready slash commands for Claude Code
From a goal to a task DAG, automatically. TypeScript-native multi-agent orchestration.
Automatic LLM router — 82% cost savings, 79.4% accuracy, 93.4% pass rate. Drop-in OpenAI proxy.
A universal git-native AI agent framework. Your agent lives inside a git repo — identity, rules, memory, tools, and skil
Yunjue Agent: A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks
A repo lists papers related to LLM based agent
[ICLR'25] OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?
A transparent, minimal, and hackable agent framework. ~300 lines of readable code. Full control, no magic.
Skill Scan Agent — Automated scanning, identification, and assessment of SKILL security risks.
ML-Dev-Bench is a benchmark for evaluating AI agents against various ML development tasks.
Fully autonomous AI Agents system capable of performing complex penetration testing tasks
A high-performance LLM inference API and Chat UI that integrates DeepSeek R1's CoT reasoning traces with Anthropic Claud
How real engineers run Claude Code and Codex: spec-driven planning, enforced TDD, persistent memory, and quality enforce
50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime.
Build, Improve Performance, and Productionize your LLM Application with an Integrated Framework
Provider-neutral Agent Skill for Codex, Claude Code, and agentic harness design.
⚡️ Open-source AI Gateway — Use any SDK to call 100+ LLMs. Built-in failover, load balancing, cost control & end-to-end
Multi-agent autonomous SDLC framework. Spec to deployed app. PRD, GitHub issue, OpenAPI/JSON/YAML, or one-line brief. 5
AgentNotch is a sleek macOS menu bar app that lives in your Mac's notch, providing real-time visibility into your AI cod
All-in-one agent harness for OpenAI Codex CLI — Boss meta-orchestrator, 400+ agents, 200+ skills, 3 MCP servers. Install
A self-hosted AI workspace unifying chat, code execution, parallel multi-agent orchestration, and project management. Ea
An open-source Digital Worker platform for reliable execution and continuous co-evolution.
LLM QA, Observability, Evaluation and User Feedback
历年ICLR论文和开源项目合集,包含ICLR2021、ICLR2022、ICLR2023、ICLR2024、ICLR2025.
♾️ Private Agent Fleet with Spec Coding. Each agent gets their own GPU-accelerated desktop. Run Claude, Codex, Gemini an
The missing DevTools for Claude Code — inspect session logs, tool calls, token usage, subagents, and context window in a
A simple, observable code-writing agent builder in TypeScript.
Official companion repository for our survey "A Survey of the OpenClaw Ecosystem: From Platform Extensibility to Constra
Run Claude agents in secure cloud sandboxes — via API, CLI, or Slack. One call. Full agent. Zero infrastructure.
Turn your markdown vault into a compounding knowledge wiki (Karpathy inspired). Six agent skills - knowledge grows with
MeshCore Hub provides a complete solution for monitoring, collecting, and interacting with MeshCore mesh networks