225 packages found
A repo lists papers related to LLM based agent
It is a comprehensive resource hub compiling all LLM papers accepted at the International Conference on Learning Represe
A curated list of Generative AI tools, works, models, and references
All-in-one Web Agent framework for post-training. Start building with a few clicks!
[NeurIPS 2024 D&B] GTA: A Benchmark for General Tool Agents & [arXiv 2026] GTA-2
ML-Dev-Bench is a benchmark for evaluating AI agents against various ML development tasks.
历年ICLR论文和开源项目合集,包含ICLR2021、ICLR2022、ICLR2023、ICLR2024、ICLR2025.
Awesome papers involving LLMs in Social Science.
Multi-agent orchestration system for Claude Code with parallel execution, automated quality gates, Board of Directors, a
HealthFlow: Automating electronic health record analysis via a strategically self-evolving multi-agent framework
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and e
A curated list of tools, papers, and datasets for applying AI to cybersecurity tasks. This list primarily focuses on mod
18 mental models and critical thinking frameworks for Claude Code - First Principles, Bayesian, Systems Thinking, OODA,
Code and Data for "MIRAI: Evaluating LLM Agents for Event Forecasting"
Framework and toolkits for building and evaluating collaborative agents that can work together with humans.
SPEC-First Agentic Development Kit for Claude Code — 24 AI agents + 52 skills with TDD/DDD quality gates, 16-language pr
A Systematic Survey of Deep Research
From a sample document to a running Milvus / Zilliz Cloud search app in minutes — an AI-guided scaffold delivered as a A
Claude Skills for Governance, Risk, & Compliance (GRC): Expert-level compliance guidance for ISO 27001, SOC 2, FedRAMP,
Official companion repository for our survey "A Survey of the OpenClaw Ecosystem: From Platform Extensibility to Constra
AI agent platform for building multi-agent systems with orchestration, memory, RAG, workflows, and enterprise observabi
Agentic-RAG explores advanced Retrieval-Augmented Generation systems enhanced with AI LLM agents.
🧙🏻 Code and benchmark for our Findings of ACL 2024 paper - "TimeChara: Evaluating Point-in-Time Character Hallucinatio
Multi-agent cryptocurrency intelligence system built with Claude Code. 6 AI agents, 66 MCP tools, Agent Teams, zero orch
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural langu
Run AI coding agents unattended for hours and ship PRs worth merging. Cybernetics-based multi-agent orchestration + cros
Open Source Generative Process Automation (i.e. Generative RPA). AI-First Process Automation with Large ([Language (LLMs
Claude Cowork plugin for job seekers. 9 AI skills: evaluate job postings, generate ATS-optimized resumes, scan company c
Official Implementation for the Paper [AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent](https://ar
MASSW is a comprehensive text dataset on Multi-Aspect Summarization of Scientific Workflows. MASSW includes more than 15
Deterministic, resumable, tournament-based orchestrator for LLM-driven software development. Turns Claude Code or Cursor
Monocle is a framework for tracing GenAI app code. This repo contains implementation of Monocle for GenAI apps written i
[ICLR'25] OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?
AI agent firewall that intercepts tool calls (file, shell, network) and enforces deterministic policies at sub-microseco
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
🗜️ Codebase-digest is your AI-friendly codebase packer and analyzer. Features 60+ coding prompts and generates structur
[ICML2025 Oral] LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models
Awesome LLM Papers and repos on very comprehensive topics.
The benchmark tasks and evaluation harness for "PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments".
Source code for our paper: "ARIA: Training Language Agents with Intention-Driven Reward Aggregation".
Agent memory for LLMs: 30 runnable Jupyter notebooks covering conversation buffers, vector stores, knowledge graphs, epi
Windows Agent Arena (WAA) 🪟 is a scalable OS platform for testing and benchmarking of multi-modal AI agents.
💻 A curated list of papers and resources for multi-modal Graphical User Interface (GUI) agents.
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE
MobileUse: an open-source mobile GUI agent for Android phone automation, AndroidWorld/AndroidLab evaluation, hierarchica
Towards Large Multimodal Models as Visual Foundation Agents