63 packages found
xLAM: A Family of Large Action Models to Empower AI Agent Systems
A general purpose scientific writer
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE
A simple yet versatile context engineered for scalable online data collection
🔥 An autonomous AI agent that runs your deep learning experiments 24/7 while you sleep. Zero-cost monitoring, Leader-Wo
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
An automated AI research-paper writer based off Google's PaperOrchestra paper's implementation through a skills - bench
MASSW is a comprehensive text dataset on Multi-Aspect Summarization of Scientific Workflows. MASSW includes more than 15
Official Implementation of UA^{2}-Agent and other baseline algorithms of "Towards Unified Alignment Between Agents, Huma
HumanStudy-Bench: Towards AI Agent Design for Participant Simulation
Official companion repository for our survey "A Survey of the OpenClaw Ecosystem: From Platform Extensibility to Constra
SkillOrchestra: Learning to Route Agents via Skill Transfer
Hypernetworks that update LLMs to remember factual information
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural langu
Official Repo for ICML 2024 paper "Executable Code Actions Elicit Better LLM Agents" by Xingyao Wang, Yangyi Chen, Lifan
Agentic-RAG explores advanced Retrieval-Augmented Generation systems enhanced with AI LLM agents.
AI-Powered Predictive Maintenance & Fault Diagnosis through Model Context Protocol. An open-source framework for integra
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
[NeurIPS 2024] ReEvo: Large Language Models as Hyper-Heuristics with Reflective Evolution
[SOSP'25] Automatic checker synthesis for system-level static analysis
Conversational & memory-enabled AI research partner for multi-omics analysis. CLI + Desktop App (installers in Releases)
[NeurIPS 2024 D&B] GTA: A Benchmark for General Tool Agents & [arXiv 2026] GTA-2
ConcoLLMic: the first language- and theory-agonistic concolic execution engine via LLM agents
RepairAgent is an autonomous LLM-based agent for software repair.
AlgoTune is a NeurIPS 2025 benchmark made up of 154 math, physics, and computer science problems. The goal is write code
[🏆 CHI26 Best Paper] CoBRA: Reproducible control of LLM agent behavior via classic social science experiments
[NeurIPS 2024 D&B] VideoGUI: A Benchmark for GUI Automation from Instructional Videos
Official Implementation for the Paper [AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent](https://ar
HealthFlow: Automating electronic health record analysis via a strategically self-evolving multi-agent framework
OrcaLoca: An LLM Agent Framework for Software Issue Localization [ICML 25]
[ICLR2026] The official repository for the CodeGym project: "Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
Awesome papers involving LLMs in Social Science.
LLM Agent that leverages cheminformatics tools to provide informed responses.
A theoretical reconstruction of the Claude Mythos architecture, built from first principles using the available research
DeepSeek CLI, a command-line AI coding assistant that leverages the powerful DeepSeek Coder models
syftr is an agent optimizer that helps you find the best agentic workflows for your budget.
Yunjue Agent: A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks
ML-Dev-Bench is a benchmark for evaluating AI agents against various ML development tasks.
Open-source Claude Design alternative. One-click import your Claude Code / Codex API key. Prompt → prototype / slides /
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
🌸 Best framework to build web agents, and deploy serverless web automation functions on reliable browser infra.
🤖 24/7 AI agent that maximizes Claude Code Pro usage via Slack. Auto-processes tasks, manages isolated workspaces, crea
Agent memory for LLMs: 30 runnable Jupyter notebooks covering conversation buffers, vector stores, knowledge graphs, epi
AAAI24(Oral) ProAgent: Building Proactive Cooperative Agents with Large Language Models
[ICLR'25] OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?
[EMNLP'24] EHRAgent: Code Empowers Large Language Models for Complex Tabular Reasoning on Electronic Health Records
AgentHER: Hindsight Experience Replay for LLM Agents