69 packages found
A repo lists papers related to LLM based agent
A curated list of Generative AI tools, works, models, and references
[NeurIPS 2024 D&B] GTA: A Benchmark for General Tool Agents & [arXiv 2026] GTA-2
ML-Dev-Bench is a benchmark for evaluating AI agents against various ML development tasks.
Awesome papers involving LLMs in Social Science.
Code and Data for "MIRAI: Evaluating LLM Agents for Event Forecasting"
A Systematic Survey of Deep Research
🧙🏻 Code and benchmark for our Findings of ACL 2024 paper - "TimeChara: Evaluating Point-in-Time Character Hallucinatio
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural langu
[ICLR'25] OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
The benchmark tasks and evaluation harness for "PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments".
Source code for our paper: "ARIA: Training Language Agents with Intention-Driven Reward Aggregation".
[ICML2025 Oral] LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models
Awesome LLM Papers and repos on very comprehensive topics.
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE
Towards Large Multimodal Models as Visual Foundation Agents
scAgent: No-code single-cell analysis for every biologist
🌱 A little course on Reinforcement Learning Environments for evaluating and training Language Models
Official Repo for ICML 2024 paper "Executable Code Actions Elicit Better LLM Agents" by Xingyao Wang, Yangyi Chen, Lifan
[ICML 2024] LLMCompiler: An LLM Compiler for Parallel Function Calling
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
Yunjue Agent: A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks
[NeurIPS 2024] Official implementation for "AgentPoison: Red-teaming LLM Agents via Memory or Knowledge Base Backdoor Po
DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
[Up-to-date] A curated list of resources on graph-empowered agents and agent-facilitated graph learning (Graphs Meet Age
A general purpose scientific writer
OrcaLoca: An LLM Agent Framework for Software Issue Localization [ICML 25]
Asynchronous LLM Agent playing games of Mafia against human players
A curated list of awesome things related to Anthropic Claude
"DeepCode: Open Agentic Coding (Paper2Code & Text2Web & Text2Backend)"
Ask the oracle when you're stuck. Invoke GPT-5 Pro with a custom context and files.
Luann (fka TypeAgent) allows you to create many LLM based agent(Various types of agent,scale up)
Hypernetworks that update LLMs to remember factual information
Odyssey: Empowering Minecraft Agents with Open-World Skills
[ICLR 2025 Oral] This is the official repo for the paper "LLM-SR" on Scientific Equation Discovery and Symbolic Regressi
ICML 2026 · Plug-and-play long-term memory for LLM agents
Elixir implementation of a LangChain style framework that lets Elixir projects integrate with and leverage LLMs.
🔴 VERY LARGE AI TOOL LIST! 🔴 Curated list of AI Tools - Updated 2026
Nemp - The memory plugin for Claude Code that remembers everything.
YantrikDB memory provider for NousResearch/hermes-agent — self-maintaining memory with canonicalization, contradiction t
Discourse AI now lives in the discourse/discourse repo
A Claude Code plugin that stops vague prompts before they run. Detects your stack, asks targeted questions, and forges a
Sotopia: an Open-ended Social Learning Environment (ICLR 2024 spotlight)
A powerful AI assistant integrated into KOReader.
Your ai-powered everything assistant. Oboto is an AI assistant that runs on your computer and in the cloud. Oboto rememb