71 packages found
A repo lists papers related to LLM based agent
A curated list of Generative AI tools, works, models, and references
[NeurIPS 2024 D&B] GTA: A Benchmark for General Tool Agents & [arXiv 2026] GTA-2
ML-Dev-Bench is a benchmark for evaluating AI agents against various ML development tasks.
Awesome papers involving LLMs in Social Science.
Code and Data for "MIRAI: Evaluating LLM Agents for Event Forecasting"
A Systematic Survey of Deep Research
🧙🏻 Code and benchmark for our Findings of ACL 2024 paper - "TimeChara: Evaluating Point-in-Time Character Hallucinatio
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural langu
[ICLR'25] OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
Awesome LLM Papers and repos on very comprehensive topics.
Source code for our paper: "ARIA: Training Language Agents with Intention-Driven Reward Aggregation".
[ICML2025 Oral] LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models
The benchmark tasks and evaluation harness for "PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments".
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE
Towards Large Multimodal Models as Visual Foundation Agents
scAgent: No-code single-cell analysis for every biologist
🌱 A little course on Reinforcement Learning Environments for evaluating and training Language Models
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
Yunjue Agent: A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks
A general purpose scientific writer
Official Repo for ICML 2024 paper "Executable Code Actions Elicit Better LLM Agents" by Xingyao Wang, Yangyi Chen, Lifan
[NeurIPS 2024] Official implementation for "AgentPoison: Red-teaming LLM Agents via Memory or Knowledge Base Backdoor Po
[ICML 2024] LLMCompiler: An LLM Compiler for Parallel Function Calling
DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
[Up-to-date] A curated list of resources on graph-empowered agents and agent-facilitated graph learning (Graphs Meet Age
LLM-as-judge evaluation framework for assessing AI agent output quality
Asynchronous LLM Agent playing games of Mafia against human players
OrcaLoca: An LLM Agent Framework for Software Issue Localization [ICML 25]
"DeepCode: Open Agentic Coding (Paper2Code & Text2Web & Text2Backend)"
A curated list of awesome things related to Anthropic Claude
RAG system design — chunking, embedding models, retrieval strategies, evaluation
ICML 2026 · Plug-and-play long-term memory for LLM agents
[ICLR 2025 Oral] This is the official repo for the paper "LLM-SR" on Scientific Equation Discovery and Symbolic Regressi
Odyssey: Empowering Minecraft Agents with Open-World Skills
Hypernetworks that update LLMs to remember factual information
Elixir implementation of a LangChain style framework that lets Elixir projects integrate with and leverage LLMs.
🔴 VERY LARGE AI TOOL LIST! 🔴 Curated list of AI Tools - Updated 2026
Nemp - The memory plugin for Claude Code that remembers everything.
Ask the oracle when you're stuck. Invoke GPT-5 Pro with a custom context and files.
Luann (fka TypeAgent) allows you to create many LLM based agent(Various types of agent,scale up)
YantrikDB memory provider for NousResearch/hermes-agent — self-maintaining memory with canonicalization, contradiction t
Discourse AI now lives in the discourse/discourse repo
A Claude Code plugin that stops vague prompts before they run. Detects your stack, asks targeted questions, and forges a
Sotopia: an Open-ended Social Learning Environment (ICLR 2024 spotlight)
A powerful AI assistant integrated into KOReader.