27 packages found
xLAM: A Family of Large Action Models to Empower AI Agent Systems
A simple yet versatile context engineered for scalable online data collection
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE
A general purpose scientific writer
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural langu
Official Repo for ICML 2024 paper "Executable Code Actions Elicit Better LLM Agents" by Xingyao Wang, Yangyi Chen, Lifan
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
[NeurIPS 2024 D&B] GTA: A Benchmark for General Tool Agents & [arXiv 2026] GTA-2
RepairAgent is an autonomous LLM-based agent for software repair.
[🏆 CHI26 Best Paper] CoBRA: Reproducible control of LLM agent behavior via classic social science experiments
OrcaLoca: An LLM Agent Framework for Software Issue Localization [ICML 25]
[ICLR2026] The official repository for the CodeGym project: "Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
Awesome papers involving LLMs in Social Science.
LLM Agent that leverages cheminformatics tools to provide informed responses.
Yunjue Agent: A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks
ML-Dev-Bench is a benchmark for evaluating AI agents against various ML development tasks.
[ICLR'25] OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?
AgentHER: Hindsight Experience Replay for LLM Agents
Sotopia: an Open-ended Social Learning Environment (ICLR 2024 spotlight)
[CVPR 2025 🔥]A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
[ICLR 2026] Meta-RL Induces Exploration in Language Agents
The benchmark tasks and evaluation harness for "PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments".
MR. Video: MapReduce is the Principle for Long Video Understanding
Official Implementation of UA^{2}-Agent and other baseline algorithms of "Towards Unified Alignment Between Agents, Huma
Hypernetworks that update LLMs to remember factual information