126 packages found
It is a comprehensive resource hub compiling all LLM papers accepted at the International Conference on Learning Represe
A curated list of tools, papers, and datasets for applying AI to cybersecurity tasks. This list primarily focuses on mod
Code and Data for "MIRAI: Evaluating LLM Agents for Event Forecasting"
历年ICLR论文和开源项目合集,包含ICLR2021、ICLR2022、ICLR2023、ICLR2024、ICLR2025.
A Python library powered by Language Models (LLMs) for conversational data discovery and analysis.
MASSW is a comprehensive text dataset on Multi-Aspect Summarization of Scientific Workflows. MASSW includes more than 15
Official Implementation for the Paper [AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent](https://ar
LLM Agent paired with Image Captioning and Yolov8 models plays God of War
A repo lists papers related to LLM based agent
💻 A curated list of papers and resources for multi-modal Graphical User Interface (GUI) agents.
[ICML2025 Oral] LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models
Awesome papers involving LLMs in Social Science.
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural langu
🧙🏻 Code and benchmark for our Findings of ACL 2024 paper - "TimeChara: Evaluating Point-in-Time Character Hallucinatio
Connect RStudio to Claude Code, Codex, Gemini, and other LLM agents via MCP. Multi-agent orchestration, automated manusc
A curated list of Generative AI tools, works, models, and references
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
From a sample document to a running Milvus / Zilliz Cloud search app in minutes — an AI-guided scaffold delivered as a A
[ICLR'25] OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?
Official companion repository for our survey "A Survey of the OpenClaw Ecosystem: From Platform Extensibility to Constra
HealthFlow: Automating electronic health record analysis via a strategically self-evolving multi-agent framework
xLAM: A Family of Large Action Models to Empower AI Agent Systems
A Pair Programming Framework for Code Generation via Multi-Plan Exploration and Feedback-Driven Refinement, ASE 2024 (D
[NeurIPS 2024] Official implementation for "AgentPoison: Red-teaming LLM Agents via Memory or Knowledge Base Backdoor Po
Framework and toolkits for building and evaluating collaborative agents that can work together with humans.
[NeurIPS 2024 D&B] GTA: A Benchmark for General Tool Agents & [arXiv 2026] GTA-2
Official Implementation of Dynamic LLM-Agent Network: An LLM-agent Collaboration Framework with Agent Team Optimization
All-in-one Web Agent framework for post-training. Start building with a few clicks!
[ICLR 2025 Oral] This is the official repo for the paper "LLM-SR" on Scientific Equation Discovery and Symbolic Regressi
A multi-user LLM agent whose knowledge, skills, and tools live in the database — grows at runtime, no redeploys. Web cha
AMG-RAG (Agentic Medical Graph-RAG) is a comprehensive framework that automates the construction and continuous updating
[EMNLP'24] EHRAgent: Code Empowers Large Language Models for Complex Tabular Reasoning on Electronic Health Records
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google, ...). 10+
Asynchronous LLM Agent playing games of Mafia against human players
[WWW-2025] ColaCare: Enhancing Electronic Health Record Modeling through Large Language Model-Driven Multi-Agent Collabo
A collection of 2025 agentic workflows built in n8n. Showcases manual multi-model orchestration, RAG-to-SQL, and autonom
✨✨Latest Advances on Neuro-Symbolic Learning in the era of Large Language Models
SkillOrchestra: Learning to Route Agents via Skill Transfer
Towards Large Multimodal Models as Visual Foundation Agents
[CVPR2024 Highlight] Editable Scene Simulation for Autonomous Driving via LLM-Agent Collaboration
👾 Open source implementation of the ChatGPT Code Interpreter
AlgoTune is a NeurIPS 2025 benchmark made up of 154 math, physics, and computer science problems. The goal is write code
Official Repo for ICML 2024 paper "Executable Code Actions Elicit Better LLM Agents" by Xingyao Wang, Yangyi Chen, Lifan
A framework that uses multi-agents to enable users to perform a systematic data science pipeline with just two inputs.
Odyssey: Empowering Minecraft Agents with Open-World Skills
This repo covers LLM, Agents, MCP Tools, Skills concepts with sample codes: LangChain & LangGraph, AWS Strands Agents, G
ML-Dev-Bench is a benchmark for evaluating AI agents against various ML development tasks.