53 packages found
Source-backed GPT-5.6 use cases for coding, agents, creative work, integrations, benchmarks, and practical limits.
A repo lists papers related to LLM based agent
A curated list of Generative AI tools, works, models, and references
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE
[Up-to-date] A curated list of resources on graph-empowered agents and agent-facilitated graph learning (Graphs Meet Age
[NeurIPS 2024 D&B] GTA: A Benchmark for General Tool Agents & [arXiv 2026] GTA-2
A Systematic Survey of Deep Research
[ICML2025 Oral] LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models
[ICML 2024] LLMCompiler: An LLM Compiler for Parallel Function Calling
Awesome papers involving LLMs in Social Science.
Awesome LLM Papers and repos on very comprehensive topics.
xLAM: A Family of Large Action Models to Empower AI Agent Systems
[ICLR 2025 Oral] This is the official repo for the paper "LLM-SR" on Scientific Equation Discovery and Symbolic Regressi
The benchmark tasks and evaluation harness for "PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments".
Yunjue Agent: A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks
Agent skills for Claude Code where every entry ships with receipts: accuracy-gated benchmarks vs baseline AND placebo. R
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
Odyssey: Empowering Minecraft Agents with Open-World Skills
"DeepCode: Open Agentic Coding (Paper2Code & Text2Web & Text2Backend)"
ML-Dev-Bench is a benchmark for evaluating AI agents against various ML development tasks.
LLM Agent that leverages cheminformatics tools to provide informed responses.
Official Implementation of UA^{2}-Agent and other baseline algorithms of "Towards Unified Alignment Between Agents, Huma
An experimental game engine for robots (and their humans), made by robots (and their humans)
YantrikDB memory provider for NousResearch/hermes-agent — self-maintaining memory with canonicalization, contradiction t
[CVPR 2025 🔥]A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
ICML 2026 · Plug-and-play long-term memory for LLM agents
Code and Data for "MIRAI: Evaluating LLM Agents for Event Forecasting"
Agentic Theorem Prover for Rocq for Program Verification
🧙🏻 Code and benchmark for our Findings of ACL 2024 paper - "TimeChara: Evaluating Point-in-Time Character Hallucinatio
A general purpose scientific writer
A curated list of awesome things related to Anthropic Claude
A human-in-the-loop control plane for reliable agentic coding—plan, challenge, implement, verify, and preserve progress.
🔴 VERY LARGE AI TOOL LIST! 🔴 Curated list of AI Tools - Updated 2026
[ICLR'25] OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?
[ACL 2024 Findings] MedAgents: Large Language Models as Collaborators for Zero-shot Medical Reasoning https://arxiv.org/
The pretty much "official" DSPy framework for Typescript
🤖 Awesome list of AGI Agents. Agents 精选资源合集.
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
Frontier models sell confidence. FABULA ships proof — an agent harness where any model is a swappable chip and every fin
The official repository of "SmartAgent: Chain-of-User-Thought for Embodied Personalized Agent in Cyber World".
Official Repo for ICML 2024 paper "Executable Code Actions Elicit Better LLM Agents" by Xingyao Wang, Yangyi Chen, Lifan
Structured deep research skill for Claude Code/Open Code/Codex with human-in-the-loop control
Towards Large Multimodal Models as Visual Foundation Agents
RepairAgent is an autonomous LLM-based agent for software repair.
[ICLR2026] The official repository for the CodeGym project: "Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
Code repo for the paper: Attacking Vision-Language Computer Agents via Pop-ups
MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)