31 packages found
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE
A curated list of Generative AI tools, works, models, and references
A map of what AI tokens actually cost, and where they're wasted vs. well spent. Tools, research, practices, and copy-pas
xLAM: A Family of Large Action Models to Empower AI Agent Systems
Source-backed GPT-5.6 use cases for coding, agents, creative work, integrations, benchmarks, and practical limits.
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
Non-custodial Hyperliquid execution + tracking CLI for autonomous LLM agents. Signs EIP-712 actions with an agent wallet
Awesome LLM Papers and repos on very comprehensive topics.
[NeurIPS 2024 D&B] GTA: A Benchmark for General Tool Agents & [arXiv 2026] GTA-2
The GEP-powered self-evolving engine for AI agents. Auditable evolution with Genes, Capsules, and Events. | evomap.ai
Open-source AI gateway — pool your provider keys behind one OpenAI-compatible endpoint with load balancing, failover, ci
Vlad's Playbook — the operator's field manual where every artifact is live, clickable, and forwardable. 39 chapters · 25
Towards Large Multimodal Models as Visual Foundation Agents
Lightweight, auditable Python code agent (~1500 LOC) — ReAct + Planner + Reflexion + Hybrid RAG, with SWE-bench Lite e
Deterministic quality scorer for AI agent instruction files — 8-dimension scoring with security, multi-format (SKILL.md,
Awesome papers involving LLMs in Social Science.
~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google, ...). 10+
Multi-agent code review mesh — orchestrates AI agents from multiple providers to review code in parallel, cross-review e
🧙🏻 Code and benchmark for our Findings of ACL 2024 paper - "TimeChara: Evaluating Point-in-Time Character Hallucinatio
HumanStudy-Bench: Towards AI Agent Design for Participant Simulation
A repo lists papers related to LLM based agent
Telegram bot for Claude Code and Codex CLI with MCP routing, multi-agent orchestration, cron jobs, and access controls.
All-in-one Web Agent framework for post-training. Start building with a few clicks!
CI for Claude Skills — lint, eval, and regression-test SKILL.md files across a model matrix, with a self-growing eval lo
It is a comprehensive resource hub compiling all LLM papers accepted at the International Conference on Learning Represe
大型語言模型(LLM)的讀書會內容code部分整理,此github提供簡單、免費、易實作code範例,並由淺入深帶你一步步了解如何基於python建置LLM流程、部屬功能等
Turn any file into a beautiful, interactive, shareable HTML — WhatsApp logs, PDFs, transcripts, code, anything.
Platform for stateful agents: AI with advanced memory that can learn and self-improve over time.
OrcaLoca: An LLM Agent Framework for Software Issue Localization [ICML 25]
Yunjue Agent: A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frame