13 packages found
Production-tested commands, skills, and workflow patterns for Claude Code. Developed through 6+ months of daily use. Inc
A curated list of Generative AI tools, works, models, and references
HumanStudy-Bench: Towards AI Agent Design for Participant Simulation
All-in-one Web Agent framework for post-training. Start building with a few clicks!
CivAgent is an LLM-based Human-like Agent acting as a Digital Player within the Strategy Game Unciv.
[NeurIPS 2024 D&B] GTA: A Benchmark for General Tool Agents & [arXiv 2026] GTA-2
🪢 Open source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Int
Awesome papers involving LLMs in Social Science.
LLM-as-judge evaluation framework for assessing AI agent output quality
一个生产级的深度研究 Agent 系统,从零构建多智能体编排、Red-Blue 对抗降噪、 语义级上下文压缩、跨 Agent 共享记忆四大核心能力,配套 165 次独立实验 + Bootstrap 统计显著性检验的完整评测体系。
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE
LLM QA, Observability, Evaluation and User Feedback