2 packages found
Governed GPU inference ops (vLLM + Ray Serve): latency RCA, scaling, drain, 30 MCP tools (preview)
An MCP server that gives a language model a structured, persistent scratchpad for hard thinking — staged reasoning + sel