4 packages found
MCP server for LLM quantization. Compress any model to GGUF/GPTQ/AWQ in one tool call. First MCP server for model compre
Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server, 65 tok/s Qwen 3.5 122B,
The highest-scoring AI memory system ever benchmarked that isn't reliant on LLM reranking. And it's free & burns less to