440 packages found
A powerful RAG (Retrieval-Augmented Generation) system built with LangChain, designed as an MCP server for Cursor, VS Co
MCP server for Claude Code — Apple Vision OCR & image analysis, fully offline, no API keys
for AI Agents - give your text-only model eyes. Use any OpenAI-compatible vision model API through MCP tools for image
A Model Context Protocol (MCP) server that provides vision capabilities to analyze image and video
Advanced Claude Code CLI toolkit - agents, hooks, skills, MCP servers, phased development, site intelligence dev-scan, a
Claude Code hook toolkit that gives vision-blind models a text description of pasted/tool-produced images.
Awesome LLM Papers and repos on very comprehensive topics.
Give non-multimodal Claude Code main models the ability to see pasted screenshots — a ~200-line UserPromptSubmit hook.
Transform YouTube videos into a compounding knowledge base with transcripts, vision analysis, and agentic search. Works
历年ICLR论文和开源项目合集,包含ICLR2021、ICLR2022、ICLR2023、ICLR2024、ICLR2025.
Private on-device AI suite for Android. Fork of Google AI Edge Gallery with llama.cpp, whisper.cpp, stable-diffusion.cpp
stdio MCP server to read or index markdown with images
The agent-native LLM router for OpenClaw. 41+ models, <1ms routing, USDC payments on Base & Solana via x402.
clawdcursor compiles whatever's on screen into one UI map — accessibility tree and OCR fused into stable, addressable el
One MCP server, 51 tools for AI agents on mobile + canvas UIs: iOS & Android automation, Maestro E2E, evidenced assertio
List of Permanent Free LLM API (API Keys)
Every smart contract deserves intelligence, not just data. MCP server for on-chain calculated indicators (EMA, RSI, VWAP
Undetectable Chrome for AI agents, web scraping, and automation. REST API + MCP server + VNC debugging.
MCP server for accessing reMarkable tablet data - sync files, extract text from highlights, and browse your reMarkable c
RPG Maker MV Ultimate — AI copilot (MCP): generate maps, edit the database & events, and understand your project (valida
💻 A curated list of papers and resources for multi-modal Graphical User Interface (GUI) agents.
Computer Experience Layer — MCP server for native computer use on macOS. Structured perception, reliable actions, contin
Give your AI agent eyes and hands on any desktop — cross-platform accessibility API with MCP server
Fully autonomous AI Agents system capable of performing complex penetration testing tasks
Speaker is a Codex skill project for academic presentations: read real.pptx, combine text extraction, PPTX structure par
🔴 VERY LARGE AI TOOL LIST! 🔴 Curated list of AI Tools - Updated 2026
Run local AI models, search your files and code, and crawl the web, all in one program. Cited answers, local-first, with
Give Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio a
A repo lists papers related to LLM based agent
The open source agent.
MCP server that prevents LLM vision downscaling, tiles large images & screenshots so Claude, GPT-4o, and Gemini see ever
Open-source 3D AI agent framework — GLB/glTF avatars with LLM brains, memory, emotions, and autonomous payments. MCP ser
Write-capable PDF toolkit for any MCP client: 22 tools to read, create, render, encrypt, and transform PDFs. Vision rend
A Security-centric MCP Server providing enterprise-grade filesystem powers to AI assistants—read, write, edit, and manag
SGPT is a command-line tool that provides a convenient way to interact with OpenAI models, enabling users to run queries
OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) wi
Whiteboard tldraw diagrams (.tldr) from natural language with 6 presets and vision-based self-check. PNG/SVG export, mul
MCP server for xAI's Grok API with agentic tool calling, image and video generation, vision, and file support.
MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)
It is a comprehensive resource hub compiling all LLM papers accepted at the International Conference on Learning Represe
Turn a phone flip-through video, photos, or a PDF of a book into a clean, proofread EPUB/PDF/Markdown — local vision-LLM
All-in-one Web Agent framework for post-training. Start building with a few clicks!
Clone any .pptx into your own deck — OpenAI gpt-image-2 mimics the layout, you supply the content. 10 bundled styles. |