Are you the author? Sign in to claim
MCP server that turns a folder of PDFs and Markdown into an AI-friendly document shelf. Convert, split by chapter, auto-
Put your manuals on a shelf, hand the AI the index.
📖 Docs & landing page: https://ignatenkofi.github.io/docshelf-mcp/
_ _ _ __
__| | ___ ___ ___| |__ ___| |/ _|
/ _` |/ _ \ / __/ __| '_ \ / _ \ | |_
| (_| | (_) | (__\__ \ | | | __/ | _|
\__,_|\___/ \___|___/_| |_|\___|_|_|
MCP server for AI-friendly doc shelves
An MCP server that turns a folder of PDFs and Markdown into a chat-project-friendly document collection: AI agents see a single INDEX.md and pull individual sections by raw GitHub URL on demand — instead of choking on a 4 MB datasheet.
You have 30 hardware manuals, or 200 cooking recipes, or a stack of research PDFs.
You want Claude / ChatGPT / whatever to be able to answer questions across them — but:
docshelf-mcp solves it like this:
INDEX.md.INDEX.md to your Claude project. When the model needs a section, it fetches it via raw.githubusercontent.com.Result: a 5 KB index pointing at a 50 MB collection. The model reads exactly the chapter it needs.
From PyPI (once the first tagged release is published):
# uv (recommended)
uv pip install docshelf-mcp
# or plain pip
pip install docshelf-mcp
Or straight from main (always-latest, no PyPI required):
pip install "git+https://github.com/ignatenkofi/docshelf-mcp"
Optional high-quality PDF engine (pulls ~2 GB of PyTorch — only if you need it):
pip install "docshelf-mcp[high-quality]"
Optional input formats beyond PDF/Markdown — DOCX, HTML, EPUB (lightweight):
pip install "docshelf-mcp[formats]" # or [docx] / [html] / [epub]
Drop this into the Custom Instructions of any Claude project that consumes
a docshelf-style INDEX.md:
This project uses the docshelf pattern.
INDEX.mdis the entry point. When answering: read INDEX → fetch ONLY the needed section file via its GitHub raw URL (use WebFetch / fetch / curl). Don't load full source files into context. For large manuals split into chapters, follow INDEX → chapter SUBINDEX → section file.
Medium (~150 words) and full (~400 words) versions, plus how-to snippets for
Claude Code, Claude Desktop, and the Anthropic API, live in
docs/PROJECT_PROMPT.md.
from docshelf_mcp import Shelf
shelf = Shelf("~/Documents/my-homelab-docs").init(
name="My HomeLab Docs",
remote="https://github.com/me/my-homelab-docs",
default_categories=["routers", "switches", "psu", "motherboards"],
)
shelf.add_document(
"~/Downloads/MIKROTIK_RouterOS.pdf",
category="routers",
title="Mikrotik RouterOS — full manual",
description="Official RouterOS reference, split by chapter.",
)
# → docs/routers/mikrotik-routeros-full-manual.md + docs/routers/.../001-..md, 002-..md, ...
# → INDEX.md is regenerated automatically.
Then in the shelf directory: git add . && git commit -m "docs: add RouterOS" && git push.
In your Claude project, attach only INDEX.md. Done.
Edit ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%/Claude/claude_desktop_config.json (Windows):
{
"mcpServers": {
"docshelf": {
"command": "docshelf-mcp",
"env": {
"DOCSHELF_ROOT": "/Users/me/Documents/my-homelab-docs"
}
}
}
}
Restart Claude Desktop. You now have eleven new tools available:
| Tool | What it does |
|---|---|
docshelf_init_shelf | Bootstrap a new shelf directory. |
docshelf_add_document | Add a file (MD/PDF/DOCX/HTML/EPUB). Converts, splits, re-indexes. |
docshelf_add_directory | Add every supported file (MD/PDF/DOCX/HTML/EPUB) in a folder in one call. Re-indexes once. |
docshelf_read_document | Read a document/section's content over MCP (works on private shelves). |
docshelf_remove_document | Remove a document, its sections, and metadata. Re-indexes. |
docshelf_rename_document | Retitle / recategorize a document (moves file, sections, meta) — no re-conversion. |
docshelf_rebuild_index | Regenerate INDEX.md from disk. |
docshelf_doctor | Check shelf integrity; optionally auto-fix safe drift. |
docshelf_search | Plain-text search across the shelf, with raw URLs. |
docshelf_list_documents | List documents by category. |
docshelf_convert_pdf | Standalone PDF → Markdown (no shelf). |
The shelf files are also exposed as read-only MCP resources, so a client can browse and attach them natively — see MCP Resources below.
claude mcp add docshelf -- docshelf-mcp
# Optional: set the default shelf
claude mcp add docshelf --env DOCSHELF_ROOT=/path/to/shelf -- docshelf-mcp
# Sanity check — should print the server version then wait on stdin
docshelf-mcp
Alongside the tools, every shelf file is exposed as a read-only MCP resource, so an MCP client (Claude Desktop, Claude Code, …) can browse and attach shelf content natively — no tool call required.
docshelf:///<relative-path>, e.g. docshelf:///INDEX.md or docshelf:///docs/routers/mikrotik/003-firewall.md.INDEX.md plus every document and every split section under docs/ — one resource each. A split document exposes both its whole-file parent and its individual section files.docshelf_read_document tool, which pages the rest.add_document, add_directory, remove_document, rename_document, rebuild_index, init_shelf) — so newly added files appear and removed ones drop out. Reads are confined to the shelf root.Resources are only registered for an initialized shelf (one that has a .docshelf.json); a non-shelf DOCSHELF_ROOT simply exposes none.
my-shelf/
├── .docshelf.json ← shelf metadata: name, remote, category order
├── INDEX.md ← auto-generated navigation (your chat-project file)
├── .gitignore
└── docs/
├── routers/
│ ├── .meta.json ← per-document title/description overrides
│ ├── mikrotik-routeros.md (full document, lightly cleaned)
│ └── mikrotik-routeros/ (auto-split sections)
│ ├── SUBINDEX.md (per-document navigation page)
│ ├── 001-overview.md
│ ├── 002-bridging.md
│ └── 003-firewall.md
└── switches/
└── cudy-gs1010pe.md
Everything in docs/ is committed; everything is fetchable via raw URL once you push to GitHub.
A document is split when both conditions hold:
.docshelf.json:split_threshold_bytes).## (H2) headings.The splitter:
NNN-<slug>.md so they sort naturally and survive title changes.SUBINDEX.md navigation page into the split directory (title,
description, per-section links) — regenerated on every rebuild_index.In INDEX.md, split documents with up to 10 sections list every section
inline; bigger splits get a single link to their SUBINDEX.md so the index
stays small. Control this via .docshelf.json:
"index_style": "auto" | "inline" | "subindex" and
"subindex_threshold_sections": 10.
If you want to keep a document whole, pass split=False.
See the examples/ directory for three concrete use cases:
examples/homelab/ — original use case, hardware manuals for a home lab.examples/recipes/ — a cookbook with one recipe per file.examples/research-papers/ — academic PDFs with abstracts in .meta.json.Each example shows the directory layout and the INDEX.md you'd end up with.
The default engine (pymupdf4llm) is fast and good enough for ~95% of technical documents. For papers with complex tables, math, or scanned content, install the marker-pdf backend:
pip install "docshelf-mcp[high-quality]"
Then pass quality="high":
shelf.add_document("paper.pdf", category="research", title="...", quality="high")
⚠️ marker-pdf pulls in PyTorch (~2 GB) and is significantly slower (10–60 s per document on CPU). The library import is deferred — if you don't use quality="high", the dependency is never loaded.
Why GitHub raw URLs and not embeddings / RAG? Because it's dead simple, costs nothing to host, and the AI is already good at chasing links. You can layer embedding search on top later if you want — the on-disk shape is a normal git repo.
Does this work with private repos?
Partly. The raw-URL trick needs a public repo — raw.githubusercontent.com won't serve private ones without auth. But docshelf_search and docshelf_read_document both work over MCP on private (or purely local, non-git) shelves: the model searches, then reads the exact section's content directly from the server, no raw URL required. You only lose the ability to hand a bare INDEX.md to a chat project and have it fetch by URL — with the MCP server attached, the full flow works either way. Make the doc repo public if you want the URL-fetch path too.
Do I have to use GitHub?
No. Set provider in .docshelf.json (or at init_shelf): github (default), gitlab, gitea, custom, or none. The github provider also covers GitHub Enterprise Server: a self-hosted github.<company>.com remote gets the GHES raw form (https://<host>/<owner>/<repo>/raw/<branch>/<path>) automatically. custom takes a url_template with {owner}, {repo}, {branch}, {path} placeholders, so you can point at S3, Cloudflare R2, GitLab/Gitea raw, a GHES deployment on a fully custom domain, or any static host — the generated URLs are correct everywhere, no post-processing. none renders relative links in INDEX.md, which stay navigable offline / in a local checkout.
Does it edit the source PDFs?
No. PDFs are converted on add_document and the source is left in place. The shelf only writes inside its own directory.
What about non-English documents?
Slugify is Unicode-aware (NFKD-normalized, with \w under re.UNICODE). Cyrillic / CJK titles slug down to ASCII-ish forms; the body Markdown is preserved as-is.
Can I use it without MCP?
Yes — from docshelf_mcp import Shelf and use the class directly. See docs/USAGE.md.
INDEX.mds.docshelf_search.INDEX.md on disk, but the caller (you, or an agent) is responsible for git add / commit / push. This is intentional — staying out of git's way keeps the tool safe to call from agents.Measured on two real shelves (24 hardware manuals; a full novel split by chapter): answering a question the docshelf way costs ~3.7K tokens vs 1.2M to dump the collection — 99.7% fewer — and the biggest manual (RouterOS, ~1.05M tokens) doesn't even fit in a 200K context window, while a section fetch always does.
📊 Full write-up with the numbers, chart, and a reproducible benchmark:
docs/demo.md (run benchmarks/token_savings.py on your own shelf).
For a deeper dive, see docs/ARCHITECTURE.md — module layout, data flow, design rationale.
INDEX.md in context and recalls sections on demand. Born as RFC-0001 in this repo; uses docshelf as its storage/index layer.Bug reports and PRs welcome. To set up a dev env:
git clone https://github.com/ignatenkofi/docshelf-mcp
cd docshelf-mcp
uv pip install -e ".[dev]"
ruff check src tests
pytest -v
MIT — see LICENSE.
docshelf-mcp started life as a 350-line Python script (homelab-encyclopedia.py) that managed a single homelab manuals repo. The split / index / clean logic is the same code, generalised to work for any category-organised document collection.
Run Claude Code as an MCP server so any agent can delegate coding tasks to it
Browser automation using accessibility snapshots instead of screenshots
Google's universal MCP server supporting PostgreSQL, MySQL, MongoDB, Redis, and 10+ databases
Official GitHub integration for repos, issues, PRs, and CI/CD workflows