Are you the author? Sign in to claim
RuvNet Brain — a downloadable, source-grounded brain for Claude Code over Reuven Cohen's (rUv's) RuvNet stack: RuVector/

A portable, source-grounded brain over Reuven Cohen's (rUv's) RuvNet stack — delivered as a Claude Code plugin that makes Claude use the stack instead of fighting it.
🧭 North Star
rUv has built dozens of genuinely powerful capabilities that are effectively invisible — not undocumented, but undiscovered. Ruflo's own docs call cross-project IPFS pattern transfer "the substrate plugin's most underused capability." The author knows people can't find his own work.
This project exists to close that gap — a CTO on your shoulder that says "you have this, it's off, here's what turning it on buys you."
Retrieval was never the product. Proactive capability advocacy is the product.
The test we hold ourselves to: a developer who solves a hard problem with this tool and later learns they'd been sitting on a capability that would have made it trivial — that is a failure of this project, not of the user. Knowing which question to ask is the scarce thing; supplying that question is the job.
Every feature is judged against this. If a surface can detect something useful and doesn't volunteer it, it is broken — see ADR-027 and DDD-0004.
Three independent things version separately here — by design, not drift. Headline claims are regenerated and checked by the claims ledger (
scripts/claims-verify.mjs); other numbers below are hand-stamped and dated:
plugin(badge above) — the Claude Code plugin itself: SKILL.md, the grounding hooks, the MCP server. Read live fromplugin/.claude-plugin/plugin.json. Updates often — this is where behavior fixes land.installer (npm)(badge above) — thenpx ruvnet-brainsetup script. Read live from the npm registry. Only moves when the installer script itself changes — rare.- Brain Release (the downloadable knowledge bundle, linked from the "download" badge above) — always resolves to
releases/latest(the nightly publishes fresh bundles as the corpus grows). Only moves when the underlying knowledge base is rebuilt — separate again from the two above.- On an old version? One line makes you current — and, with
--auto, keeps you current forever:hljs language-sqlnpx ruvnet-brain@latest --update --auto--updatepulls the latest plugin + knowledge (backs up first, re-verifies, fails loud instead of half-applying). Adding--autoenrolls you in Evergreen — the brain keeps itself up to date from then on, so you never run this again. Drop--autofor a one-time update:npx ruvnet-brain@latest --update.- Turn Evergreen off any time:
npx ruvnet-brain --disable-nightly(Linux/Windows get the cron line documented in the bundle'sforge-update.mjs).- Your copy only advances when a new Release is published — the updater pulls
releases/latest, so running it between releases is a safe no-op.
Built by Stuart Kerr at Isovision.ai · free & fair use, to help everyone leverage the high end of agentic coding.
An interactive, animated walkthrough of what you can actually build — click the preview to open it.
Building toward L4/L5 (3.9.x, dev). The mechanisms for the top two rungs of the proactivity
ladder are built and wired — but they are not yet verified to 4.0's bar, which requires all five
of ADR-028's test classes green under independent grading. Until then, ADR-028
grades the shipped, verified state at L2–L3, and -dev is the honest version. 4.0 is claimed
only when L3+L4+L5 are each measured, not when the code merely exists.
null until enough offers resolve, never a fabricated
score. A dismissal is evidence about fit, not importance — asymmetric budget: a suggestion dies
on one dismissal, an important finding needs three. What's still owed for L5: real precision data
from use, and the cross-project promotion test.Shipped 2026-07-22. Every gate in this project fires on an action — a write, a push, a claim. That is what makes it enforceable. But stopping is the absence of an action, so the most expensive failure of the night had no trigger at all, and a system built to prevent it did not.
continuation-gate.mjs runs when a turn ends. If work you committed to is unfinished, it
makes that the last thing in context. It cannot force another turn — claiming otherwise would be
the fabrication this project exists to kill — but it removes the silence that let a stop pass
unnoticed.Shipped 2026-07-22. 3.6 made the machine legible. 3.7 closes the loop: a lesson you taught is read by a gate at the moment it applies, and — once you ratify it — refuses the action.
lesson-gate.mjs — a gate asks the store "what applies right now?" and gets your own words
back, with the evidence and how many times you had to say it. Wired into the real pre-push gate.lesson-ratify.mjs — see every lesson, ratify it, or delete it. Demotion is sticky: it
survives every future mining run, because a reject the nightly quietly undoes is worse than none.Shipped 2026-07-22. 3.5 made the brain speak. 3.6 makes it legible: a console panel that answers the question the owner has been asking for weeks — "what is actually turned on?"
UNKNOWN is a first-class state and is visually separated from OFF in a
non-colour channel, because collapsing the two is precisely the lie this release exists to kill.default ∉ escalates is asserted as a test — so the rule
holds for settings added later instead of relying on a comment nobody re-reads.ruflo hooks list renders a table column keyed enabled
against a payload that has no such key, so it prints "No" 26 times. The learner had 457
trajectories and had adapted 106 minutes earlier. We scraped a human-readable table instead of
reading state, which is the exact mistake this project keeps writing rules about. The detector now
reads the learner's own state file and stays silent when it cannot tell — ADR-028 fixes the
false-alarm rate at zero and calls it non-negotiable.capability-registry.mjs
had zero call sites; the client referenced it only in comments. Built-tested-unwired is this
project's signature failure, and it happened three times in one night.Honest limit: this is still a dev release, not 4.0. 4.0 requires levels 3–5 of ADR-028's ladder
(contextual, anticipatory, compounding) shipped and independently graded ≥95. The mechanisms now
exist — L3's delivery seam (the chokepoint), an L4 goal surface, and L5 cross-project promotion (lessons
promoted to your global brain, and survival across --update is now tested rather than argued —
tests/unit/promoted-lessons-survive-update.test.mjs runs the real updater against a sandboxed HOME
and asserts the promoted block is byte-identical afterward) — but the ladder still grades L2–L3,
because L3 is built rather than live (deploy-gated) and the five acceptance metrics aren't all measured.
See docs/4.0-READINESS.md and docs/4.0-EXECUTIVE-BRIEFING.md (last independent grade: 70/100 overall).
Shipped 2026-07-22. For three weeks this thing had indexed rUv's entire learning stack — ReasoningBank, SONA, MoE, the ADR-174 distillation pipeline — and could have answered any question about any of it. It never once said the only sentence that mattered:
"Your learning system is installed, and it is switched off."
Stuart found it himself. The measurement: 1,884 captured events sat undelivered while the learner held 5 trajectories and hadn't trained in six days. Draining the queue took it to 412 in one command. Three weeks of learning had been sitting there, available for the asking, and nobody knew to ask.
That gap is the product. A brain that waits to be asked is a search box with good manners. 3.5 is the version where it speaks first.
ruflo wasn't installed to anyone whose npm prefix wasn't the author's. It answered "nothing to undo" about a database repair for which it held a backup. It reported a distillation that wrote 684 patterns as having done nothing (the check used a read-only connection that can't see another process's WAL). And its installer threw away the stderr that explained every crash — so a missing module and a slow download looked identical. All four found by running the thing rather than trusting it.The honest limit: this is 3.5, not 4.0. The advocacy surface is live and the engine behind it is proven, but score-delta alarms aren't built and ADR-027 has not yet survived its own required adversarial review. A major version marks the release where the product becomes a different thing to the person using it. This is the release where it starts talking.
Shipped 2026-07-17. 3.3 made every card lead with a point. Then Stuart looked at the two pages meant to teach the stack and scored them 55/100: the graphics were mediocre, the one page that should show how the pieces fit was a wall of text, and two diagrams were literally unreadable. 3.4 is the fix — the harness's invisible work, finally drawn, and drawn so you can actually read it.
scripts/check-legibility.mjs) — it measures the effective pixel size in the live page, proven to catch the known-bad case before it was trusted. Every diagram label now clears 12px on a phone.Shipped 2026-07-17. 3.2 made the stack visible. Visible turned out not to be the same as useful: the console could tell you 21 things and leave you no smarter. Stuart, looking at his own product: "I have no idea what message it's supposed to tell me… it seems to be facts without purpose." 3.3 is the answer to that — every card now leads with a verdict, not a census.
tests/init-test fixture), and two were echo strings printing advice. Real exposure: zero. It now says so: "Every rUv tool here resolves to one known version." A card that shows amber for a problem you don't have is worse than no card.FRONTIER — 0 read like missing data. It's the punchline: the expensive model never had to fire. Now it says that, in that order.Shipped 2026-07-16. rUv's tools do their best work invisibly — which meant nobody could see them working, working stale, or working in conflict. The 3.x line makes the machinery visible, and everything it shows you is measured, never projected.
/rvbc puts your whole stack on one page: what's installed, how it's wired, what your AI learned. Every warning arrives paired with a one-click, undoable fix, and the page re-checks itself after every change so you always see the after state.
Open it any time with /rvbc (RuvNet Brain Console) — and the first time you load the Brain, it offers to open it for you.
New gate with this release: narrative versions are tested. If any public page says "What's new in X" where X isn't the shipping version, CI fails — because this README sat on 2.5 while 3.1 shipped, and nobody's eyes are a gate.
Shipped 2026-07-13. Two hard lessons, both fixed at the root.
The 2.4 router was hand-rolled: our own heuristic with a placeholder policy, presented as "the MetaHarness router." rUv had already shipped the real thing. 2.5 replaces it with @metaharness/router — his actual learned cost-optimal router (k-NN over labelled embeddings, the productized DRACO Phase-2 finding, ADR-040/043).
Our code now does one honest job: a price transform. A model your subscription already covers becomes costPerMTok = 0, and rUv's router does the rest natively — a $0 model that clears the quality bar simply is the cheapest sufficient candidate. Cost-optimal routing and "already paid for" compose; they never competed.
npm run substitution:check— a new CI gate that fails the build if any code implements a capability rUv already ships and wears his name without either using the real package or openly disclosing the hand-roll. Verified to fire on the exact commit where we got this wrong. You may hand-roll. You may never hand-roll silently.
launchctl reports exit 0 for a job that has never run — byte-identical to success. So "check the exit code" cannot tell triumph from total absence, and a nightly job sat unfired for its entire life while every surface read green.
config/scheduled-jobs.json) of what must run — so unloaded, deleted, or never-fired is a violation, not a silence.Result: 11 jobs supervised, every one producing a fresh successful receipt. One had been totally blind — writing zero bytes on a healthy day, so "ran fine" and "never ran" were indistinguishable. Cured without changing a line of its logic.
A subagent inherits your session's model unless something says otherwise. Ten agents on an Opus session are ten Opus agents; on a Fable session that's $10/$50 per Mtok — up to 10× what the same mechanical work costs on Haiku. That single default was the biggest cost leak in the harness, and an advisory rule did not fix it (the router's entire first life saved $0.018).
So 2.5.1 makes it a wall, not advice: a PreToolUse gate that blocks any subagent dispatch that doesn't declare a model, tells you which tier the task actually needs, and logs every allowed dispatch so routing is auditable rather than merely claimed. Forks still inherit — that's what a fork is.
npm run falsify— the adversary. Every question that had to be asked of this project ("is the nightly actually running?", "is that really rUv's code?", "why is my quota still burning?", "is CI actually green?") is now a check that fails on an unproven claim, not merely on broken code. Because tests you wrote yourself passing is circular evidence.
2.0 is the release where the brain got bigger — and, more importantly, stopped taking its own word for anything. Every number below regenerates from an artifact on disk; the claims ledger (node scripts/claims-verify.mjs) re-checks the advertised ones in CI:
| v1 (0.x–1.x) | v2.0 | |
|---|---|---|
| Corpus | 24 repos built | 69 repos built (of 248 in the org), each verified by a live retrieval query |
Depth (flagship ruvector) | 18,491 passages · 0 full source bodies | 28,018 passages · 2,996 full bodies — depth also restored to agent-harness-generator (8,896/715), ruview (7,434/765), open-claude-code (195/69) |
| Corpus QA gate | none | every store must prove embeds correctly + reads correctly — vector count == passage count, depth floors, a 3-passage self-retrieval round-trip per store — 72/72 store-variants PASS, wired fail-closed into the nightly publish |
| Retrieval eval | 12 frozen questions | 120 frozen, hash-pinned questions across 5 strata; promotion gated on Wilson lower bounds, fail-closed — it blocked a real release this morning, which is the feature working |
| Token cost | 6,183 bytes injected per hook turn · zero self-measurement | 684 bytes (~90% cut) with an eval-PASS proving zero quality loss · a live token meter measuring real bytes/tokens per prompt class · ~27% faster repeat queries via the KB cache |
| rUv's gists | not indexed | 444 gists indexed with per-chunk freshness/provenance banners, refreshed nightly with cost-disciplined skip |
| Reliability | claims were prose | claims ledger (6 marketing claims mechanically re-verified in CI) · integration tests in CI incl. Linux · a Windows CI job · an honest coverage denominator (all source files) |
| Autonomy | hooks asked questions to an empty room — the #1 real-user complaint | /loop contract — checkpoint / resume / done-criteria; hooks detect autonomous mode and stop asking — 10 mutation-verified tests |
| Publishing | manual npm token · a nightly release PATH bug | self-renewing npm token (launchd daemon, proven end-to-end) · PATH bug root-caused and cured |
| Memory layer | silent SQLite corruption · searches returning 0 · invisible flywheel patterns | 3 root-caused fixes (an ABI-mismatched binary falling back to WAL-blind whole-file overwrites; keys that were never scored; a split-store display bug) — plus 6 exact patches queued upstream to ruflo |
The depth jump wasn't tuning — it was two pipeline root-causes fixed for good: missing --full hints, and a day-one v2/ skip-dir bug that had made two repos undepthable since the first build.
| Dimension | v1 (2026-07-09) | v2.0 (2026-07-10) | Δ | What moved it |
|---|---|---|---|---|
| End-user experience | 54 | 83 | +29 | One-command install now offers nightly self-updates (default yes); the page lives on isovision.ai; publishing renews itself |
| Knowledge corpus | 71 | 88 | +17 | 24→69 verified repos; a 72/72 embeds-and-reads QA gate; full source depth restored — flagship went 0→2,996 source bodies |
| Effectiveness (eval-proven retrieval) | 58 | 88 | +30 | 120-question Wilson-bound gate — it blocked a bad release, then passed the fix above the old baseline |
| Acting like rUv | 38 | 72 | +34 | Memory layer root-caused and fixed with proofs; 6 exact patches queued upstream; real multi-agent swarm operations |
| Developer smarter | 62 | 84 | +22 | /brain-score runs this same scorecard on any repo; honest tool announcements; per-answer source receipts |
| Token cost efficiency | 41 | 82 | +41 | ~90% smaller per-turn injection, eval-proven free; a live token meter — measured, not guessed |
| Engineering reliability | 61 | 86 | +25 | 312 tests green; a claims ledger re-verifies marketing claims in CI; fail-closed publishing; Windows CI added |
| Safety & privacy | 74 | 82 | +8 | Private-store fence held under full rebuild; secrets never transit chat; autonomy hard fence |
| Overall | 55 | 83 | +28 | — |
These are self-scores under a hard rule: every deduction requires specific evidence, and a known architectural flaw caps a dimension at ≤70 until the flaw is fixed — acting like rUv spent the morning capped for its memory-layer flaw; the flaw was root-caused and fixed with proofs, the cap lifted, and it re-scored 72. Overall 55 → 83 in two days, each number regenerating from a stored receipt — and the same scoring that produced these blocked a release mid-day. That's why they're credible: scores you can trust beat scores that flatter.
The honest small print, kept visible: warm query is ~21 s on the large corpus (it grew with depth; candidate dedup/pruning is the next optimization) · the 8 newest repos are findable by name but don't yet have primers/capability cards for described-need routing (coming) · the Windows CI job is new and unproven until its first green run.
Reuven Cohen (rUv) builds about nine months ahead of the state of the art. His RuvNet building blocks — Ruflo, RuVector, AgentDB, agentic-flow, SPARC and ~20 more — are the working prototypes of what becomes mainstream AI tooling three quarters later. This is the actual front edge.
But Claude was trained on classical software development. Point it at rUv's stack and it doesn't recognize the work: it drifts, it doubts real capabilities (“ruflo can't edit files,” “RuvNet has no vector DB — use Pinecone”), and it quietly falls back to the patterns it knows (pgvector, LangChain, a hand-rolled cosine loop). A newcomer gets the worst of both worlds — a revolutionary toolset with no instruction manual, and an assistant that talks them out of using it.
RuvNet Brain is the missing instruction manual. It reads rUv's real source, hands Claude the answer key, and removes Claude's permission to make things up about the stack. Install it once, aim it at any repo, and a newcomer can build ~9 months ahead — without being rUv.
The novelty is structural grounding, not plain retrieval. Plain RAG only decides what to add to context. This ships a UserPromptSubmit hook that injects a grounding directive on every RuvNet-relevant turn — the harness consumes that stdout structurally, so the directive is always present, not a decline-able suggestion. It's a strong, always-on nudge — Claude is pointed at the real source and told to ground before asserting on every relevant turn — not a hard block on the model's output. RAG decides what to add; this makes grounding the default the model has to actively argue its way out of.
You know the failure mode. You ask Claude to build with Ruflo or RuVector, and instead of reading rUv's actual code it reaches for its training priors: “let's just use pgvector,” “I'll hand-roll cosine similarity,” “I don't think ruflo can edit files.” It skims, it guesses, it doubts tools that work perfectly well — and the result drifts off the very stack you chose.
WITHOUT the brain — drift | WITH RuvNet Brain — grounded
|
You: build with Ruflo/RuVector | You: the same request
| | |
v | v
Claude falls back to priors | Enforcement hook grounds the turn
| | |
v | v
"just use pgvector" | search_ruvnet -> cited rUv source
hand-rolls cosine similarity | (whole files, labeled repo + path)
"ruflo can't edit files" | |
| | v
v | Builds ON the stack (RVF/HNSW, swarms)
Drifts OFF the stack |
npx ruvnet-brain
That single command runs the whole setup, narrating what it's doing and why at each step: it downloads the brain from the latest GitHub Release, unpacks it to ~/.cache/ruvnet-brain/kb, installs its local reader (no cloud calls, no API keys), and wires the Claude Code plugin — the search_ruvnet MCP tool + the UserPromptSubmit grounding hook — at user scope. The installer and the search_ruvnet tool run on macOS, Linux, and Windows; the grounding/enforcement hooks are POSIX shell, so they fire on macOS, Linux, and Windows-via-WSL/Git-Bash (on native Windows without WSL the search tool still works, but the auto-grounding hooks don't fire — a Node port of the hooks is on the roadmap). It's safe to re-run, and the brain itself always fetches the current Release regardless of which install path you use — you install once; you don't keep re-downloading.
Want the bleeding-edge installer, even ahead of the last npm publish?
npx github:stuinfla/ruvnet-brainalways runs straight off the latest GitHub commit.
Anonymous usage counts — opt-in, counts only. At the end of the install you're asked once: "Share anonymous usage counts (installs/searches — never your queries or code)?" If you say yes, the only things ever transmitted are event counts (
install/search/session) plus the bundle version — batched to at most one ping per machine per day. Never your queries, code, repo names, or paths; the entire client is one readable file (kb/telemetry-ping.mjs). Your answer is a plain-text file you can read or flip any time (~/.cache/ruvnet-brain/.telemetry-consent), andnpx ruvnet-brain --no-telemetrydeclines without being asked. Found it useful? Star the repo or leave feedback in Discussions.
claude plugin marketplace add stuinfla/ruvnet-brain
claude plugin install ruvnet-brain@ruvnet-brain --scope user
Registers the search_ruvnet MCP tool, the grounding skill, and the UserPromptSubmit enforcement hook — globally, at user scope. The plugin expects the brain at ~/.cache/ruvnet-brain/kb (or point RUVNET_BRAIN_KB at your own copy). The first install may show a one-time trust prompt for the hook.
Then just ask. “How does Ruflo orchestrate agent swarms, and what implements it?” The hook grounds the turn, Claude calls search_ruvnet, and it answers from cited source — down to the function body.
You install once. After that, three mechanisms keep you on the current brain without you having to remember an update command.
Consent-gated auto-update heartbeat (the SessionStart hook, plugin/scripts/session-start.sh). The first time the plugin runs on a machine it asks you once whether it may keep itself updated in the background — a security-conscious opt-in, because self-update can change the model's own instructions. Your answer is remembered (~/.cache/ruvnet-brain/.auto-update-pref) and never asked again. On each session start it does a rate-limited (~15 min) 3s-capped check of the live GitHub plugin.json. If a newer plugin version exists and you opted in, it downloads it in the background through Claude Code's own trusted marketplace path — but the new version is staged, not active: Claude Code only loads plugins at process start, so this session keeps running the version it started with until you restart (claude --continue brings your conversation right back on the new version). If you declined, it just tells you the command to run. The knowledge bundle is handled more conservatively — detect + notify only, never auto-applied, because the bundle isn't cryptographically signed yet and applying it would overwrite executable tool files (SEC-0010 #6).
Grounding receipt line (the UserPromptSubmit gate, plugin/scripts/ground-ruvnet.sh). When the brain engages on a prompt, the answer ends with one dim line stating what it actually did — either it read rUv's real source and names the file, or it says plainly that it didn't:
🧠 RuvNet Brain jumped in · cited agentic-flow/docs/adr/ADR-076-reposition-agentic-flow-as-agentic-meta-harness.md · v3.4.18-dev
🧠 RuvNet Brain jumped in · guidance only, no source read · v3.4.18-dev
An unearned citation is worse than no citation, so the line may only name a path the tools genuinely returned — and on a prompt where nothing fires, it stays silent rather than manufacture a receipt. The version shown is the one actually loaded in memory for this session; if a newer one is staged awaiting a restart, the line says so plainly (… vX staged, restart to load). So you never have to wonder whether the brain is on, which version is acting, or whether an answer was grounded or guessed.
Nightly publish → releases/latest chain (scripts/self-update.mjs --publish, run by the deploy/com.ruvnet.brain-nightly.plist LaunchAgent at 03:15). The nightly rebuilds only the repos whose upstream changed, and if anything was rebuilt it bumps the product version, cuts a GitHub Release, and advances releases/latest. Plugin and knowledge bundle move under one version number, so the heartbeat above picks up both automatically. (The LaunchAgent is not auto-installed — enabling a system scheduler needs explicit owner approval.)
Earlier bundles knew only the docs and architecture. v0.5 began re-indexing the code-rich repos to full function bodies, against each repo's real source layout — so “how is this actually implemented?” returns the implementation, not a summary. The table below is that v0.5 depth jump, kept as the before/after receipt; 2.0 went further still — flagship ruvector alone now carries 28,018 passages and 2,996 full bodies (see the "what 2.0 proved" foldout near the top):
| Repo | Full-body code passages (v0.5) | |
|---|---|---|
agentic-flow | 296 → 984 | model-routing, ReasoningBank |
ruv-fann | 52 → 779 | ruv-swarm, cuda-wasm, neuro-divergent |
qudag | → 531 | post-quantum crypto core, exchange, MCP |
daa | → 266 | orchestrator, economy, rules, swarm |
safla | 0 → 136 | the metacognitive / self-modification layer |
Plus: the “take the wheel” behavioral pipeline (below), a 4-level behavioral test harness, a build-integrity fix (empty files no longer pollute the index), and the private-store fence that keeps unpublished repos out of the public bundle.
The expensive work happens once, at build time: every covered repo is deep-walked (whole files, full function bodies, plus a symbol index), embedded into two vector variants (MiniLM-384 for edge/portability, bge-768 for depth) stored on-disk in RVF / HNSW, and distilled into a concepts + capability layer of per-repo primers and cards. That's 149,930 source chunks. At query time, search_ruvnet searches every repo's store at once, pools the hits, and runs them through one cross-encoder rerank on a common scale — so the truly relevant file wins regardless of which repo it lives in — then returns whole source files, each labeled by repo and path.
Grounding is injected every turn, not left to chance. On a RuvNet-relevant prompt the UserPromptSubmit hook injects a directive into context; Claude calls search_ruvnet, gets whole source files labeled by repo and path, and answers from them. Because the hook's stdout is consumed by the harness every turn, the directive is always there — a strong grounding nudge on every relevant turn. (It's a nudge, not a hard gate: the hook can't rewrite Claude's output, so it steers rather than blocks — see ADR-0005 for exactly what ships.)
And on a build request, the brain doesn't wait to be told each step — it runs the process the way Ruv would:
Propose the architecture first → one go/no-go → then execute end-to-end: SPARC (Spec → Pseudocode → Architecture → Refinement → Completion) · model the domain (DDD) and capture ADRs · spin up parallel Ruflo swarms · persist decisions to AgentDB · treat design as a build step (frontend-design + AI image generation, never “working but ugly”) · test → score 1–100 → loop to ≥98 · and it asks for an API key when a step needs one instead of silently skipping it.
The downstream effect: Claude prefers RuvNet building blocks over generic defaults (RVF/HNSW over pgvector/Pinecone, Ruflo swarms over hand-rolled orchestration) and drives the whole assess → build → verify → score loop rather than answering one question at a time.
The brain answers both kinds of questions. Name the repo or ask something specific and it resolves to the right repo (47/48, 98%). Describe a need without naming the repo — the way a newcomer would — and it still lands the right repo (26/28, 93%). That newcomer path used to be the weak spot (33% before the fix); adding capability cards — a capability-phrased passage per building block — closed the gap.
69 of rUv's repos in the ruvnet org — the reusable building blocks you'd actually compose into a system — each deep-walked and embedded in both variants. The core blocks below also carry symbol indexes and capability cards (the 8 newest repos are findable by name; their capability cards are coming).
| Repo | What it gives you | Repo | What it gives you |
|---|---|---|---|
ruflo | Agent orchestration / swarms | ruvector | RVF + on-disk HNSW vectors |
agentdb | Agent memory + graph/Cypher | rulake | Vector cache layer |
ruview | Camera-free WiFi/CSI sensing (presence, pose, vitals) | agentic-flow | Cheap multi-provider model routing |
agenticow | Copy-on-write agent memory (branch 1M vectors in ~162 B) | sparc | 5-phase build methodology |
qudag | Quantum-resistant DAG messaging | safla | Self-aware feedback loop |
ruv-fann | Fast neural nets (Rust/WASM) + ruv-swarm | daa | Decentralized autonomous agents |
synthlang | Prompt compression (~75% token cut) | rupixel | On-device visual embeddings |
dspy.ts | DSPy-style programmable LLM pipelines in TS | fact | Fast-Access Cached Tools + circuit breaker |
cve-bench | Security-fix benchmark | metaharness | Harness scaffolding / metaharness |
rvm | Proof-gated capability microhypervisor | rUv-dev · open-claude-code | Dev workflow + agent tooling |
Not in the public brain: rUv's private Cognitum One repos (seed, v0-appliance, platform-docs) are fenced out of the download by design — verified zero-leak in every build. Helix (rUv's local-first health app) is a finished product, not a building block, so it's out too.
Everything below is re-runnable — the proof is the output of a command, not a claim.
bash scripts/gate.sh # the pass/fail routing gate
node scripts/behavioral-l1-l4.mjs # the 4-level behavioral harness
node plugin/test/run-tests.mjs # full plugin QA over real JSON-RPC
| Test | Result | What it proves |
|---|---|---|
| Named / specific routing | 47 / 48 (98%) | name the tool → right repo |
| Described-need routing | 26 / 28 (93%) | describe the need, no name → right repo (was 33% before capability cards) |
| Context-scenario routing | 7 / 8 (88%) | full-scenario prompts route correctly |
| L1–L4 behavioral harness | all pass | route · deep-recall (returns code) · implement (cites the API) · orchestrate (the hook drives the full pipeline) |
| Plugin QA | 26 / 26 | manifests, hook firing, MCP initialize/tools/list, capability battery |
| Clean-room install | 3 / 3 | download the published bundle fresh → unzip → query → grounded, cited answers |
| Unit tests | 548 passing, 169 todo · 14% of ALL source covered | npm run test:cov regenerates both — the coverage floor fails CI if it slips (claims:verify re-derives the %, it is not a hand-typed badge). 14% is the honest number over every shipped file; the previous "75%" measured a hand-picked 8-file subset |
| Grounding proof | npx ruvnet-brain --doctor | asks a real question, then checks the cited path really exists in the on-disk store; a citation that doesn't resolve is reported as NOT grounded |
| Held-out eval | grounded 100/100 · routed 63/80 | npm run eval — 120 frozen, hash-pinned questions across 5 strata, never used for tuning, graded on ground truth, never by a model |
The suite also carries 169 it.todo stubs — a written backlog, each naming an untested behavior and what it would take to cover. They are deliberately not counted as tests: a stub proves nothing, and a number that flatters is worse than no number.
Two honest residuals, not hidden: one described question (“route to cheaper models to cut cost”) routes to open-claude-code instead of agentic-flow; one “methodology” question routes to agent-harness-generator (in PROOF.md) or cognitum-cogs (in HELIX-DEMO-NOHELIX.md) instead of sparc/ruflo. Proof reports land in PROOF.md, DESCRIBED-PROOF.md, and HELIX-DEMO-NOHELIX.md.
Most of the numbers above were tuned against. evals/held-out.json was not: 120 frozen, hash-pinned questions across 5 strata (named, described, scenario, adversarial, provenance), phrased the way a newcomer would ask, each with the owning repo chosen from first principles before the brain ever saw them. npm run eval scores, among other things, two claims — both without asking a model's opinion:
npm run eval:gate fails the build (exit 1) if either score drops below evals/baseline.json — and also if there is no baseline at all, since you cannot promote against nothing. Baselines are only ever written deliberately, with npm run eval:record; a baseline that silently follows the code is a ratchet with no teeth.
Why not let an LLM grade it? Because on this very repo an LLM panel scored a zero-citation answer 98/100. Model-as-judge is blind to the one failure that matters here.
Current baseline (n=120, evals/baseline.json): grounded 100/100 · routed 63/80 · abstain 18/20 · banner 20/20, promotion gated on Wilson lower bounds. The routing misses are recorded, not tuned away — the recurring shapes: "generate tests and find coverage gaps" cites agentic-qe's real test generator, which lives vendored inside the ruflo repo, so the answer is right and the repo label is a corpus-attribution artifact; "spend less money on model calls" cites a genuine agentic-qe/docs/guides/cheaper-model-eval-lanes.md, the same cost/orchestration overlap noted above.
Query it directly (CLI):
cd kb
export KB_MODEL_CACHE=/path/to/models-cache
node forge-ask-all.mjs --dir . --q "How does RuVector implement HNSW vector search?" --k 3
Cross-repo answer quality requires the cross-encoder, which requires
npm iin the bundle. Without it, search falls back to raw vectors and ranks poorly.
This project versions in the open (see the live badge up top for the exact plugin version; the downloadable knowledge bundle is a separate track) — we don't claim “done,” “complete,” or “zero hallucinations.” Where it stands:
npx ruvnet-brain (short form); npx github:stuinfla/ruvnet-brain always tracks the latest commit if you want it even fresher.kb/ — the brain: per-repo .rvf + .big.rvf stores, full-passage sidecars, symbol indexes, per-repo primers, the concepts store, and the forge-* query tools (CLI + MCP search_ruvnet).plugin/ — the Claude Code plugin (MCP server, grounding skill, UserPromptSubmit enforcement hook, marketplace manifest, test suite).scripts/ — gate.sh (routing gate), behavioral-l1-l4.mjs (behavioral harness), prove.mjs, build-bundle.mjs, brain-stamp.mjs.docs/ — VISION.md (the why), adr/ (locked decisions incl. ADR-0008), DDD.md.explainer/ — the source of the live explainer.SPEC.md · PROGRESS.md — the master spec and the living, timestamped build log.The brain binaries ship via the Release, not git — a fresh clone is lightweight; npx fetches the full bundle.
npx ruvnet-brain --feedback prefills a Discussion with your version + a 3-line health summary (you see exactly what's in it; never your queries, code, or paths) and opens it in your browser.CONTRIBUTORS.md — the two-signal hook gate and the /brain-build contract both started as user field reports. House rules: CODE_OF_CONDUCT.md; build/test map: CONTRIBUTING.md.Run Claude Code as an MCP server so any agent can delegate coding tasks to it
Browser automation using accessibility snapshots instead of screenshots
Google's universal MCP server supporting PostgreSQL, MySQL, MongoDB, Redis, and 10+ databases
Official GitHub integration for repos, issues, PRs, and CI/CD workflows