# MM3 Knows in ~20 ms. Learns in ~500 ms. The memory and decision layer for coding agents. Every answer is citable. [CI] [Apache-2.0 license] [Node 22.13 or newer] [Figure: How MM3 works. MM3: make and model. Six verbs in two bands and three columns. MAK³, make, use what is proven: view is a free lookup of the ledger, class gives one verdict for one subject, replay rechecks after a fix. MDL³, model, learn what is missing: scan sweeps to find where to look, drill digs into one weak spot, loop vets a design before code. The columns are Know, Judge and Prove. One ledger sits under all six and learns.] ## Install In Claude Code, from your project: ```text /plugin marketplace add mvp-scale/mm3 /plugin install mm3@mvp-scale ``` Pick **project** scope. Claude asks for a TypeSafe API key (masked, optional): press Enter on "TypeSafe API key", paste, Enter, then "Save configuration". Leave it empty to add one later with `mm3 init`. The plugin carries its own copy of MM3 and gives Claude the `mm3` tool, so Claude can run every command (`mm3 config --load` included) without anything else installed. It is updated through Claude Code, like any plugin. In your own terminal (optional, needs Node 22.13+). Run it without installing anything: ```bash npx @mvpscale/mm3 config ``` or install the command once, pinned to the version you ran: ```bash npm install -g @mvpscale/mm3 mm3 init ``` `mm3 doctor` says which `mm3` your terminal finds and warns when its version differs from the one running, with the fix (`mm3 init`). After you update the plugin, run `mm3 init` once to bring the terminal copy along. `mm3 init` never uses sudo and never edits your shell profile: it installs into your user folder (`~/.local`) and prints the one `export PATH=...` line if that folder is not on your PATH. **Status: beta.** It works well and we use it ourselves; formal benchmarks are coming. The plugin and the npm package are both out. ### Try it locally Bring a TypeSafe API key. MM3 wraps TypeSafe's API (Jev at `api.typesafe.ai`), and any TypeSafe endpoint works. Your key stays out of your project: Claude Code keeps it in its secure storage, `mm3 init` in your OS keychain (or a 0600 file). To switch endpoints: ```bash mm3 config # see the settings in effect, the endpoint included mm3 config --write # create .mm3/config.yaml, then uncomment baseURL: and set it mm3 config --load # check the file and record the load in the ledger TYPESAFE_BASE_URL=https://api.example.com mm3 class review.yaml # or for one run only ``` Then just ask your agent. The [MM3 skill](https://github.com/mvp-scale/mm3/blob/main/skills/mm3/SKILL.md) tells it when to reach for MM3, which verb fits and how to write the request; you read the verdict. ## See it run **The challenge:** you host n8n yourself and want it to run faster. You ask your agent where to start. It narrows the question to one place, the node loader's cleanup in `directory-loader.ts`, and asks MM3 twelve questions about it in one call: nine yes/no, two decisions and the goal itself. [Figure: A real MM3 quick class run on n8n (MM3-0008). On the left, a coding agent is asked where n8n could be faster and runs one mm3 class request: 12 yes/no and decision questions, one call, 321 ms, about $0.00006, and a next move. On the right, the run's real YAML: the request opens folded (one claim, with concerns, decisions and an mdl block collapsed to fit) and a pointer unfolds each in turn; then the whole response (verdict, gates and odds per concern, consensus and the next command) and what the ledger now holds, where looking the same request up again is free.] The request the agent wrote (45 lines): [Figure: The request the agent wrote for MM3-0008: one goal on one file, three concerns of three yes/no questions each, two decisions, and an mdl block saying why it asked.] The response, **real output · jev-1.13.0 · api.typesafe.ai · 321 ms · ~$0.000063**: [Figure: The response to MM3-0008: the gate fails; a gate and odds per concern and decision, consensus SPLIT, escalate true, what the mdl block recorded, and the next command.] - Every concern is written so that "no" is healthy: a `pass: no` answer clears the bar at 0.30 or below. - **Goal: pass** (0.79). Yes, this cleanup could be faster. - **Availability: fail.** At 0.88, the `realpathSync()` calls on line 606 block the event loop. - **Design: fail** on all three: a rescan on every `unloadAll()` call (0.78), a try-catch that silently skips errors (0.90), and caching would help (0.84). - **Design-risk: unsure.** None of 0.40, 0.46, 0.64 lands clearly either way. - **Decisions:** measure first passes (0.89); severity is unsure (low, 0.56). - **Next:** the concerns disagree, so consensus is SPLIT, `escalate` is true, and `next:` points at a drill into availability. The `mdl:` line lists what the ledger recorded about why the agent asked, so later runs on this code start from it. ## Why MM3 An agent can ask a fast classifier a yes/no about your code, but on its own that answer is untraceable and never reused. MM3 asks the same way every time, scores the answer and keeps it, so every check adds to what you know about your codebase. **Three angles, three kinds of knowledge.** Every concern is asked three ways, so no decision rests on one look. Every request keeps what you asked, the scored verdict and the problem you were working on. Know what you know, judge fairly, prove it. The name is **MAK³** (make) plus **MDL³** (model), each across Know, Judge and Prove. It started as Sidewise; we hope you like the new name. **MAK³ uses what is proven.** Know what's been done here before you act, judge one claim, prove the change. [Figure: MAK3 make. view: your codebase memory, every past check searchable in milliseconds, free. class: turns is this OK into a question set you can measure, reuse and cite. replay: tests for your judgments, rerun any past check on any two commits.] **MDL³ learns what is missing.** In code you don't know yet: take a census, follow the flag, prove the design before anyone builds it. [Figure: MDL3 model. scan: a census of the present, same questions every file, comparable answers. drill: turns a flag into a fix target and hands it back to class and replay. loop: proves the design before a line of code exists.] view knows the past, scan the present, loop the future. Any order; every run feeds the next. **The ledger is our favorite part.** One append-only JSONL file. The runs you'd make anyway become a heat map at 10, a reference architecture at 20, a knowledge graph at 50. And `mdl:` is open: keep the stock fields, add your own, free, and your agent paints any of it back from the log or the SQLite index. ```yaml mdl: # stock why: find # validate · find · debug area: [build, hosting] # where in the stack stage: operate # design … operate problem: n8n startup scans node_modules with sync FS calls # one line, your words uses: # C4 chains - person:self-hoster -> system:n8n -> container:n8n-main - container:n8n-main -> component:core/nodes-loader -> code:unloadAll blast: container # how far a failure reaches # yours: any key, any lens hypothesis: caching symlink roots cuts cold start by a third # developer user-feels: first workflow after a restart is slow # product cost-driver: every pod restart pays the full scan # business pager-risk: none, slow not down # on-call tags: ["#perf", "#cold-start", "@platform-team"] # common dev tags: #topic, @owner refs: [BUG-1042, "test:startup-bench failed"] # bugs, test results, anything to link ``` Once a field settles over a few runs, `mm3 report fields --accept tags` promotes it: every tag becomes a node in the knowledge graph, tied to the runs that carried it. [Figure: The MM3 ledger. One append-only JSONL file, filled by runs you were making anyway. The same 10 runs give you a layered heat map of your architecture; 20, a reference architecture with a heat map; 50, a knowledge graph of your whole system. The mdl block is yours: add any field to any request, such as standard or owner, for free. Your agent reads the log or the SQLite index in milliseconds and paints charts, maps and graphs. Where your agents are strong, and where they are not.] Same questions on unchanged code come back from the ledger: no call, no cost. `mm3 report` shows where your agents keep going wrong, and a budget cap that only you raise stops runaway spend. ## Architecture That's all of it: one engine on your machine, one folder in your project, one call out. Simple on purpose, built to grow. [Figure: MM3 architecture. Any coding agent calls MM3 through its MCP tool or the mm3 CLI. MM3 runs on your machine: it reuses answers it already has and pays only for new questions, sent to a TypeSafe-compatible classifier endpoint (typesafe.ai or self-hosted), its only outside call. In your project it reads code and git as evidence, appends every run to log.jsonl, the append-only ledger and the only source of truth, and keeps index.db, a read-only hot cache of the ledger that rebuilds itself and answers lookups in under a millisecond, tested to 100,000 runs. No server, no MM3 account, no telemetry. config.yaml sets the budget cap, endpoint and model, reuse rules and your own mdl fields, and is the one file in .mm3/ that git keeps. MM3's whole footprint in your project is the .mm3/ folder.] ## Journeys Each journey asks one question: **can a smaller model, with MM3, learn more about a real codebase for less?** One agent, a $1 cap that covers everything, the same three beats (Know, Judge, Prove). Every run is on file in [mm3-journeys](https://github.com/mvp-scale/mm3-journeys), so you can rerun it and compare. | | The hook | Status | |---|---|---| | 1 | [WordPress](#journey-1-wordpress): make every page render per reader, core untouched | done: 8 of 12 right, 28k tokens written, under $1 | | 2 | Turn one dial: what changes when a decision gets more questions? | planned | ### Journey 1: WordPress **The hook:** take WordPress 7.1.2 as it ships and teach it to render a page differently for each reader, as a plugin, core untouched. [Figure: Journey 1, WordPress 7.1.2: make every page render per reader, as a plugin, core untouched. Agent Claude Sonnet 5.5; 28k tokens written and 2.6M read from cache; MM3 14 runs, 15 calls, about a tenth of a cent; about 15 minutes; 8 of 12 right against our answer key. Know: four runs map the boot layers, user meta and render chain (MM3-0001 to 0004). Judge: the plan and the definition of done each had gaps, a drill found four missing checks, and the fixed version passes them (MM3-0006 to 0011). Prove: nine behaviour checks met, a header-after-output risk found and fixed, a replay shows zero regressions (MM3-0012 to 0014). A gap is a requirement MM3 found: gap is fail, open is unsure, met is pass.] **The recap:** one agent mapped WordPress, checked its plan before writing code, then committed a small plugin and proved what held. MM3 cost about a tenth of a cent. The agent cost about $0.98 at today's list prices, so the whole journey came in under the $1 cap. **How to read it.** A gap is a requirement MM3 found, not a failure: gap is fail, open is unsure, met is pass. The agent's figures are counted from its session's token usage and priced at list; the score is against our own answer key. Every verdict came from reading code, not running WordPress. It missed WordPress's Block Bindings API, and `replay` can't yet compare against a file that didn't exist before (it's beta). **What we prompted** (the goal, verbatim; the beats and budget follow in the [full brief](https://github.com/mvp-scale/mm3-journeys/tree/main/wordpress/run-3-each-beat/brief.md)): > **Goal:** a page in WordPress renders differently depending on a simple persona of the current user, built from what WordPress can already see about them (what they read, what they like, what they adjusted). It only has to prove the page can change by persona; how accurate the persona is doesn't matter. Build it as a plugin under `src/wp-content/plugins/`; don't edit WordPress core or the bundled themes. **All of it is public:** the brief, the notes for each beat, the diff, the score against our answer key and the full ledger for each of the three runs: [run 1, bare brief](https://github.com/mvp-scale/mm3-journeys/tree/main/wordpress/run-1-bare-brief), [run 2, plugin guidance](https://github.com/mvp-scale/mm3-journeys/tree/main/wordpress/run-2-plugin-guidance), [run 3, each beat spelled out](https://github.com/mvp-scale/mm3-journeys/tree/main/wordpress/run-3-each-beat). > [!IMPORTANT] > **Lessons learned.** The prompt decides whether your agent uses MM3 as a map or as a spot check: 2 runs on a bare brief, 14 when each beat was spelled out. > > **Corrective action.** If you want the map, spell the steps out in your agent's brief; the run 3 brief linked above is a starting point. MM3's own guidance stays short on purpose, because one `class` call is often the whole job. `mm3 init --agents` adds a pointer to MM3's command cards to `AGENTS.md` and points `CLAUDE.md` at it. ## Run it Let your agent drive. `mm3 agent` prints every command and rule in one dense card built for agents, and `mm3 agent ` gives one verb's rules with good and bad examples. Claude agents read it and run the commands just fine, down to Haiku: it drove both of our stories end to end. For humans, `mm3 help` is the same contract in plain words, and `mm3 template ` prints a filled-in request with its rules as comments. Four jobs to start with: - **Check a change before you merge it.** One `class` call, three angles per concern, a verdict for each. Start with `mm3 template class`. - **Prove a fix actually worked.** `replay` re-asks a past run's own questions across two commits, so you check the fix without re-checking everything. Start with `mm3 template replay`. - **Find where a problem lives.** `scan` sweeps a folder and ranks the files that most need a look. Start with `mm3 template scan`. - **Check a design before any code exists.** `loop` puts a plan through the same checklist before anyone writes it. Start with `mm3 template loop`. Run it by hand Every line below runs in order in a fresh project: ```bash # mm3-quickstart mm3 template class > review.yaml mm3 class review.yaml mm3 view src mm3 report mm3 outcome MM3-0001 held --by you mm3 budget ``` Add `--dry-run` to any request to validate it and count its questions without a call. With no key set, MM3 falls back to a built-in sample provider so these lines still run; its answers are canned and labelled, never evidence. ## Limits and alternatives - **Advice, not action: it gives the odds, you make the call.** Delete, deploy, drop and pay stay human. - **A pass is a probability: a [calibrated](https://github.com/mvp-scale/mm3/blob/main/docs/numbers.md#what-calibrated-means) 0.9 is wrong one time in ten.** `unsure` is a real answer, and `mm3 outcome` shows which verdicts held. - **Run linters, scanners and tests first: they're free and exact.** MM3 takes the questions they can't ask, like "does this handler check the caller?" - **First principles still apply: map your architecture first.** A few cheap checks name the layers (tag requests with `mdl.uses`) and `mm3 report graph` draws them. Every later question then lands in a known place, and reading it back is free. - **Use a full review for open questions.** A model that reads the whole codebase answers anything. MM3 answers yes/no. - **Beta: works well in our own use and on an intentionally vulnerable app (OWASP NodeGoat).** Formal benchmarks are coming. ## Docs and contributing - [The contract](https://github.com/mvp-scale/mm3/blob/main/docs/contract.md): every claim the answer format makes - [Evidence for each claim](https://github.com/mvp-scale/mm3/blob/main/docs/evidence/README.md) and [where the numbers come from](https://github.com/mvp-scale/mm3/blob/main/docs/numbers.md) - [The agent skill](https://github.com/mvp-scale/mm3/blob/main/skills/mm3/SKILL.md), and `mm3 agent` / `mm3 help` in your terminal - Contributing: see [AGENTS.md](https://github.com/mvp-scale/mm3/blob/main/AGENTS.md) for commands, test tiers and rules. Open pull requests against the `nightly` branch. ## License [Apache-2.0](https://github.com/mvp-scale/mm3/blob/main/LICENSE) · [Contributing](https://github.com/mvp-scale/mm3/blob/main/AGENTS.md) · [Issues](https://github.com/mvp-scale/mm3/issues) Last updated 2026-10-06. ## Pattern playbook A pattern is a small, named way of using a verb. The same six verbs do more as you climb: you start by using them, then share them, chain them, join them to your other tools, and finally let them run on a schedule. Multipliers are a mental model, not a benchmark. "Anticipate" means reading your own history early: MM3 records and reports, and does not forecast. Live version: https://mm3lab.dev/patterns How a sweep decomposes: loop, scan and drill share one shape. (1) decompose: over: part -> story, you write the tree; (2) ask one checklist per layer, asked of every item; (3) sweep, one call per layer, a verdict per item; (4) read the next: line, which points at the worst item; (5) go down with drill --from , listing the next layer yourself for ideas (over: call: [a, b, c]) or using each for code; (6) repeat until you reach code, then class and replay. loop works on ideas you write, scan on code MM3 reads (file, function, call), drill goes one layer down from one flagged item. ### 1× Coding agent / Vibe coder (Use): "It runs. Ship it." Finds value in the first use, and again in every reuse. - view / Memoize: Ask the ledger before you ask a model. The same question on unchanged code comes back free. ``` mm3 view src/handlers ``` - class / Smoke test: One quick verdict on what the agent just wrote, before you accept it. ``` mm3 template class > review.yaml mm3 class review.yaml ``` - replay / Regression test: After a fix, re-ask the same questions and see whether it held. ``` mm3 replay --parent MM3-0042 \ --compare main..HEAD --expect injection ``` - scan / Census: Same questions across every file, so the answers are comparable. ``` mm3 template scan > scan.yaml mm3 scan scan.yaml ``` - drill / Root cause: Don't guess at a failure. Go down into the one thing that flagged. ``` mm3 template drill --parent MM3-0060 \ --from src/handlers/user.ts/findUser ``` - loop / Decompose: Break an idea into parts and stories, then check every one in a single run. Add --dry-run to count the questions first. ``` mm3 template loop > plan.yaml mm3 loop plan.yaml --dry-run ``` ### 5× Senior engineer (Standardize): "Show me." Sees that the pattern matters as much as the tool: how checks are shared between people and between runs. - view / Provenance: Every claim carries its run id, so anyone can follow it back to the evidence. ``` mm3 view MM3-0042 # cite MM3-0042 in the PR description ``` - class / Quality gate: The same checklist before every merge. Keep the request file in the repo. ``` mm3 class review.yaml # read the failing category, then the next: line ``` - replay / Golden master: Pin a verdict you trust and compare every later change against it. ``` mm3 replay --parent MM3-0042 \ --compare v1.2.0..HEAD --expect injection ``` - scan / Triage: Add a scale question so findings rank by severity, worst first, instead of an unordered map. ``` decisions: severity: pass: [none, low] 10: scale: How severe is {function}? levels: [none, low, medium, high, critical] ``` - drill / Runbook: Write down which flag leads to which drill, so the next person doesn't improvise. ``` access fails -> mm3 template drill \ --parent --from access ``` - loop / Architecture review: Make the parts the components and ask each one the same questions: responsibility, dependency, testability. A whole design reviewed before code. ``` over: part: [gateway, payments, ledger] ask: part: concerns: responsibility: pass: yes 1: Does {part} own one clear responsibility? ``` ### 10× Agentic engineer (Chain): "I direct. I verify." Links the patterns and the calls together, so the output of one becomes the input of the next. - view / Cross-check: A helper agent says it did the work. Check its report against the ledger instead of taking its word. ``` mm3 agent delegate # give helpers the rules mm3 view MM3-0042 # verify each cited run ``` - class / Escalate: A fail is not the end. Follow the next: line straight into a drill. ``` mm3 class review.yaml # next: mm3 template drill --parent MM3-0001 --from access ``` - replay / Suite: Replay a whole sweep, not one question. You get fixed, still and regressed for every item. ``` mm3 replay --parent MM3-0060 \ --compare main..HEAD --expect injection # MM3-0060 is a scan: fixed / still / regressed per item ``` - scan / Fan-out: Give each agent its own folder. Merge what they find by run id. ``` mm3 scan scan-api.yaml # agent A mm3 scan scan-web.yaml # agent B mm3 report hits # worst first ``` - drill / Closed loop: Class finds it, drill pins it, replay proves the fix. One loop, three verbs. ``` class -> drill -> (fix) -> replay ``` - loop / Stepwise refinement: A loop inside a loop. Drill into the weakest story and list its sub-steps yourself: the next layer, checked the same way. ``` mm3 loop plan.yaml mm3 template drill --parent MM3-0003 \ --from "gateway/guest checkout" # then in drill.yaml: # over: {call: [cart-form, payment-token, receipt]} ``` ### 25× Forward Deployed Engineer (Integrate): "Drop me in anywhere." Brings in more data and more tools, so decisions across a complex system get easier to make. - view / Join: Walk one run's lineage and see every drill, replay and outcome that hangs off it, in one place. ``` mm3 view MM3-0001 # lineage: scan, the drill below it, # the replay that proved the fix ``` - class / Decision record: A choice question turns a verdict into a decision among named options, kept with its evidence. ``` decisions: route: pass: [ship] 5: choice: Where should this go? options: [ship, fix, block] ``` - replay / Contract test: Replay the other team's checklist against your branch before you hand it over. ``` mm3 replay --parent MM3-0091 \ --compare main..worktree --expect auth # MM3-0091 holds their questions ``` - scan / C4 map: Tag each request with the layers it touches. The report draws the architecture for you. ``` mdl: uses: - container:api -> component:auth -> code:verify # then: mm3 report graph component:auth ``` - drill / Postmortem: Drill the failure, then record the outcome and the bug it belongs to. ``` mm3 drill drill.yaml mm3 outcome MM3-0042 held --by the-owner # someone other than the asker confirms it ``` - loop / Fitness function: Turn the design's key claims into questions on the real code, and rerun them every release. ``` mm3 loop plan.yaml # the claims mm3 class claims-on-code.yaml # each release, on the real code ``` ### 50× Growth hacker (Anticipate): "Make the loop run itself." Automates all of the above and reads the history early, so trouble is seen coming and not only handled after. - view / Hotspot: Let history show where trouble clusters. A place that keeps failing is where the next problem will be. ``` mm3 report web # writes .mm3/viewer.html: heat map and consensus map ``` - class / Calibrate: Record whether verdicts held. Trust the kinds of verdict that do, and doubt the ones that don't. ``` mm3 outcome MM3-0002 held --by the-owner mm3 report history ``` - replay / Canary: Replay on every commit and watch the trend. A drift toward regressed is the early warning. ``` # in CI, on every commit mm3 replay --parent MM3-0042 \ --compare main..HEAD --expect injection mm3 report history ``` - scan / Drift watch: Re-sweep on a schedule against a baseline scan, and see what is moving before it breaks. ``` # weekly: MM3-0060 is the baseline scan mm3 replay --parent MM3-0060 \ --compare main..HEAD --expect injection ``` - drill / Pareto: A few weak spots cause most of the failures. Rank them and fix those first. ``` mm3 report hits # worst first mm3 report patterns # what keeps failing ``` - loop / Pre-mortem: Before building, assume it failed and ask why. Healthy answers are no. ``` failure: pass: no # a healthy answer is no 1: Could "{story}" fail silently in production? 2: Could "{story}" lose a customer's money? ``` ## Agent support Status key: verified = tested end to end; early testing = rounds run, results still coming in; expected to work = follows the open standards (AGENTS.md, SKILL.md, hooks and plugins where supported), not yet tested. Live version: https://mm3lab.dev/agents - Tier one: Claude Code (reads CLAUDE.md, imports AGENTS.md): verified; working: Plugin install, the mm3 tool, the AGENTS.md pointer, the hook nudge; known gaps: None recorded - Tier one: Codex (reads AGENTS.md): early testing - Tier one: Gemini CLI (reads GEMINI.md, AGENTS.md): early testing - Tier one: GitHub Copilot (reads AGENTS.md, skills): early testing - Tier one: Cursor (reads AGENTS.md, rules): early testing - Open source: OpenCode (reads AGENTS.md): early testing - Open source: Aider (reads AGENTS.md): early testing - Open source: Hermes (reads AGENTS.md): early testing - Open source: Goose (reads AGENTS.md): expected to work - Open source: Cline (reads AGENTS.md): expected to work - Open source: Kilo Code (reads AGENTS.md): expected to work - Open source: Pi (reads AGENTS.md, skills): expected to work - Open source: OpenClaw (reads AGENTS.md): expected to work - Other agents: Amp (reads AGENTS.md, skills): early testing - Other agents: Factory Droid (reads AGENTS.md): early testing - Other agents: Kiro (reads AGENTS.md): early testing - Other agents: Antigravity (reads AGENTS.md): early testing - Other agents: Any agent that reads AGENTS.md (reads AGENTS.md): expected to work