Your plan rests on beliefs nobody has checked. Track the ones that could kill it, and nothing else.

A ledger for the beliefs your business runs on. Evidence graded by what people did, not what they said. Work in the app, or let Claude and Cursor keep the ledger from your meetings.

No card needed. AI agents connect on every plan.

The Assumption Mapper dashboard: 44 open assumptions, peak risk 5.0 of 5.0, a ranked list of top risks with risk and confidence scores, and a category-by-domain coverage cross-tab.

The rule

Fewer, better

Most tools reward you for writing more down. This one pushes back. Restatements of something you already track are merged on ingest, and a source that proposes more than twelve new beliefs is refused outright.

10–30
Live assumptions per project
The working range. Eighty is a list nobody triages.
12
New beliefs per source, maximum
Past that the ingest is refused. Trim it or merge it.
16
MCP tools
For Claude Code, Claude Desktop and Cursor.

How it works

Four moves, run every week

Your agent can drive all of them, or you can work the queue by hand. Same ledger, same scoring.

01Capture

It enters the ledger while it is still a belief

A transcript pasted into the app, a claim typed straight into the queue, or your agent handing over the meeting it just sat through. Each one arrives with a category, a domain and an importance — not as a bullet in a doc nobody reopens.

The triage queue: assumptions sorted by risk, each with a risk score and a P1–P5 priority, and a peek panel open on the right.
02Weigh

Evidence is graded by what people did, not what they said

Six rungs, from repeated behaviour at the top down to reasoning by analogy at the bottom. Ten enthusiastic quotes are still worth less than one renewal, and the ledger says so out loud.

Evidence cards on an assumption dossier, each tagged with its strength rung and weight: repeated behavior 1.00, opinion 0.15.
03Monday

The week opens with three lists, not a dashboard

What moved since last week. Needs you, overdue tests first. On your plate, the beliefs and tests you personally owe. Everyone gets their own Monday email, and owners get a nudge on Friday.

The Monday view: what moved this week above a ranked “needs you” queue with risk and confidence badges.
04Decide

Record the call, keep the receipts

An assumption is ready once enough real evidence has landed, and a Critical call needs more of it than a Medium one. Decide it and the reasoning stays on record — including for your agent. The one page puts the whole project in front of a board, as a read-only link you can share.

An assumption dossier: the Make decision button above a trajectory chart plotting risk and confidence across four snapshots.

The wedge

One command, and the agent you already have keeps the ledger

A published MCP server with sixteen tools. Add it to Claude Code, Claude Desktop or Cursor, and your agent captures assumptions from a meeting, grades the evidence, plans a test, and reports what moved.

your project — zsh
$ claude mcp add assumption-mapper \
    --env ASSUMPTION_MAPPER_API_KEY=am_your_key \
    -- npx -y assumption-mapper-mcp

Create a key in the app under Settings → Account → API keys — read-only, scoped to one project, with an expiry if you want it — then restart the client. Claude Desktop and Cursor take the same server as a four-line JSON block.

Claude Code — assessment brief, excerpt

You

Transcript from this morning’s Delphi call. Log it against NovaKit.

Claude

→ ingest_transcript(projectId: novakit, extraction: 2 assumptions, 2 evidence)

Ingested delphi-2026-07-24.txt · 1 source · 2 evidence items

Assessment:
  Evidence logged: 2 items · 1 independent source
  Strength mix: stated intent ×2
  Assumptions moved (1):
    • confidence 50 → 59 · risk 4.00 → 3.20 · 2 new (2 supports)
      Usage-based pricing beats per-seat for expansion revenue
  Drift flags (1) — stated as settled fact, still undecided:
    • "EU data residency is table stakes for mid-market" —
      named as a requirement, not as a hypothesis.
  Cautions:
    ⚠ All 2 new items are stated intent from a single source —
      one independent data point, not 2.

Both signals are stated intent from one call, so the pricing assumption stays in triage. To test it, give me the kill criterion first — I can’t log an experiment without one.

Works on every plan

Your agent extracts with its own model, so our server makes no AI call and agent-side ingest has no plan gate. Pasting a transcript into the app is the Pro and Team path.

Scores stay server-side

Your agent submits raw evidence: a quote, a direction, a strength, a source. Risk, confidence and evidence mass are computed in Postgres — the write schema has no field for them.

Why it holds up

Two things a doc cannot do for you

A Notion page will hold any belief you put in it, at any confidence you claim, for as long as you like. That is the whole problem.

Evidence that can’t flatter you

Supporting and contradicting evidence are counted separately, and stronger evidence counts for more. Six warm quotes and two hard contradictions do not average out to fine: contradiction raises urgency, so the belief stays loud until you decide.

Evidence strength

Weight

  • Repeated behavior1.00

    Renewals, recurring usage or purchases.

  • Committed behavior0.80

    Pre-order, signed LOI, paid pilot, deposit.

  • Engagement0.60

    Return visits, referrals, unsolicited requests.

  • Stated intent0.30

    “I would use it / pay for it / switch.”

  • Opinion0.15

    Expert, advisor or team opinion. Secondhand reports.

  • Analogy0.05

    Reasoning by analogy, with no external data.

Weight fades with age — half at six months, never below a quarter — so last year’s research stops counting as this year’s proof.

The experiments block on an assumption: “No experiment planned. Commit to a kill criterion before you test.”

Kill criteria, pre-committed

An experiment cannot be recorded without a written kill criterion: “we abandon this if <observable> by <date>”. It is a required field on the API and an instruction the agent is told not to satisfy on your behalf.

“We already do this in Notion.”

Notion never refuses. It will hold an experiment with no exit condition, and a criterion agreed after you have seen the result is not a criterion.

Where your data sits

Row-level security on every path
An API key is exchanged for a short-lived, RLS-scoped token, so it sees exactly what its owner sees — and it cannot mint, list or revoke keys.
Your transcript can stay with your model
When your agent extracts, our server makes no AI call. Pasting a transcript into the app is the one path that sends text to Anthropic, our AI subprocessor.
No ad trackers, no data sale
Encrypted in transit and at rest, operated from the EU. Delete your account and your data goes within 30 days.

Full detail in the privacy policy.

Pricing

Free today

Paid tiers buy volume and seats, never the method: the ladder, the scoring, kill criteria and the Monday lists are on every plan. Pro and Team are built, but there is no payment integration yet — nothing to buy today.

Free

Available now
€0forever
  • 30 active assumptions
  • 1 project · 3 members
  • MCP server and agent-side ingest
  • Monday lists, reports and export

Pro

€6/month
  • Unlimited assumptions
  • 20 projects
  • AI transcript import in the app

Coming soon

Team

€29/month
  • Everything in Pro
  • 20 members
  • Shared library and admin

Coming soon

The free cap counts active assumptions, so deciding one frees a slot.

Questions

The ones that actually come up

What counts as evidence?
Anything you can attribute to a source: a quote from a call, an analytics figure, a signed pilot, an advisor’s take. Each item gets a direction — supports, contradicts, ambiguous — and a strength rung, from repeated behaviour at 1.00 down to analogy at 0.05; unclassified counts as stated intent, 0.30. Weight fades with age. Readiness scales with importance: a Critical call needs 1.4× the evidence of a Medium one.
Why does it refuse to record everything I bring it?
A project that works carries 10–30 live assumptions, not 80; a list nobody can read is a list nobody triages. Claims that restate something you already track are merged on ingest, and more than 12 new beliefs from one source is refused. The fix is to trim, not to override.
Which agents work with it?
Claude Code, Claude Desktop and Cursor are documented, and any MCP client that speaks stdio works — one npx command or a four-line JSON block. Sixteen tools, across capture, assess, act and recall. There is also a companion skill that teaches your agent the method, not just the tool names, and a CLI (am) for piping a directory of transcripts in from the terminal.
Can my team use it without an agent?
Yes. The app is the shared surface: the Monday lists, the triage queue, dossiers with evidence trajectories, experiments, the decision log, and three reports — the one page, coverage (never tested, gone quiet, gaps) and your tests. The agent is an ingestion path, not a dependency. The free plan covers three members on one project.
What happens to my transcripts?
Each ingest is stored as a source record, the provenance behind every piece of evidence, so months later you can ask what a call taught you and get the quotes back. The content is hashed, so re-ingesting the same file returns the original brief instead of double-counting it. Encrypted in transit and at rest; removed within 30 days of account deletion.
Do you train on my data?
No, and we build no models. Paste a transcript into the app and the text goes to Anthropic for extraction — named as a subprocessor in the privacy policy alongside Supabase, Vercel, PostHog, Sentry and Resend. If your own agent extracts, our server makes no AI call and forwards nothing. No ad trackers, and we do not sell personal data.

Name the beliefs your business is betting on. Then let something other than optimism grade them.

Start with one project and the handful you already argue about.

€0 forever · 30 assumptions · your agent on every plan