Docs

Everything about setting up Samethru, the formats it reads and what it reports.

For samethru 0.1.0.

Contents

In short: Samethru checks that what your app shows on screen matches what's saved in your database. Every value makes a trip: it's saved in the database, sent by your server (the API), and shown on the page. Samethru reads it at all three stops, and when they don't match, it tells you where it went wrong and which line of code to look at.

Samethru checks that a value is the same in your database, in your API's response, and on the rendered page. It runs locally as an MCP server inside AI coding tools (Claude Code, Cursor, OpenAI Codex, Gemini CLI). Your coding assistant reads your code to work out where each value on screen comes from. Samethru then runs the checks for real, against your running app, and reports every mismatch with evidence.

It catches the bugs that type checks and unit tests miss because every layer is "right" on its own:

  • a price stored in cents, shown as dollars ($1,999.00 instead of $19.99)
  • a date that's a day early for everyone west of UTC
  • a dashboard number served from a cache long after the database changed
  • a list that silently stops at the first page, a count that ignores the filter, null shown as text

How it works

  1. Your assistant traces. Following Samethru's audit workflow, it opens each page with Samethru, reads your code to trace every value back to its API field and database column, and saves what it found as a small check spec in .samethru/checks/.
  2. Samethru checks. For each spec it reads the database (read-only), then the API, then the page in a real browser, compares the values in order, and works out where the data first goes wrong.
  3. You get a report. Each check ends in exactly one verdict:
Verdict Meaning
PASS The check ran and the values matched.
FAIL The check ran and found a mismatch. It comes with the value at each layer, where each came from, and the bug pattern it matches.
NOT_RUN The check couldn't run (for example, the app wasn't reachable). The reason is given.
INCONCLUSIVE The check ran but couldn't decide. It says what evidence was missing.

A check that didn't run is never counted as a pass. Every report also says how many of the values on each page (and fields in each API response) a check reads, and lists the ones none does, so nothing unchecked looks fine by omission.

Every failure names one of 12 bug patterns with a plain explanation, such as "Cents vs dollars", "Timezone shift", "Stale cached value" or "List cut off by pagination". The full list, with how each is detected, is in the bug catalogue.

Quickstart

You need Node.js 22.13 or later, Google Chrome (or Playwright's Chromium), an app you can run on your machine, and one of Claude Code, Cursor, Codex or Gemini CLI.

  1. Set up your project. In its folder, run:

    npx samethru init
    

    It finds your AI tools and asks which to set up. Then it shows every change as a diff and asks before making it: the MCP server entry for each tool, two .gitignore lines, a starting .samethru/config.yaml, the audit skill, and (for Cursor and Codex) a short section in AGENTS.md. --dry-run shows the changes without making them.

  2. Point it at your app. Edit .samethru/config.yaml: set baseUrl, and add your database. Samethru only ever opens it read-only, and passwords come from environment variables:

    version: 1
    baseUrl: http://localhost:3000
    browser:
      timezone: America/New_York
    databases:
      main:
        driver: postgres            # or sqlite, with path: db/development.sqlite3
        url: ${env:DATABASE_URL}
    
  3. Start your app, the way you do when developing.

  4. If pages need a login, run npx samethru login. A browser window opens, you log in as usual (SSO and two-factor work), and the session is saved to .samethru/auth.json, readable only by you and gitignored.

  5. Check the setup: npx samethru doctor. It starts the MCP server, launches the browser, connects to each database (and verifies it's read-only), reaches your app, and says what to fix.

  6. Ask your assistant to "run a Samethru audit" (or to audit one page). In the Claude Code CLI you can also type /mcp__samethru__audit, and in Gemini CLI /audit. Claude Code asks you once to approve the new server.

  7. Re-run the saved checks whenever you like, with no AI involved: npx samethru run. That's also how to run them in CI.

Commands

Command What it does
npx samethru init Sets up the MCP server for your AI tools, .gitignore, a config, the skill and AGENTS.md. --tools claude,cursor,codex,gemini (or all, none), --scope user (your user config instead of the project's), --dry-run, --yes.
npx samethru login Saves a logged-in session for checks of pages behind a login. --url /login, --yes.
npx samethru doctor Checks the whole setup and says what to fix. --json.
npx samethru run Runs the saved checks, prints the report and writes it to .samethru/runs/. --ids, --tags, --min-coverage <percent>, --format md|json.
npx samethru mcp The MCP server itself. Your AI tool starts it; you don't need to.

Every command takes --root <dir> for a project in another folder.

Where init puts the MCP server

Tool In the project (default) With --scope user
Claude Code .mcp.json claude mcp add --scope user samethru -- npx -y samethru mcp
Cursor .cursor/mcp.json ~/.cursor/mcp.json
Codex .codex/config.toml (read only in trusted projects) ~/.codex/config.toml
Gemini CLI .gemini/settings.json ~/.gemini/settings.json

An existing samethru entry is never modified: init says it left it unchanged and prints both entries. Other entries, comments and formatting in these files stay as they were.

doctor

doctor prints one line per finding (✓ ok, ! warning, ✗ problem), each problem or warning with its next step, then a count. It exits with 1 if anything is ✗, otherwise 0. It checks Node.js, the config and saved checks, that .samethru/runs/ is gitignored, each AI tool, the MCP server (all seven tools must be there), the browser, each database (connected, read-only verified, SELECT 1), the app, and the saved login if there is one. For Postgres it warns when the role is a superuser or can write to any table.

Exit codes

samethru run exits with:

Code Meaning
0 Every check ran and passed.
1 At least one check failed.
2 Nothing failed, but something didn't run or couldn't decide, no check ran at all, or page coverage is below --min-coverage.
3 A config or usage error.

When the code isn't the obvious one, the reason goes to stderr, for example "1 NOT_RUN and 0 INCONCLUSIVE: not a pass."

Running in CI

Start your app (and a database with test data), then:

- run: npx samethru run --min-coverage 60

The report is printed to the log, and --format json prints the full result (its schema is schemas/report.schema.json in the package). Each run is also written to .samethru/runs/<runId>/ as run.json and report.md.

Configuration

.samethru/config.yaml lives in your project and is committed. init writes a starting one. Every field:

version: 1
baseUrl: http://localhost:3000          # page + API URLs in specs are relative to this
allowedHosts: []                        # non-loopback hosts you allow, e.g. [staging.example.com]
browser:
  timezone: America/New_York            # default: system timezone. Pin it so date checks reproduce.
  locale: en-US
  viewport: { width: 1280, height: 800 }
  thirdPartyRequests: allow             # allow (default) | block. Every outside host is listed either way.
auth:                                   # optional, for logged-in pages
  storageState: .samethru/auth.json     # written by `npx samethru login` (gitignored)
  loginUrl: /login                      # where `login` opens the browser (default: baseUrl)
  check:                                # an element that only exists when logged in
    url: /account
    selector: "[data-testid=user-menu]"
api:
  headers:
    Authorization: "Bearer ${env:APP_API_TOKEN}"
  allow:                                # GET and HEAD are always allowed. Any other method must be listed.
    - POST /api/search
    - POST /api/reports/*/preview       # * = one path segment, ** = any number of segments
    - POST /graphql                     # paths ending in /graphql are GraphQL: queries only, never mutations
    - { endpoint: POST /api/gql, graphql: true }   # mark other GraphQL endpoints explicitly
databases:
  main:
    driver: sqlite                      # sqlite | postgres
    path: ./data/app.db
  warehouse:
    driver: postgres
    url: ${env:DATABASE_URL}
statementTimeoutMs: 5000
defaults:
  money:
    decimals: [0, 2]                    # optional: this site shows whole amounts without decimals ("$24")
output:                                 # what goes back to the assistant
  maxRows: 20                           # query_db default rows returned inline
  previewChars: 300                     # response previews and element text
coverage:
  ignore:                               # listed in reports as "excluded" with the reason, never as covered
    - { url: /products, selector: footer, reason: static copy }
redact:                                 # added to built-in defaults (authorization, cookie, set-cookie, password, token, secret…)
  headers: [x-internal-key]
  jsonKeys: [ssn]
  • Secrets are referenced as ${env:NAME} and resolved only when used. An unset or empty variable is an error, never an empty string. Validation warns when a literal password appears in a database URL or a secret-looking header, and when something looks like a malformed reference (${NAME}, $NAME).
  • Defaults: browser.timezone is the system timezone, browser.locale is en-US, output.maxRows 20, output.previewChars 300, statementTimeoutMs 5000, auth.storageState .samethru/auth.json.
  • Endpoint patterns are METHOD path. The path is relative to baseUrl, or a full URL on an allowedHosts host. Query strings are ignored when matching.
  • The project root resolves in this order: the --root flag, then CLAUDE_PROJECT_DIR (set by Claude Code), then the client's MCP roots, then the current directory.
  • Editors can validate the file as you type: # yaml-language-server: $schema=../node_modules/samethru/schemas/config.schema.json.

In your project, Samethru only touches .samethru/: config.yaml and checks/ are committed; auth.json and runs/ are gitignored.

Check specs

The assistant writes the specs, but they're plain YAML you can read, edit and commit. One file per check, in .samethru/checks/<id>.yaml. A check binds one value across two or three layers:

# .samethru/checks/product-7-price.yaml
version: 1
id: product-7-price
title: Price of the Walnut Desk Lamp
type: money
currency: USD
page: { url: /products/7, selector: "[data-testid=price]" }
api:  { request: { url: /api/products/7 }, jsonPath: $.price_cents, unit: minor }
db:   { connection: main, sql: "SELECT price_cents FROM products WHERE id = $1", params: [7], column: price_cents, unit: minor }
traced: { page: src/pages/product.tsx:42, api: src/routes/products.ts:18, db: db/schema.sql:7 }

Types are money, number, count, text, enum, datetime, date and list (with pagination). There is no code in a spec: parsing uses declared parsers and optional regexes.

Common fields

Field Required Meaning
version yes 1
id yes kebab-case, unique, equals the file name
title yes plain-language description of the value
type yes money, number, count, text, enum, datetime, date or list
page / api / db at least 2 of 3 the layer bindings (below)
currency for money ISO 4217 code. Expected decimals default to its minor units.
decimals no decimals the page may show, for money and numbers: a number, or a list such as [0, 2]
labels for enums raw value → expected display label, from the source of truth (backend enum, i18n file), not the rendering code under test
order for lists any (default) or same
fields for lists per-item fields and their types, e.g. price: { type: money, currency: USD }
pagination for lists intended pagination: { pageSize, pager }
missing no acceptable page renderings of a null value. Default: ["", "—", "–", "-", "N/A"]
context no extra evidence for the bug patterns: filter, unfilteredCount, updatedAt
traced no where the assistant found each binding (file:line). Shown in reports, never executed.
tags no for filtering runs

Layer bindings

page: url, then either selector (a single value) or items and key (and fields) for lists. read is text (the default, trimmed innerText), attr:<name>, value or count. pick is only (the default; more than one match makes the check INCONCLUSIVE), first, last or an index. parse is per type, for example money { locale }, datetime { format, zone }, count { pattern: "(\\d+) results" }. waitFor takes a selector.

api: request: { method?, url, query?, headers?, body? }, then jsonPath (a single value) or items and key (and fields). For lists the API paginates, add total (a JSONPath to the full count) or more (a JSONPath to a next cursor, next link or has-more flag). unit: major|minor for money, and zone for timestamps without an offset. source is auto (the default), page or direct: auto uses the response the page actually received when the check has a page layer, and a direct call otherwise.

db: connection, sql, params (required for literal values; there is no string interpolation), then column (a single value; exactly one row expected) or key (and fields) for lists. unit, and zone (default UTC) for timestamps without an offset.

Pagination

Most real lists are paginated on purpose. A list spec declares the intended pagination so a correct first page passes:

pagination:
  pageSize: 10                               # items per page the UI intends to show
  pager: "[data-testid=pager] a[rel=next]"   # the control that leads to the remaining items

A paginated page passes when it shows the correct first page and the pager is present, visible and enabled. It fails as pagination-cutoff only when items are missing and there is no way to reach them. Samethru checks that the pager is usable; it doesn't click it. Later pages that have their own URL (/customers?page=2) can get their own specs.

More examples

# .samethru/checks/customers-list.yaml (a list that is paginated on purpose)
version: 1
id: customers-list
title: Customer list shows the first page of customers, newest first, with a working pager
type: list
order: same
pagination:
  pageSize: 10
  pager: "[data-testid=pager] a[rel=next]"
page:
  url: /customers
  items: "[data-testid=customer-row]"
  key: "[data-testid=customer-email]"
api:
  request: { url: /api/customers }
  items: $.items
  key: $.email
  total: $.total
db:
  connection: main
  sql: SELECT email FROM customers ORDER BY created_at DESC
  key: email
# .samethru/checks/order-1001-placed-at.yaml
version: 1
id: order-1001-placed-at
title: When order 1001 was placed
type: datetime
page:
  url: /orders/1001
  selector: "[data-testid=placed-at]"
  parse: { format: "MMM d, yyyy, h:mm a" }    # interpreted in browser.timezone unless zone is given
api:
  request: { url: /api/orders/1001 }
  jsonPath: $.placed_at                        # ISO-8601 with offset
db:
  connection: main
  sql: SELECT placed_at FROM orders WHERE id = ?
  params: [1001]
  column: placed_at
  zone: UTC

Values are compared at the coarsest precision present: a page that shows minutes is compared to the minute. Datetimes are compared as instants, dates as calendar dates. Money and number values use exact decimals, never floats.

JSONPath

A small, safe subset: $, .key, ['key'], [n], [*], .* and an equality filter [?(@.id == 7)] with a string, number, boolean or null literal. No script evaluation, recursive descent or slices. A field missing from an object that exists is a real "no value". A path that can't be followed means the binding is probably wrong. In YAML's inline { … } style, quote paths that contain [*].

Validation

Specs are validated when they're saved and before they run. Problems come as errors and warnings, each with a dotted path and, for files, a line and column; unknown fields get a "did you mean" suggestion (jsonpath → jsonPath).

  • Errors include unknown fields, wrong types and missing required fields; fewer than two layers; an id that doesn't match the file name; an unknown currency, timezone or locale; a regex that doesn't compile or has more than one capture group; a JSONPath that doesn't parse; an unknown connection; a placeholder count (? for SQLite, $n for Postgres) that doesn't match params; pagination without a page layer; and SQL, page URLs or API requests that the safety rules would refuse.
  • Warnings include a missing traced entry, brittle selectors (nth-child, generated class names, deep chains), order: same without ORDER BY, pagination with order: any, enum labels shared by two values, a datetime format with no time of day, and a missing entry that looks like a real zero.

Editors can validate specs as you type: # yaml-language-server: $schema=../../node_modules/samethru/schemas/check.schema.json.

Verdicts and comparisons

  • Read order is database, then API, then page, as close together as possible, and each read is timestamped.
  • Comparisons run along the chain: database → API and API → page, or database → page when there's no API layer. That shows where the data first goes wrong.
  • Page against API uses the response the page actually received, not a fresh call. If the page never requested the endpoint, the comparison is INCONCLUSIVE ("the binding may be wrong"); it never silently falls back to a fresh call.
  • NOT_RUN covers infrastructure: a connection refused, a timeout, no browser, the login page, an API 401/403 or 5xx, a database that won't open or a query that times out, and specs that don't validate (reported, never skipped).
  • INCONCLUSIVE covers bindings that don't resolve: no matching element, an ambiguous element, a SQL error, 0 rows or more than 1, a JSONPath that matches nothing, another 4xx, a non-JSON body, text that won't parse, or a page that never requested the endpoint. Text that is null, undefined, NaN, [object Object] or Invalid Date is an observed value, because that is the bug.
  • The check's verdict is FAIL if any comparison fails. Otherwise it is PASS only if every comparison passed; otherwise NOT_RUN if any comparison didn't run, else INCONCLUSIVE. A partial run is never a PASS.
  • Money on the page must show an allowed number of decimals: the spec's decimals, else the config's defaults.money.decimals, else the currency's digits. So $12.5 fails even though it equals 12.50. With [0, 2], $24 for 24.00 passes, but $25 for 24.99 fails.
  • Other numbers on the page are compared at the precision shown, accepting half-up or half-even rounding but never truncation.
  • Enums: the page's label must equal labels[stored value].
  • Missing values: no value on both sides passes, whatever form each takes. A value on one side only fails.
  • Re-read guard: on a mismatch involving the database, the database is read once more after the page. If the value changed during the check, the verdict is INCONCLUSIVE.
  • Cache evidence: when the API or page differs from the database, Samethru collects the cache headers, the response's generation time (from Date − Age or Last-Modified), the database's context.updatedAt if the spec binds it, and a cache-busted re-fetch. Server-side caches ignore no-cache, which is why the timestamps are collected too.

Evidence for each run is written to .samethru/runs/<runId>/: run.json, the API response for each check, and for failures with a page, a screenshot of the element and the page and the element's HTML.

Coverage

Every report states how many values per page and fields per API response a check reads, and lists the uncovered ones. Coverage that couldn't be measured is reported as not measured, never as 0% or 100%.

  • Page values are found as visible text that looks like data (numbers, money, dates, percentages) or that equals a value in a JSON response the page received. Repeated items, such as a list column, are grouped: a repeated value counts once, and is covered only when every occurrence is read. Every report states this method.
  • API fields are every leaf JSON path in every response seen during the run. A field is covered when a check's API binding references it.
  • Excluded values come from coverage.ignore in the config. They're listed with their reasons and never count as covered.
  • Coverage is of the checks in the run: a value read only by a check that wasn't selected (--ids, --tags) isn't covered in that run's report.
  • samethru run --min-coverage <percent> exits with 2 when page coverage is below the minimum, or when a page couldn't be measured.

Bug catalogue

Every FAIL names the pattern it matches, with a plain-language explanation built from the real values, or says "unclassified mismatch" when none fits. When several patterns match, the strong match with the lowest precedence wins and the others are listed as alternatives.

# Pattern id Applies to
1 Rounding error rounding-error money, number
2 Cents vs dollars cents-vs-dollars money
3 Wrong decimal places wrong-decimal-places money, number
4 Timezone shift timezone-shift datetime
5 Date off by one day date-off-by-one date
6 List cut off by pagination pagination-cutoff list
7 Wrong count after filtering wrong-filtered-count count
8 Sort order differs between layers sort-order-mismatch list
9 Stale cached value stale-cache every value type
10 null, undefined or NaN shown as text null-as-text money, number, count, text, enum, datetime, date
11 Missing value shown as 0 missing-as-zero money, number, count
12 Wrong label for a status wrong-enum-label enum

1. Rounding error (rounding-error)

A price or total is off by a cent (or one unit in the last digit shown) because of float maths or the wrong rounding.

What it looks like

  • The cart total shows $35.79 but the items add up to $35.80.
  • A tax amount is a cent higher than the stored one.

How it's detected. The values differ by at most one unit in the last decimal place shown (strong), or two to three units (possible).

Evidence attached: Both exact values; The difference; The decimals shown.

Common causes

  • Adding prices as floating-point numbers (0.1 + 0.2 = 0.30000000000000004)
  • Truncating (Math.floor) instead of rounding
  • Rounding each line instead of the total, or the reverse

Precedence: 11.

2. Cents vs dollars (cents-vs-dollars)

An amount stored in minor units (cents) is shown as if it were in major units (dollars), or the reverse.

What it looks like

  • A $19.99 item shows as $1,999.00.
  • A $250 item shows as $2.50.

How it's detected. One value is exactly 10^n times the other, where n is the currency's minor digits (100 for USD).

Evidence attached: Both values; The exact factor; The units each layer declares.

Common causes

  • Passing price_cents to a formatter that expects dollars
  • Dividing by 100 twice, or not at all

Precedence: 4.

3. Wrong decimal places (wrong-decimal-places)

The number is right but shown with the wrong number of decimals.

What it looks like

  • A discount shows as $12.5 instead of $12.50.
  • A price shows four decimals.
  • $24.99 shown as "$25" on a site that only drops decimals for whole amounts.

How it's detected. The page shows a number of decimals that is not allowed (the spec, project default or currency), or drops digits the amount has.

Evidence attached: The text shown; The allowed decimals; The decimals shown.

Common causes

  • Building the text by hand ("$" + amount, toString())
  • A formatter missing minimumFractionDigits

Precedence: 10.

4. Timezone shift (timezone-shift)

A time is off by whole hours because a UTC time is shown as local time, or the reverse.

What it looks like

  • An order placed at 2:30 PM shows 6:30 PM.
  • Every time on the page is five hours early.

How it's detected. The times differ by a whole number of hours (or a 30/45-minute offset), at most 14 h. Strong when the shift equals the browser timezone's offset from UTC on that date.

Evidence attached: Both instants in UTC; The shift; The browser timezone and its offset.

Common causes

  • Dropping the "Z" or offset before parsing
  • Storing local time in a UTC column
  • Converting twice

Precedence: 5.

5. Date off by one day (date-off-by-one)

A date shows one day early or late because it was converted across midnight between time zones.

What it looks like

  • A delivery date of 17 March shows as 16 March.
  • Birthdays show the day before.

How it's detected. Date-only values differ by exactly one day. Strong when the browser's timezone isn't UTC, so a midnight conversion explains it.

Evidence attached: Both dates; The browser timezone and its offset.

Common causes

  • new Date("2026-03-17") reads the date as midnight UTC
  • Formatting a UTC midnight in a local timezone

Precedence: 6.

6. List cut off by pagination (pagination-cutoff)

A list stops at the first page of results and gives the visitor no way to see the rest.

What it looks like

  • The page shows 20 products; the shop has 23.
  • An API returns 50 rows and no total or next link.

How it's detected. Items are missing downstream and nothing leads to them (no total or next marker from the API, no usable pager on the page). Strong when the count shown is a typical page size or the limit the request asked for.

Evidence attached: Counts at each layer; The missing items; The request and its limit; The pager state; A screenshot.

Common causes

  • Calling a paginated API once and rendering only the first page
  • A pager that is hidden or disabled

Precedence: 7.

7. Wrong count after filtering (wrong-filtered-count)

A result count does not match the filtered results, often because it shows the unfiltered total.

What it looks like

  • "23 results" above a list of 5 lighting products.

How it's detected. A count check with a declared filter disagrees with the filtered count. Strong when the count shown equals the unfiltered total (context.unfilteredCount).

Evidence attached: The filter; The count at each layer; The unfiltered count.

Common causes

  • Fetching the count without passing the filter
  • Counting before filtering

Precedence: 8.

8. Sort order differs between layers (sort-order-mismatch)

A list has the right items in the wrong order.

What it looks like

  • Orders listed "newest first" with 9 March above 29 March.

How it's detected. Two lists hold the same items in a different order, and the spec says order matters.

Evidence attached: The first position that differs; The items expected and found there.

Common causes

  • Sorting by the displayed text (formatted dates sort alphabetically)
  • Re-sorting data the API already sorted

Precedence: 9.

9. Stale cached value (stale-cache)

A cache keeps serving an old value after the data changed.

What it looks like

  • The dashboard still says 4 low-stock products after a restock brought it to 2.

How it's detected. The API (or page) differs from the database, and either a cache-busted re-fetch returns the database value (an HTTP cache), or the database changed more than a second after the response was generated, by Date − Age or Last-Modified (a server-side cache). Cache headers alone make it possible.

Evidence attached: Cache headers; When the response was generated; When the data changed; The value after cache-busting.

Common causes

  • An in-memory or Redis cache not invalidated on writes
  • A CDN or proxy with a long max-age

Precedence: 3.

10. null, undefined or NaN shown as text (null-as-text)

A missing or broken value is written out literally instead of being left empty.

What it looks like

  • "Brand: null"
  • "$NaN"
  • "Delivery: Invalid Date"

How it's detected. The page shows null, undefined, NaN, [object Object] or Invalid Date.

Evidence attached: The text shown; The value at the source; A screenshot.

Common causes

  • Interpolating a nullable field into a template
  • Arithmetic on a missing value

Precedence: 1.

11. Missing value shown as 0 (missing-as-zero)

"No value" is shown as zero, which looks like a real number.

What it looks like

  • A product with no reviews shows a rating of 0.0.
  • An unknown shipping cost shows $0.00.

How it's detected. The source has no value (null or absent) and the page shows zero.

Evidence attached: The source value; The text shown.

Common causes

  • value ?? 0 or value || 0
  • A database default of 0 for "unknown"

Precedence: 2.

12. Wrong label for a status (wrong-enum-label)

A status or category is shown with the wrong label.

What it looks like

  • A refunded order shows "Cancelled".
  • A status shows the raw code "on_hold".

How it's detected. The page's label differs from the source of truth for the stored value. Strong when it is another value's label or the raw code.

Evidence attached: The stored value; The label expected; The label shown; Where the labels come from.

Common causes

  • A copy-pasted entry in a frontend label map
  • A missing translation falling back to the code

Precedence: 12.

MCP tools and the audit workflow

Your AI tool starts npx samethru mcp and talks to it over stdio. It offers seven tools and one prompt. Everything a tool returns goes to your assistant, so results are small by default, with file paths rather than contents wherever possible.

Tool What it does
status Reports the project root, the config and any problems with it, the databases, the browser, the saved login and the saved checks. With probe, it also tries the app, each database (read-only verification included) and the browser.
capture_page Loads a page in the browser and returns the JSON responses it received, the values it shows, and hints for where each value appears in that JSON.
call_api Makes a request (GET or HEAD unless the endpoint is allowed) and returns the status, cache headers, a short preview and the path of the full body.
query_db Runs one read-only SELECT. It returns 20 rows by default (at most 200 inline) and writes up to 10,000 to a file.
save_check Validates a check spec and saves it to .samethru/checks/.
run_checks Runs the saved checks and returns each verdict, its findings and the coverage counts.
get_report Returns a run's report as Markdown or JSON: the summary, the failures, the coverage or all of it.

The audit prompt gives the assistant Samethru's workflow. It's /mcp__samethru__audit in the Claude Code CLI and /audit in Gemini CLI; Codex doesn't support MCP prompts, so the same workflow also reaches every assistant through the server's instructions and the audit skill that init installs.

The workflow's rules: never call a value fine unless its check ran and passed; don't change the app's code unless asked; prefer the returned file paths; and Samethru only reads. Its steps:

  1. Call status. If the config is missing or invalid, show the problem. If pages need a login and none is saved, ask for npx samethru login.
  2. Find the pages in the routes.
  3. Call capture_page on each page.
  4. Trace each value in the code: component, state, fetch, handler, query, column. Record each file:line in traced.
  5. Work out the type, units, timezones, deliberate pagination, filtered counts and last-changed columns from the code. Take enum labels and formats from the source of truth, not the rendering code under test.
  6. Spot-check bindings with query_db and call_api.
  7. Write one spec per value with save_check, fixing errors and addressing warnings.
  8. Call run_checks, then get_report.
  9. Report each FAIL (pattern, values per layer, where it first goes wrong, code location, evidence), each NOT_RUN and INCONCLUSIVE with what would resolve it, the values found but not checked, and Samethru's counts exactly.

Privacy

Samethru makes no network calls of its own and adds no new place your data goes. It has no telemetry, does no update checks, and calls no AI models. It talks only to the app, API and databases you configure. Non-local addresses need your explicit opt-in.

What it returns to your coding assistant (query rows, page text, short response previews) is handled like any other tool output. Most assistants send tool output to their model provider, so treat it the way you treat your assistant reading your files. To keep that small, Samethru returns 20 rows by default and keeps previews short. Full responses, screenshots and evidence are written to .samethru/runs/ on your disk, and only their paths are returned. Secrets (environment-variable values, the saved login's cookies) are removed from everything it outputs.

Safety

  • Databases are read-only at more than one level. Only a single SELECT gets through Samethru's SQL check. SQLite connections are opened read-only and that is verified, not assumed. Postgres queries run in read-only transactions. doctor warns when the database role could write, because a read-only role is the real protection.
  • API calls are GET and HEAD unless you allow specific endpoints (POST /api/search) in the config. GraphQL endpoints accept queries only, never mutations.
  • Only local addresses are contacted unless you add a host to allowedHosts.
  • The saved login (.samethru/auth.json) is written readable only by you, gitignored by init, checked by doctor, and never printed or returned to your assistant.

Limitations

  • v1 only reads values that are reachable by URL. It doesn't click, switch tabs, fill in forms or press "load more". A value that only appears after an interaction can't be checked yet.
  • Lists split across pages are checked as paginated lists: the first page must be correct and the pager must be usable. Later pages can be checked if they have their own URL.
  • Databases: PostgreSQL and SQLite. MySQL is planned.
  • Pages are read in Chrome or Chromium only.
  • How well values are traced back to their API field and database column depends on your coding assistant. Samethru runs the checks deterministically, but it can only check the bindings it's given.
  • Coverage counts the values a page shows that look like data (numbers, money, dates, and text that appears in the page's JSON) or that a check reads. Every report states this method; plain prose isn't counted.
  • CI tests Linux and macOS. Windows should work (init writes cmd /c npx there) but isn't tested yet.

Requirements

  • Node.js 22.13 or later
  • Google Chrome, or Playwright's Chromium (init and doctor offer the download, and never download without asking)
  • PostgreSQL or SQLite
  • macOS or Linux
  • One of Claude Code, Cursor, Codex or Gemini CLI

Samethru is free and MIT-licensed.