AI assistants now hand real work to software written by strangers: look up a price, book a slot, move money. Usually the only thing telling you that software is what it claims to be is that software. So we built a checker that measures it instead, publishes what it found, and gives you the arithmetic to redo the answer without us. Point it at any server. It needs nothing from you but a URL.
Paste the URL of any MCP server: yours, ours, or one you have never met. The verdict comes back with a SHA-256 that this page recomputes in your browser and compares. If the two do not match, we altered the verdict and you just caught us.
curl -s -X POST https://gate.horizonshield.dev/check \
-H 'content-type: application/json' \
-d '{"endpoint":"https://your-server/mcp"}'
One machine, one register, one set of five conditions. What you want from it depends entirely on which side of the connection you are standing on.
No account, no application, nobody to persuade. A badge that is drawn when it is asked for, so it stops being green the moment your row does. And a plain list of what to change when a condition is not met.
社内にAIエージェントを入れる前に、繋ぎ先を第三者が読み取り専用で測ります。誰が運営しているか、金の出所はどこか。稟議に添付できる形で返します。
The check is read only, so it works on a server we do not own. We only ever pointed it at volunteers, and that was the mistake. The method is published. There is no number yet.
On the first run, our own servers did not meet a single one of the rules we had just published. We fixed what we could and left the rest visible. What the table below says about us today is whatever the register says today. This sentence deliberately does not tell you, because a sentence typed once cannot, and a page that states a verdict in fixed type goes stale while the thing it describes keeps moving. The checker itself is in that table under the same rule. Until 2026-08-15 it could not reach its own hostname at all, refused to count that unmeasured condition as a pass (the same rule it applies to everyone else), and its row sat as held · http 522. Then an outside runner found the mirror image of that limit: on-demand checks of our own servers were failing with 522 too, while the nightly sweep quietly passed. We mapped the boundary, routed those probes over the public edge through a relay, and every verdict now names the path that measured it in a probed_via field. On 2026-08-23 the relay was moved off Cloudflare entirely, and the gate's own row went green for the first time: reachable: true · verified, on demand and on the sweep alike. The held · http 522 records stay in the history rather than being tidied away, because a register that edits its own past is not a register.
If the people who write the test are the only ones who pass it, the test is decoration. That is the whole reason this thing is worth pointing at your server: it is allowed to come back red, including for us.
The ledger behind this register accepts verification walks from anyone. Acceptance is mechanical: schema, size, rate, duplicate, and signature validity if a signature is present. There is no route in the code by which the operator declines a schema valid submission. Until 2026-08-18 that was a description, because the only record in the pool had been written by the operator, about a system he had been invited to measure.
On 2026-08-18 the operator of invinoveritas submitted his own walk of the same target, taken from his own infrastructure, and wrote into his vantage field that it is not an independent network. Two records now sit in the pool from two positions. Both callers asserted an identity they do not hold: a reserved example domain in one case, the Ethereum burn address in the other. Neither choice was discussed beforehand. Both mismatches are permanent and queryable by anyone.
What this does not prove. Both records are unsigned and say so. Both concern one target, and three of its two hundred and forty four entries. Neither witness audited the other's code. A pool of two is not a track record. It is the smallest number at which the design stops being a claim about itself.
No human reviews an application, so there is no one to persuade, no queue to jump, and no reason to charge you. What is on the left is the question in plain words. What is on the right is exactly how it is measured, so you can dispute it.
| 01 | Is it actually there, and does it do anything? | Responds to initialize and tools/list with at least one tool. |
| 02 | Does it say who it is, in a place a machine can read? | /.well-known/agent-card.json returns JSON with a name and a description. |
| 03 | Does it say who is paying it? | The card declares paid_by, referral_fee, listing_fee. Only silence disqualifies. We do not judge whether the answer is flattering. |
| 04 | Does it answer the same way twice? | Identical input returns identical output. Not measured unless you ask: it means executing one of your tools, and the first tool a server lists may well be destructive. |
| 05 | Can you check our answer without trusting us? | Every verdict carries a record_sha256. Drop that field, stringify the rest in key order, hash it. It must match. |
A badge that quietly means more than it measured is worse than no badge. So here is the ceiling, stated at the same size as the claim.
The count and the breakdown are drawn from the register above rather than written here, because a number typed into a page is wrong the moment the register moves. That is not a soft launch we are dressing up. It is the state of the register, and hiding it would make every other row here worthless. Rows are re-measured on a schedule and the history stays public, so a change means a condition actually flipped, not that a fresh verdict was issued. The sweep calls no tool on anyone, us included.
Answers "who verifies the verifier" by going first, and then failed its own test in public until 2026-08-23. It now speaks MCP at /mcp, so the tool that lets you distrust the issuer is provided by the issuer. For a while its endpoint row carried pass: true on a condition it openly had not measured, which made its own verdict read verified. It was the only server this gate was lenient with, so we removed the leniency and the verdict fell to pending. On 2026-08-15, after babyblueviper1 reproduced our check from his own network and caught on-demand probes of our own zone failing with 522, those probes were rerouted over the public edge, and the gate measured its own endpoint for the first time: reachable: true, still pending, because determinism stays unmeasured without consent, including for us. The row stays visible either way.
Endpoint, card and disclosure all pass, measured. Determinism stays not measured, because the sweep calls no tool on any server and that includes ours. So it reads in process, not verified. We did not relax the rule for ourselves.
Both in progress. No listed party may call itself verified before it passes. In-process entries are marked pending and turn green only on passing. That is the core of the design and it is not relaxed for anybody.
This table is read from the gate's public register at /register every time this page loads. A member that joins appears here on the next load, with its latest verdict and full public history. Loading the register…
Everything measuring this page is public code: github.com/ogasurfproject-jpg/horizon-shield. If a register you cannot bribe is worth having, star it. Stars are the one metric on this page we happily accept from strangers.
You are not applying for approval. You are starting a clock, and it only runs forward.
We did, on every condition, first run. Ten seconds tells you which one and why. Finding out here is free and private until you choose otherwise. Finding out because a customer found it first is neither.
Every row on this list today is ours or a member firm we brought in. The first outside name on a register is remembered in a way the fiftieth is not, and it is the same one command either way. The cost of being early here is zero.
Six months from now a stranger opens a URL you do not control and reads every measurement of your server since the day you registered. Dated. Recomputable. Including the week you broke something and fixed it, which is the part that convinces them. You did nothing that day. It was already there.
Traffic. Our own listing on another MCP directory produced 93,983 impressions, 24 clicks, 421 profile views and zero tool calls in the thirty days to 2026-08-14, real numbers off our dashboard, not a guess, and dated because a number without a date quietly stops being true. That is ninety-four thousand chances to be chosen, and not one of them ended in a tool call. Until 2026-08-14 this line read 54,911 impressions as of 2026-08-09; that was a running total and this is the dashboard’s rolling thirty days, so the two are not a growth rate and we are not going to let them read like one. A directory is not a distribution channel, and anyone telling you otherwise has not measured it. What this gives you is a record. A record is worth something precisely because it takes time to make, which is the same reason it cannot be bought later.
A fix repairs one casting. The mould that cast it goes back in the drawer, ready to cast the same flaw into the next hundred, unless somebody goes looking. So there is a second ledger here, and it records four things about a fix: the class of assumption behind it, where the author searched for that same assumption, what they found at each place, and at what volume each casting failed.
A record whose search list is empty is accepted and published as such, marked in amber. An instance fixed with no class search is the exact thing this ledger exists to make visible. It is not rejected, not hidden, and not ranked below a thorough one, because nothing here is ranked at all. A ledger that only accepted diligence would record nothing but diligence and be worth nothing.
Since 2026-08-20 a record can also carry the commands that re-run its own search, and the commit they were run against. This gate still does not run them and does not vouch for them. What changed is that a stranger can run them, compare their own hit list against the locations the record names, and establish that the record is wrong without asking anyone. Records that carry no such commands say so of themselves.
Writing to it needs no key and no account: you open an issue, and the gate fetches that issue from GitHub itself. There is no shared write token to hand out, so there is nothing to forge and nothing to leak. Reading the ledger.
Read the ledger and record one → · GET https://gate.horizonshield.dev/mould
The category was named by Federico Blanco Sanchez-Llanos in "The Mould, Not the Letter", 2026-08-20. Every record in it so far is one of our own, because we were not going to ask anyone else to go first.
Everything above is also published as a repository that rebuilds itself from the same API once a day: mcp-conduct-register. A script writes the table, so nobody chooses the rows there either. The same run publishes a register.json snapshot and an llms.txt, and the repository carries a citation file, so the register can be quoted and cited without going through any page we control.
The live API underneath all of it answers directly, with no key: curl -s https://gate.horizonshield.dev/register
The checker speaks MCP itself, at https://gate.horizonshield.dev/mcp. Five tools, all read-only, no key: get_conditions, check_conformance, verify_verdict, lookup_server, is_verified. Connect it once, then ask in plain words. Nothing on this path differs from the curl above: same gate, same record, same hash, and your assistant can recompute it too.
One line in the terminal. Then ask Claude to check a server.
claude mcp add --transport http wedjat https://gate.horizonshield.dev/mcp
Settings, Connectors, Add custom connector. Name it anything. The URL is the only field that matters.
https://gate.horizonshield.dev/mcp
One click installs it. The link carries the config in the URL itself, nothing else.
{"mcpServers":{"wedjat":{"url":"https://gate.horizonshield.dev/mcp"}}}
One click, or paste the block into .vscode/mcp.json.
{"servers":{"wedjat":{"type":"http","url":"https://gate.horizonshield.dev/mcp"}}}
Settings, Connectors, Advanced, Developer mode, Create. Paste the URL. The reviewed app, MCP Conduct Register, is in OpenAI's queue; until it lands, developer mode is the path, and it is the same server.
https://gate.horizonshield.dev/mcp
Not an assistant, but the same loop where your server is built. Measures the endpoint, recomputes the hash inside your job, fails the build on a measured failure. Unmeasured stays unmeasured.
- uses: ogasurfproject-jpg/wedjat-check-action@v1
with:
endpoint: https://your-server/mcp
A zero dependency npm package that reads the register before an MCP client opens a connection, and applies a policy you choose: warn, measured, verified-only or off. verified is true or null, never false, so unmeasured never reads as failed.
npm i mcp-conduct
Plain JSON-RPC over HTTPS POST. No auth, no session state to keep, no SSE required. If your client can call initialize and tools/list, it can call this.
curl -s -X POST https://gate.horizonshield.dev/mcp \
-H 'content-type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'
A prompt that makes the assistant do the whole loop: measure, name the failed condition, quote the hash, and recompute it without trusting the gate.
Using the wedjat tools: run check_conformance on https://YOUR-SERVER/mcp with allow_tool_call set to false. Tell me which of the five conditions passed, failed, or were not measured, and why, in plain words. Quote the record_sha256. Then pass the full returned record to verify_verdict and tell me whether the recomputed hash matched. Do not take the gate's word for anything it did not measure.
It does not give the gate any access to your assistant, your files, or your other servers. The gate only ever reads the endpoint you name in the prompt, and it calls no tool there unless you set allow_tool_call yourself, for a server you own. Disconnecting is one click on the same settings page.
If your endpoint is on the register, you may display its current verdict on your own site. The image is generated from the register at request time and cached for five minutes, so a green cannot be kept up after the row stops being green. An endpoint that is not on the register reads not listed, which is a fact rather than a failure.
Replace the endpoint with your own and paste it anywhere:
Ours read verified, pending and not listed at the same moment this was written. For most of this register's life the amber one was the checker's own row. That is the point of showing the badge rather than a logo: the picture changes when the measurement does, including when it changes against us.
Nobody pays to appear, nobody pays for position, and the verdict itself is never sold. You do not need to pass before you start.
One command, no key, no signup. The verdict names the condition that failed and why, so you can fix it before anyone else looks. No tool on your server is called unless you add allow_tool_call: true.
Determinism is the one condition that needs a tool call, and the gate will not make one on an assertion: anyone can claim to own a server. Publish /.well-known/mcp-conduct.json on the origin with {"allow_tool_call": true} and the gate takes that as consent, because only the owner of an origin can place a file there. Every verdict then records where and when it read the file (consent_source: well_known). Optional, and skipping it only means determinism stays unmeasured, never failed. Since gate 0.2.4.
Public, with a template. Submissions are visible to everyone, including the ones that failed. There is no private queue and no fast lane. You may open one even if you failed. Ask why, and the answer is free.
We do not take your word for the result, and you should not take ours. We run it ourselves, read-only, and post the verdict with its record_sha256 in the same issue so you can recompute it. If it passes, the row goes up. If it does not, the reply says what to change, and that costs nothing either.
The verdict, the listing and the public history are free for every endpoint, with no cap and no account. What costs money is how often we measure you and whether we wake you up when something flips. Nothing you pay for changes the result.
| Free | $15 / month · $150 / year | |
|---|---|---|
| The verdict | identical | identical |
| Listing in the directory | free | free |
| Public, recomputable history | yes | yes |
| Re-measured | once a week | every day |
| Told when a condition flips | no, you go and look | webhook, within the hour |
| Endpoints covered | every endpoint we watch | up to 3 on one subscription |
The alert does not fire on a single unreachable check. A monitor that cries wolf once is never trusted again. It takes three consecutive unreachable measurements, and the message says which condition flipped and why. No setup fee, no minimum term, cancel any time. Full terms →
Our other paid layer, tracing every figure a server returns back to a primary source, exists today for Japanese construction only, because that is where the dataset, the thirty years on site, and the licensing sit. If you arrived here as an MCP developer outside that field, the honest answer is that there is currently nothing here to sell you beyond the $15 monitor, and the check itself is genuinely all you will be charged for, which is nothing.
The badge above is a picture. This is the same statement as text, for search engines and for language models that fetch this page. It is a state, not a score, and not a recommendation.
| status | verified |
|---|---|
| record sha256 | 4dffe218167d40bb5940ba6cd404225b95b3d20ca5cdac880fbd4c62f0e343fc |
| record | https://gate.horizonshield.dev/record/4dffe218167d40bb5940ba6cd404225b95b3d20ca5cdac880fbd4c62f0e343fc (the hashed bytes: sha256 of the body equals record sha256) |
| history | https://gate.horizonshield.dev/history?endpoint=https%3A%2F%2Fmcp.horizonshield.dev%2Fmcp |
| recompute recipe | https://gate.horizonshield.dev/spec |
| one read before connecting | https://gate.horizonshield.dev/register/lookup?endpoint=https%3A%2F%2Fmcp.horizonshield.dev%2Fmcp |
| as of | 2026-09-08T23:42:10.563Z |
Establishes
Does not establish