9 indexes · 9 answering · 508,124 artefacts under measurement · version 1.0 · published 2026-08-14
How Toolproof measures
Written down before the numbers, so a figure you disagree with can be argued with on the method rather than on trust.
The rule: nothing is measured from a description of itself
A listing’s own README is a claim. A registry’s install count is a claim about popularity. A vendor’s pricing page is a claim about price on the day somebody read it. None of them is evidence that the thing works, and every catalogue in this space is built out of exactly those three.
Every index under this masthead is built the other way round. It fetches the artefact itself — the skill file, the repository, the registry item, the config, the document response — and records what happened when it did. Where an artefact cannot be fetched, that is the finding, and it is published as a finding rather than dropped.
- Primary source
- The artefact, fetched. A file read from the repository that publishes it, an API response recorded from a live call, a page loaded in a real browser.
- Never a source
- Star counts, install counts, search volume, a vendor’s own description, another directory’s listing, or a figure this site published yesterday.
- Recorded per run
- What was fetched, when, and what came back. A verdict is a function of stored signals, so the same artefact scored twice gives the same answer.
What counts as a failure, and why the denominator matters more
A failure is an artefact that was successfully fetched and then did not do the one thing it exists to do — a skill file that does not parse into something an agent could load, an extraction call that returned no usable answer, a registry item whose source has gone. A network error on our side is not a failure of the artefact and is never counted as one; it is retried, and if it keeps failing it is reported as unread rather than as broken.
The part that decides whether a failure rate means anything is the denominator. A catalogue that silently omits everything it could not process publishes a healthier ecosystem than exists, and does it without stating a single false number. Every index here keeps what it could not load inside the total it divides by. That is the entire reason a headline figure on this site is a failure share rather than a count of listings: the count is a number anybody can produce, and the share is one somebody had to go and check.
An index that cannot currently be read is named on the front page with the reason, and no number is printed in its place. There is no cached last-known value anywhere in this site’s code.
What each index does to get its number
Today’s figures- SkillWorks
- Each listing's files are read from source and scored 0–100 on four weighted components. A listing that does not parse into something Claude Code could load is counted as broken, not quietly dropped from the denominator.
- StillShipping
- Commit, release and issue activity pulled from the GitHub API nightly and turned into one of three verdicts. The verdict is recomputed every night rather than recorded once, so a project that goes quiet changes its own row without anybody editing it.
- ToolDrift
- Vendor changelogs, pricing pages and store rankings are fetched on a schedule and diffed against the previous capture. A price is recorded from the page that states it, and a page that has moved is followed and the redirect published rather than silently followed.
- KitGrade
- Kits are graded from measured facts rather than from their own landing pages, and the hands-on count is published separately from the documented count — a kit somebody installed and ran is a different claim from a kit somebody read about.
- StoreReady
- Every verdict is attached to numbered evidence — a policy clause, a rejection thread, a shipped binary. A builder nobody has evidence for is published as unknown rather than given the benefit of the doubt.
- BlockDex
- Each registry is crawled and its items are read individually, so an item that disappears between crawls is counted as removed rather than forgotten. Removals over the last thirty days are published beside the total.
- RuleStack
- Config files are read out of public repositories and measured: length, whether they carry runnable commands, whether they carry code. The format shares are computed from the files found, not from a survey of what people say they use.
- CiteRank
- A question is put to a live answer engine through a real browser and the answer it returns is recorded — one row per question per engine per run. Nothing is inferred from search rank; a citation counts when the engine actually produced it.
- doc-extract-bench
- Every vendor is run against the same pinned subsets and EVERY RAW RESPONSE IS COMMITTED, so the published accuracy can be recomputed offline by anyone with no API keys and no account. It is the only one of the nine whose numbers a stranger can fully re-derive.
Conflicts of interest, stated rather than implied
Independence is a claim like any other and it is worth exactly as much as the disclosure behind it. Every comparable page checked on 14 August 2026 — including the most-cited benchmark of model performance in the industry — carries no conflict disclosure of any kind. Here are ours, in full.
- Ownership
- All nine indexes are built and operated by Kynth Studios. They are not independent of each other. What they are independent of is the tooling they measure: no vendor in any of these indexes pays for placement, and no ranking on any of them can be bought.
- A benchmark that includes our own product
- doc-extract-bench scores Kynth Core against AWS Textract, Google Document AI, Veryfi and LlamaParse. Kynth Core is a Kynth Studios product, and it wins several of the comparisons. That is a conflict and it is why every raw vendor response in that benchmark is committed to a public MIT-licensed repository: the result can be recomputed offline by anybody with no API keys and no account, and a rigged number would not survive it. If you only check one figure on this site, check that one.
- Sponsorship
- One of the nine, SkillWorks, sells sponsor placements. A sponsor slot is rendered as a sponsor slot and cannot move a score, a rank or a verdict — the score is a function of stored signals and there is no field a payment writes to. The other eight carry no paid placement of any kind, and several say so on their own pages.
- Measuring our own estate
- The same probe method CiteRank sells is run against Kynth Studios’ own sites. On 2026-08-14 it put 23 buyer questions to five answer engines; 92 of the 115 came back with an answer, and this estate was cited in none of them. That is a bad number about us, measured by us, published on the same terms as any other. If the method flattered its owner it would not produce it.
What these numbers do not claim
- Not quality
- A skill that loads is not a good skill. Every index here measures whether something functions, not whether it is worth using. Nothing on this site is a recommendation.
- Not completeness
- Coverage is what the crawlers reached. A private repository, a paywalled registry or a listing published after the last run is not in the total, and a coverage figure is a floor rather than a census of everything that exists.
- Not a moment in time you can assume
- Every figure carries the date of the run that produced it. Where an index reports its own pipeline as behind schedule, that flag is passed straight through to the page rather than hidden.
- Not a security audit
- Execution-tested means we checked that an artefact loads and behaves, not that it is safe. A skill can load perfectly and still be malicious. Treat a passing figure here as the floor it is.
Corrections, and the version of this method
This document is versioned. A change to how something is measured gets a new version number and stays listed here with its date, because a figure cited last month was produced under the method that was published last month, and silently rewriting the rules underneath a citation is how a measurement project stops being one.
- Version 1.0
- 2026-08-14 — first publication. Nine indexes brought under one stated method.
- Reporting an error
- A figure you can show to be wrong gets corrected and the correction is recorded here with its date, whichever direction it moves the number. hello@kynth.studio.
- Reuse
- Figures on this site may be quoted and republished with attribution to Toolproof and the date of the run. The method text on this page is published under CC BY 4.0.
