How the score is calculated
Six categories add up to 100 points. Thresholds do not depend on the day: the same project state always gets the same score.
What makes up the 100
a category's contribution = its weight × the share earned inside it- 20Documentation and best practices
- 20Project activity
- 20Security
- 15CI/CD
- 15Issues
- 10Code health
Three rules that protect projects from unfair scores
If a category cannot be assessed, its score is not reduced. Its weight is shared among the others. Likewise, a GitHub mirror is not penalised for missing branch and review policies: they cannot be set here. A configured policy is counted as is. A mirror means mirroring is enabled on the platform. A repository imported from GitHub whose mirroring has stopped is scored as a regular one.
We asked and got no answer. The score then becomes a “from–to” interval instead of a zero.
The heaviest category weighs 20 out of 100. A mature project stays high even with rare commits.
Six categories
tap a category to see its metrics and scalesopen a category to see its metrics and scalesWhat we look at: the repository tree and platform settings. If the platform returned a truncated tree, a missing file is marked “not checked” rather than “absent”. A GitHub mirror is not penalised for missing branch and review policies: the metric drops out.
Observation window — 365 days.
What we look at: commits, authors, releases and pull requests within the window. Responsiveness is counted from the first page of the list and only with enough objects. One or two are too few to judge, and the metric drops out.
Nothing is visible publicly: the platform shows scanner reports only to repository members (403 by permissions, measured 15.09.2026). So for a public repository the category is “no data”, and its weight goes to the others. The breakdown of findings by severity and state switches on when you sign in to your own repository.
“CI runs” is counted only in the private analysis, with the owner’s token: the window is 90 days, at most the 50 latest observations; with fewer than 5 the metric does not apply and its weight goes to the neighbours. The scale steps are a norm set on 23 September 2026, not a calibration on a corpus.
What we look at: the build config file and its triggers, and in the private analysis also run results (the share of successful runs among finished ones). What we cannot see publicly: run results — the platform shows them only to repository members (403 by permissions, measured 15.09.2026). So for a public repository the runs metric is always “collection failed” — closed by the platform — and the category is “partially scored”. If the config itself could not be read, the category is “collection failed”. We look only at the platform config (.sourcecraft/ci.yaml); external CI (GitHub Actions etc.) is not checked.
Responsiveness for “issue resolution time” is counted over a sample of up to 50. At least 3 eligible objects are needed: one or two are too few to judge, and the metric drops out. “age of open issues” — thresholds are quartiles of the catalogue: 609 repositories, measured 15 September 2026 “issue resolution time” — thresholds are quartiles of the catalogue: 240 repositories, measured 15 September 2026
What we look at: tracker reports and the time to reply to them. Not everyone has the tracker enabled: it is empty for 97.5% of the public catalogue (full-catalogue measurement, 15.09.2026). Then there is nothing to ask, and the category goes to “no data”, not to zero.
“TODO/FIXME density” — thresholds are quartiles of 372 non-zero observations out of 13,766 catalogue repositories, measured 17 September 2026
What we look at: every code file in the collected tree, with no sampling. Almost no measured repository has debt markers at all. So the middle of the scale sits on the quantiles of those that do: otherwise the scale would measure a rare exception.
Grades
lower score bound- A85+excellent
- B70+good
- C55+not bad
- D40+room to grow
- Fbelow 40starting point
An interval score is shown as a pair of letters: the grades of its lower and upper bounds, for example C–B. The colour follows the lower one.
Interval and coverage
an interval beats a point- scored84
- partially scored61–78
- not scored—
If a metric could not be measured, the score becomes an interval: from “all lost” to “all earned”. Coverage is the share of the rubric's weight we managed to apply. Below 60% the score counts as partial, even if the interval bounds coincide. An empty repository is not scored at all. A repository with long-stale data is not scored either: instead of a stale score it shows “not scored”.
This is how much weight can be applied to a public repository. Security and Issues cannot be assessed, so their weight goes to the other 4 categories. A full score is never possible on public data: CI run statuses are closed. So instead of “scored / partial” we show coverage as a number.
An interval is more honest than a point. It shows how much we could not measure and avoids false precision. The rank uses the lower bound of the interval, that is, only what is proven. We do not use the middle: otherwise a repository whose data the platform returned less of would rank higher than its measured data shows. For large repositories the range is usually wider: the platform more often returns their data incompletely — the file tree, a long history, a large tracker — so more of their weight stays unmeasured.
What is deliberately excluded from the score
Threshold calibration
The middles of the scales sit on the quartiles of the sample, and the edges are set by hand: “30 days is still fresh”, “a year is already old”.
The date and sample above apply only to scales without their own caption; the rest have their own measurement (4 metrics), named in the metric's row.
How the score adds up
four steps from a metric value to a grade letter- The metric value is mapped to 0–100 by its scale. Between nodes — linearly, beyond the outer nodes — a plateau.
- Metric scores inside a category are averaged by their weights. The sum of “score × weight” is divided by the sum of weights.
- The category's share is multiplied by its weight. The contributions of the six categories add up to a number from 0 to 100.
- The grade is taken from that number by the grade boundaries. For an interval, a letter is taken for each bound, and the score shows the pair.
A category with no data and a metric with no answer are counted differently. Both cases are described above, in the three rules.
Response and tracker: what counts as a reply
definitions taken from the current collection code- A reply
- The first comment from an identified person other than the pull request or issue author. Code-assistant comments and deleted comments do not count.
- Who is left out of the denominator
- Objects that could not be read in full, and objects whose comment author is unknown. For them we cannot say “answered” or “not answered”.
- The minimum denominator
- The threshold is shown in the caption of an expanded category, from the same response as the scales. On a smaller sample the “unanswered” share jumps around, and the score would measure the sample, not the project. When the sample falls below the threshold because it was not read to the end — the run was stopped, or the platform refused part of the targets — the metric becomes “unavailable” with an interval instead of dropping out of the score.
How often it is recalculated
The service is designed for a daily recalculation of the catalogue. Repositories that changed since the last successful collection are refreshed. The rest stay in the snapshot with their previous data and previous analysis time.
The four category states
the state chip stands in the category row of a repository card- scored
- Every applicable metric is measured. The category's contribution is counted in full.
- partial
- Some metrics could not be measured. The category keeps its full weight, and coverage shows the gap as a number.
- no data
- Nothing to assess for this repository: no applicable metric, for example the tracker is off or the data source is closed. The weight goes to the other categories, with no penalty.
- collection failed
- The metrics apply, but none could be measured. The weight counts at the lower bound, and the score becomes an interval.