DeFi Security AllianceRequest an audit
Menu

Personas and market

Crypto Security Score: Reading Methodology, Evidence and Unknowns

A rating is useful when its inputs answer the risk question you actually have. A crypto security score should lead to dated evidence about code, authority and operations; the aggregate alone cannot establish deployment safety or investment suitability.

Circular gauge with a blue needle and layered rings with gaps, representing crypto security scores and the risks a rating misses.

Key facts

Original census
28 indicator rows; 0 numerical weight coefficients in the public table
Category scope
6 categories, including code, governance, market and community signals
Reproducibility limit
Qualitative labels do not reveal exact numerical contributions
Decision boundary
Unknown evidence and disqualifying controls remain separate from the score

How scores are built

A crypto security score compresses selected observations about a project into a number or grade. The result depends on the inputs, their weights, missing-data treatment and the date of the observation. It is an output of a methodology. It is not a direct measurement of how much money a user can safely deposit or the probability that a contract will resist an attack.

Our census of CertiK's public Skynet methodology found 28 indicator rows across 6 categories. Every weight cell used a qualitative label. None disclosed a numerical coefficient. The public table therefore lets a reader inspect which subjects contribute to the score, but it does not provide enough numerical information to reproduce the aggregate from those rows alone.

That is a disclosure observation, not evidence that the score is calculated incorrectly. A proprietary methodology can still be useful. The decision for the reader is whether the available explanation supports the use they intend to make of the number. Screening projects for further investigation requires less certainty than treating a score as a calibrated estimate of loss.

The fetched Skynet methodology combines Code Security, Fundamental Health, Operational Resilience, Governance Strength, Market Dynamic and Community Trust. Some indicators concern code review. Others concern team transparency, trading activity or public engagement. Two projects can therefore arrive at a similar aggregate while presenting different underlying weaknesses.

The public weight table uses labels rather than coefficients
Weight labelIndicator rowsWhat the row count establishes
High5Count of published indicators with this label
Medium High6Count of published indicators with this label
Medium Low6Count of published indicators with this label
Medium7Count of published indicators with this label
Low4Count of published indicators with this label

A row count is not a weight allocation. Five indicators labeled high do not imply that those indicators receive equal numerical weights or that their combined share is five times that of a low-labeled row. The table also cannot reveal normalization, interactions or caps that may be applied elsewhere in the methodology. Reverse engineering those rules from labels would manufacture precision.

An analyst should preserve the score's timestamp and the displayed category breakdown together. A historical screenshot of the aggregate without its inputs may be impossible to interpret after a methodology change. Record the project identifier as well: a token symbol can refer to more than one asset, and a protocol can operate several deployments with different dependencies or administrators.

When comparing a project with its earlier score, check whether the methodology and data coverage stayed comparable. A newly available report can add evidence without changing the code, while a revised scoring rule can move the number without a project event. Keep those explanations distinct in a historical note. If the provider does not expose enough change history to separate them, record the cause as unresolved rather than attributing the movement to better or worse engineering.

The first reproducibility question is simple: could another reviewer retrieve the same evidence and understand why the rating changed? That does not require access to every proprietary coefficient. It does require enough provenance to separate a project change from a data refresh, a newly published audit or a revision to the scoring method.

Keep the underlying audit report evidence separate from the rating that incorporates it. The report's scope, reviewed revision and unresolved findings remain relevant even when the score changes. A number cannot repair a mismatch between the reviewed contracts and the deployment being evaluated.

What they measure

The word "security" covers several distinct questions in rating products. Code review asks whether an implementation contains a defect within the reviewed scope. Operational analysis asks whether the project can maintain and defend the surrounding system. Governance analysis asks who can change rules or control assets. Market and community indicators describe additional conditions that a methodology may treat as related to risk.

A useful reading starts at the category level. If the code assessment is strong but the administrator path is unclear, the aggregate should not close the governance question. If public engagement is low, that observation does not establish a contract vulnerability. The rating is more informative when each signal leads to its own evidence request.

Turn a category into a verifiable follow-up
Signal areaUseful evidence requestConclusion the signal cannot establish alone
Code reviewReport, revision, scope and remediation statusAll deployed code was examined
Operational resilienceIncident procedures and current infrastructure controlsThe response team will contain the next incident
GovernanceRole graph, signer policy and upgrade pathNo authorized actor can harm users
Market activityData source, observation window and asset identityThe contract is free of implementation defects
Community engagementMeasurement definition and manipulation controlsThe project is trustworthy or suitable for a particular investor

DeFiSafety's published Process Quality Review documentation describes a different assessment object. Its questions emphasize the documented process around a protocol, including the visibility of code, testing and security practices. That can expose evidence gaps a contract-level report does not address. It should not be read as an interchangeable implementation of another vendor's numerical scale.

The denominator matters here as much as it does in a vulnerability study. A process review can only evaluate evidence it can access under its method. A private internal procedure may exist without appearing in public documentation. Conversely, a published procedure may not be followed consistently. The observed result is about the evaluated evidence and criteria, not an omniscient view of the organization.

Separate asset, protocol and company scopes. A token rating can include holder concentration and market behavior. A protocol review may examine deployment architecture and operational controls. A company-level security assessment may include infrastructure and staff access. Using the same word for these three entities makes a comparison appear cleaner than it is.

Read the CertiK company analysis as provider context, then open the rating methodology and the project-specific evidence. The audit company analysis hub helps distinguish service models. Neither a member profile nor an award is a substitute for the inputs behind a particular project's rating.

Treat a rating's stated purpose as a boundary. If it is designed as a broad project-health signal, using it to rank the correctness of two implementations requires additional evidence. The same applies in reverse: a narrow code review cannot settle questions about exchange custody, insider access or treasury operations. The entity and the question should match before any numerical comparison begins.

What they miss

An aggregate can conceal a disqualifying condition. Suppose a hypothetical project has extensive public documentation and an active community, but an administrator can replace the withdrawal logic immediately. The correct next step is to understand that authority and its controls. Averaging it with unrelated strengths does not make the authority disappear.

Missing evidence needs its own state. A rating interface may show a blank field, an unavailable value or no report link. Those outcomes should not be translated into "no problem found." Ask whether the system treats the field as unknown, excludes it from a denominator or assigns a default. Without that explanation, two apparently similar scores may reflect different amounts of information.

Current deployments can also outrun published evidence. An audit may apply to a previous implementation, a paused product or a subset of contracts. The presence of a report is useful only after its scope is joined to the deployed addresses and revision. A methodology that records audit history is not necessarily performing that join for every component a user interacts with.

A score supports investigation only through its evidenceThe decision path branches away from the aggregate when a required fact is unknown or a control violates the reader's acceptance rule.Observed inputMethod ruleCategory resultDecision boundary
A score supports investigation only through its evidence. The decision path branches away from the aggregate when a required fact is unknown or a control violates the reader's acceptance rule.
  1. Identify the source, date and entity behind each relevant signal.
  2. Read how the methodology treats weighting and missing information.
  3. Inspect the category that corresponds to the actual risk question.
  4. Keep disqualifying conditions and unresolved evidence outside the aggregate.

A score is not a probability unless the provider demonstrates that interpretation. A number on a scale from zero to one hundred does not imply a corresponding chance of avoiding loss. Calibration would require a defined outcome, an observation horizon, a population and a method for handling projects that change or disappear. Our methodology-table census measures none of those things.

Claims that lower-scored projects experience more adverse events should be attributed to the provider unless an independent study is reproduced. Even a real association would leave questions about causation and selection. Projects with more public information may be easier to score, while incidents can change the score after the event. A retrospective association does not automatically validate a prospective investment rule.

Commercial relationships deserve specific disclosure questions. Ask whether the provider sold an audit or another service to the project, whether that relationship affects evidence availability and what separation exists between commercial work and rating decisions. Do not infer that a project purchased a higher score from the mere fact that it purchased an audit. That allegation requires evidence of its own.

Unknowns that should survive the aggregate
Unresolved itemWhy it mattersUseful next action
Reviewed revision does not match deploymentThe available report may describe different codeRequest a deployment-to-review mapping
Administrative authority is undocumentedAuthorized changes may bypass expected behaviorObtain the current role and upgrade graph
Weighting or missing-data rule is unavailableThe aggregate cannot be independently reconstructedUse the category evidence without inventing coefficients
Incident history lacks a primary accountCause and remediation may remain uncertainFind the protocol report and distinguish confirmed facts from allegations

There is also a time problem. A score observed before an upgrade cannot establish the safety of the new release. An incident resolved last year does not establish that current keys and infrastructure follow the same controls. The review record should therefore identify events that require another look: implementation changes, role changes, a new dependency or a material change in the methodology.

These limits do not make ratings useless. They define where a rating is an efficient index and where the reader must leave the index to examine evidence.

Reading a score before investing

Use the score to organize due diligence, not to determine an allocation. Start by writing the concrete question: whether a deployment matches its review, whether administrators can change balances or whether a reported issue remains unresolved. A general desire to find a "safe project" is too broad to test against a single rating.

Preserve the visible result before following links. Record the URL, timestamp, project identity, category values and methodology version where available. Then verify the particular asset and chain you intend to evaluate. A token page can be relevant background while still failing to describe the lending vault, bridge wrapper or custody arrangement through which exposure is obtained.

A review note that keeps decisions separate from scores
FieldExample of an adequate entryReason to stop and investigate
EntityNamed deployment and chain, with contract addressThe rating applies to a different token or product
Code evidenceReport revision mapped to deployed implementationNo verified mapping is available
AuthorityDocumented mint, pause and upgrade rolesA material privileged path remains unknown
MethodNamed rating methodology and observation dateThe number cannot be connected to a documented method
Unresolved conditionSpecific missing evidence with an ownerThe condition conflicts with the reviewer's acceptance rule

Read the findings rather than only the presence of an audit badge. A report can contain an accepted issue, an unreviewed fix or a scope limitation. These are different states. The crypto due diligence checklist provides a broader evidence trail, while the fake audit report guide addresses authenticity before interpretation.

When ratings disagree, compare their objects and methods before asking which number is right. One service may assess public process documentation while another combines code, governance and market signals. Different results can follow from different questions. If the methodologies claim to measure the same thing, compare data freshness, missing fields and the underlying report links.

Do not average incompatible ratings. A numerical average of a process-quality grade and a broad project-health score creates a new metric with no defined meaning. If a team needs a decision matrix, keep the criteria separate and document the acceptance rule for each. A disqualifying authority or an unknown deployment should remain visible rather than being diluted by unrelated positive signals.

A compact decision memo can end with one of three evidence states: sufficient for the stated review question, contradicted by a material finding or incomplete pending a named document. These are editorial workflow states, not investment recommendations. They make it possible for another reviewer to understand why the process stopped or continued.

Revisit the memo when a triggering change occurs. A new audit can close a known evidence gap; an upgrade can reopen one. The saved score remains a historical observation. The current decision should be based on current scope and evidence, with the unresolved items still visible.

Original disclosure census: 28 indicator weights

Method

Source and selection
Read every indicator row in the public Skynet methodology weight table. Exclude the header. A weight is numerical only when its weight cell contains a digit.
Retrieved
Observation
The complete indicator table was inspected; only its header was excluded.

Results

Recorded observations on
Disclosure observationCountDenominator
Indicator rows28All rows in the public indicator table
Qualitative weight cells2828 indicator rows
Numerical weight cells028 indicator rows
Category groups6Distinct category groups in that table

Our recorded result is 0 numerical weight cells among 28 indicators. It identifies a reproducibility limit in the published table, not a defect in a particular project or the vendor's calculations.

The appropriate use of this result is narrow. A reviewer can check whether the methodology considers a relevant subject and can ask the provider how that subject affects the score. The reviewer cannot derive exact contributions by converting labels such as high or medium into invented percentages. No project score was recalculated in this census.

Limits

  • One provider's public methodology table, retrieved on the recorded date.
  • No rating algorithm, private data, exploit outcome sample or predictive performance was measured.
  • Indicator counts and qualitative labels may change; the saved evidence fixes the observation used here.

The reproducible record is kept in the repository: seo/research/articles-42-51-2026-09-06/surveys.py and seo/research/articles-42-51-2026-09-06/survey-43.json. These file paths are not public downloads.

Frequently asked questions

Can a score be used in an automated procurement filter?

It can trigger manual review if the filter records the provider, methodology and entity. Keep an exception path for missing data and require direct evidence for any control that is a procurement condition.

What if a project has no rating at all?

Record the rating as unavailable. Evaluate the relevant reports, deployment configuration and operational evidence directly; absence of a rating does not establish either a defect or a clean result.

Can a rating change while the contract code stays identical?

Yes, a broad methodology may incorporate new documentation, governance, market or operational information. Check the category changes and methodology notes before attributing the movement to a code change.

Comments

5
  1. Lena Y.

    The missing numerical weights are a useful limit to state. A reader can inspect the disclosed categories without pretending to reproduce the exact aggregate from qualitative labels.

  2. Sara L.

    How does the score distinguish unavailable evidence from evidence that a control is absent? Those states lead to different follow-up questions even when the displayed result is similar.

  3. Oscar P.

    An immediate upgrade path deserves a separate decision in the checklist. A strong community indicator does not explain who can replace withdrawal behavior.

  4. Priya S.

    I would save the methodology version with the score snapshot. Otherwise a later change in the number could reflect a scoring change rather than a change in the protocol.

  5. Arjun P.

    The linked evidence matters more to this workflow than the headline grade. A stale report link should lead to a deployment check, not simply another comparison of scores.

Leave a comment

Share a question or observation about this article.

10 to 3,000 characters.