Personas and market
Crypto Security Score: Reading Methodology, Evidence and Unknowns
A rating is useful when its inputs answer the risk question you actually have. A crypto security score should lead to dated evidence about code, authority and operations; the aggregate alone cannot establish deployment safety or investment suitability.

Key facts
- Original census
- 28 indicator rows; 0 numerical weight coefficients in the public table
- Category scope
- 6 categories, including code, governance, market and community signals
- Reproducibility limit
- Qualitative labels do not reveal exact numerical contributions
- Decision boundary
- Unknown evidence and disqualifying controls remain separate from the score
How scores are built
A crypto security score compresses selected observations about a project into a number or grade. The result depends on the inputs, their weights, missing-data treatment and the date of the observation. It is an output of a methodology. It is not a direct measurement of how much money a user can safely deposit or the probability that a contract will resist an attack.
Our census of CertiK's public Skynet methodology found 28 indicator rows across 6 categories. Every weight cell used a qualitative label. None disclosed a numerical coefficient. The public table therefore lets a reader inspect which subjects contribute to the score, but it does not provide enough numerical information to reproduce the aggregate from those rows alone.
That is a disclosure observation, not evidence that the score is calculated incorrectly. A proprietary methodology can still be useful. The decision for the reader is whether the available explanation supports the use they intend to make of the number. Screening projects for further investigation requires less certainty than treating a score as a calibrated estimate of loss.
The fetched Skynet methodology combines Code Security, Fundamental Health, Operational Resilience, Governance Strength, Market Dynamic and Community Trust. Some indicators concern code review. Others concern team transparency, trading activity or public engagement. Two projects can therefore arrive at a similar aggregate while presenting different underlying weaknesses.
| Weight label | Indicator rows | What the row count establishes |
|---|---|---|
| High | 5 | Count of published indicators with this label |
| Medium High | 6 | Count of published indicators with this label |
| Medium Low | 6 | Count of published indicators with this label |
| Medium | 7 | Count of published indicators with this label |
| Low | 4 | Count of published indicators with this label |
A row count is not a weight allocation. Five indicators labeled high do not imply that those indicators receive equal numerical weights or that their combined share is five times that of a low-labeled row. The table also cannot reveal normalization, interactions or caps that may be applied elsewhere in the methodology. Reverse engineering those rules from labels would manufacture precision.
An analyst should preserve the score's timestamp and the displayed category breakdown together. A historical screenshot of the aggregate without its inputs may be impossible to interpret after a methodology change. Record the project identifier as well: a token symbol can refer to more than one asset, and a protocol can operate several deployments with different dependencies or administrators.
When comparing a project with its earlier score, check whether the methodology and data coverage stayed comparable. A newly available report can add evidence without changing the code, while a revised scoring rule can move the number without a project event. Keep those explanations distinct in a historical note. If the provider does not expose enough change history to separate them, record the cause as unresolved rather than attributing the movement to better or worse engineering.
The first reproducibility question is simple: could another reviewer retrieve the same evidence and understand why the rating changed? That does not require access to every proprietary coefficient. It does require enough provenance to separate a project change from a data refresh, a newly published audit or a revision to the scoring method.
Keep the underlying audit report evidence separate from the rating that incorporates it. The report's scope, reviewed revision and unresolved findings remain relevant even when the score changes. A number cannot repair a mismatch between the reviewed contracts and the deployment being evaluated.
What they measure
The word "security" covers several distinct questions in rating products. Code review asks whether an implementation contains a defect within the reviewed scope. Operational analysis asks whether the project can maintain and defend the surrounding system. Governance analysis asks who can change rules or control assets. Market and community indicators describe additional conditions that a methodology may treat as related to risk.
A useful reading starts at the category level. If the code assessment is strong but the administrator path is unclear, the aggregate should not close the governance question. If public engagement is low, that observation does not establish a contract vulnerability. The rating is more informative when each signal leads to its own evidence request.
| Signal area | Useful evidence request | Conclusion the signal cannot establish alone |
|---|---|---|
| Code review | Report, revision, scope and remediation status | All deployed code was examined |
| Operational resilience | Incident procedures and current infrastructure controls | The response team will contain the next incident |
| Governance | Role graph, signer policy and upgrade path | No authorized actor can harm users |
| Market activity | Data source, observation window and asset identity | The contract is free of implementation defects |
| Community engagement | Measurement definition and manipulation controls | The project is trustworthy or suitable for a particular investor |
DeFiSafety's published Process Quality Review documentation describes a different assessment object. Its questions emphasize the documented process around a protocol, including the visibility of code, testing and security practices. That can expose evidence gaps a contract-level report does not address. It should not be read as an interchangeable implementation of another vendor's numerical scale.
The denominator matters here as much as it does in a vulnerability study. A process review can only evaluate evidence it can access under its method. A private internal procedure may exist without appearing in public documentation. Conversely, a published procedure may not be followed consistently. The observed result is about the evaluated evidence and criteria, not an omniscient view of the organization.
Separate asset, protocol and company scopes. A token rating can include holder concentration and market behavior. A protocol review may examine deployment architecture and operational controls. A company-level security assessment may include infrastructure and staff access. Using the same word for these three entities makes a comparison appear cleaner than it is.
Read the CertiK company analysis as provider context, then open the rating methodology and the project-specific evidence. The audit company analysis hub helps distinguish service models. Neither a member profile nor an award is a substitute for the inputs behind a particular project's rating.
Treat a rating's stated purpose as a boundary. If it is designed as a broad project-health signal, using it to rank the correctness of two implementations requires additional evidence. The same applies in reverse: a narrow code review cannot settle questions about exchange custody, insider access or treasury operations. The entity and the question should match before any numerical comparison begins.
What they miss
An aggregate can conceal a disqualifying condition. Suppose a hypothetical project has extensive public documentation and an active community, but an administrator can replace the withdrawal logic immediately. The correct next step is to understand that authority and its controls. Averaging it with unrelated strengths does not make the authority disappear.
Missing evidence needs its own state. A rating interface may show a blank field, an unavailable value or no report link. Those outcomes should not be translated into "no problem found." Ask whether the system treats the field as unknown, excludes it from a denominator or assigns a default. Without that explanation, two apparently similar scores may reflect different amounts of information.
Current deployments can also outrun published evidence. An audit may apply to a previous implementation, a paused product or a subset of contracts. The presence of a report is useful only after its scope is joined to the deployed addresses and revision. A methodology that records audit history is not necessarily performing that join for every component a user interacts with.
- Identify the source, date and entity behind each relevant signal.
- Read how the methodology treats weighting and missing information.
- Inspect the category that corresponds to the actual risk question.
- Keep disqualifying conditions and unresolved evidence outside the aggregate.
A score is not a probability unless the provider demonstrates that interpretation. A number on a scale from zero to one hundred does not imply a corresponding chance of avoiding loss. Calibration would require a defined outcome, an observation horizon, a population and a method for handling projects that change or disappear. Our methodology-table census measures none of those things.
Claims that lower-scored projects experience more adverse events should be attributed to the provider unless an independent study is reproduced. Even a real association would leave questions about causation and selection. Projects with more public information may be easier to score, while incidents can change the score after the event. A retrospective association does not automatically validate a prospective investment rule.
Commercial relationships deserve specific disclosure questions. Ask whether the provider sold an audit or another service to the project, whether that relationship affects evidence availability and what separation exists between commercial work and rating decisions. Do not infer that a project purchased a higher score from the mere fact that it purchased an audit. That allegation requires evidence of its own.
| Unresolved item | Why it matters | Useful next action |
|---|---|---|
| Reviewed revision does not match deployment | The available report may describe different code | Request a deployment-to-review mapping |
| Administrative authority is undocumented | Authorized changes may bypass expected behavior | Obtain the current role and upgrade graph |
| Weighting or missing-data rule is unavailable | The aggregate cannot be independently reconstructed | Use the category evidence without inventing coefficients |
| Incident history lacks a primary account | Cause and remediation may remain uncertain | Find the protocol report and distinguish confirmed facts from allegations |
There is also a time problem. A score observed before an upgrade cannot establish the safety of the new release. An incident resolved last year does not establish that current keys and infrastructure follow the same controls. The review record should therefore identify events that require another look: implementation changes, role changes, a new dependency or a material change in the methodology.
These limits do not make ratings useless. They define where a rating is an efficient index and where the reader must leave the index to examine evidence.
Reading a score before investing
Use the score to organize due diligence, not to determine an allocation. Start by writing the concrete question: whether a deployment matches its review, whether administrators can change balances or whether a reported issue remains unresolved. A general desire to find a "safe project" is too broad to test against a single rating.
Preserve the visible result before following links. Record the URL, timestamp, project identity, category values and methodology version where available. Then verify the particular asset and chain you intend to evaluate. A token page can be relevant background while still failing to describe the lending vault, bridge wrapper or custody arrangement through which exposure is obtained.
| Field | Example of an adequate entry | Reason to stop and investigate |
|---|---|---|
| Entity | Named deployment and chain, with contract address | The rating applies to a different token or product |
| Code evidence | Report revision mapped to deployed implementation | No verified mapping is available |
| Authority | Documented mint, pause and upgrade roles | A material privileged path remains unknown |
| Method | Named rating methodology and observation date | The number cannot be connected to a documented method |
| Unresolved condition | Specific missing evidence with an owner | The condition conflicts with the reviewer's acceptance rule |
Read the findings rather than only the presence of an audit badge. A report can contain an accepted issue, an unreviewed fix or a scope limitation. These are different states. The crypto due diligence checklist provides a broader evidence trail, while the fake audit report guide addresses authenticity before interpretation.
When ratings disagree, compare their objects and methods before asking which number is right. One service may assess public process documentation while another combines code, governance and market signals. Different results can follow from different questions. If the methodologies claim to measure the same thing, compare data freshness, missing fields and the underlying report links.
Do not average incompatible ratings. A numerical average of a process-quality grade and a broad project-health score creates a new metric with no defined meaning. If a team needs a decision matrix, keep the criteria separate and document the acceptance rule for each. A disqualifying authority or an unknown deployment should remain visible rather than being diluted by unrelated positive signals.
A compact decision memo can end with one of three evidence states: sufficient for the stated review question, contradicted by a material finding or incomplete pending a named document. These are editorial workflow states, not investment recommendations. They make it possible for another reviewer to understand why the process stopped or continued.
Revisit the memo when a triggering change occurs. A new audit can close a known evidence gap; an upgrade can reopen one. The saved score remains a historical observation. The current decision should be based on current scope and evidence, with the unresolved items still visible.
Original disclosure census: 28 indicator weights
Method
- Source and selection
- Read every indicator row in the public Skynet methodology weight table. Exclude the header. A weight is numerical only when its weight cell contains a digit.
- Retrieved
- Observation
- The complete indicator table was inspected; only its header was excluded.
Results
| Disclosure observation | Count | Denominator |
|---|---|---|
| Indicator rows | 28 | All rows in the public indicator table |
| Qualitative weight cells | 28 | 28 indicator rows |
| Numerical weight cells | 0 | 28 indicator rows |
| Category groups | 6 | Distinct category groups in that table |
Our recorded result is 0 numerical weight cells among 28 indicators. It identifies a reproducibility limit in the published table, not a defect in a particular project or the vendor's calculations.
The appropriate use of this result is narrow. A reviewer can check whether the methodology considers a relevant subject and can ask the provider how that subject affects the score. The reviewer cannot derive exact contributions by converting labels such as high or medium into invented percentages. No project score was recalculated in this census.
Limits
- One provider's public methodology table, retrieved on the recorded date.
- No rating algorithm, private data, exploit outcome sample or predictive performance was measured.
- Indicator counts and qualitative labels may change; the saved evidence fixes the observation used here.
The reproducible record is kept in the repository: seo/research/articles-42-51-2026-09-06/surveys.py and seo/research/articles-42-51-2026-09-06/survey-43.json. These file paths are not public downloads.
Frequently asked questions
Can a score be used in an automated procurement filter?
It can trigger manual review if the filter records the provider, methodology and entity. Keep an exception path for missing data and require direct evidence for any control that is a procurement condition.
What if a project has no rating at all?
Record the rating as unavailable. Evaluate the relevant reports, deployment configuration and operational evidence directly; absence of a rating does not establish either a defect or a clean result.
Can a rating change while the contract code stays identical?
Yes, a broad methodology may incorporate new documentation, governance, market or operational information. Check the category changes and methodology notes before attributing the movement to a code change.