Quick Read

SPK DDMS2000:2026 permits numeric scoring in due diligence but requires organisations to document their methodology sufficiently to reconstruct underlying findings and to avoid two critical pitfalls: treating scores as more precise than the judgements they represent, and allowing serious individual findings to be obscured through mathematical averaging of unrelated low-risk factors. The standard explicitly prohibits using scoring as a substitute for analyst and reviewer judgement, emphasising that numeric conversion must remain transparent about what it reveals and what it conceals.

Why This Whitepaper Exists

Numeric scoring is close to universal in due diligence practice. A risk assessment, a questionnaire, a background check — each produces findings, and those findings are commonly converted into a number, because numbers are easy to compare, easy to aggregate, and easy to put in a dashboard. SPK DDMS2000:2026 does not discourage this. Section 6.4.6 states plainly that scoring is acceptable and often useful. What the standard insists on is that organisations remain honest about what a score actually represents, and what it quietly hides.

A numeric score can create an impression of precision and objectivity the underlying judgement calls do not always support.

What the Standard Actually Says

Section 6.4.6 requires that any scoring methodology used to convert risk assessments, questionnaires, background checks, or other due diligence activity into a numeric or coded score be documented sufficiently that the specific findings behind a given score can be reconstructed and reviewed, consistent with the evidence standard at Section 10.4.8. It further requires that scoring never be used in a way that obscures the underlying qualitative judgement behind a score, creates false precision, or allows a material finding to be diluted through averaging against unrelated low-risk factors. Critically, the clause states that scoring shall not be treated as a substitute for the analyst's and reviewer's own judgement under Section 10.22.

Annex A.2 — the standard's own illustrative worked model for risk factor scoring — carries the same caution immediately before presenting its example table: organisations should feel free to use a model of that kind, but should remain alert to its limits.

The Two Failure Modes Named in the Standard

False precision

A score of “11” looks objective in a way that “the analyst was moderately concerned about three factors” does not. But the underlying inputs to most scoring models are themselves judgement calls — was this jurisdiction risk factor a 2 or a 3, was this ownership structure genuinely opaque or merely unfamiliar. A score inherits the imprecision of its inputs; it does not remove it. Treating a score as more exact than the judgement behind it is the first failure mode the standard names directly.

Dilution through averaging

This is the more dangerous failure mode, because it can hide a genuinely serious finding rather than merely overstate confidence in a minor one. If a scoring model averages, or otherwise aggregates, several unrelated factors into a single number, one seriously concerning factor can be mathematically offset by several low-risk factors — producing a moderate overall score that never surfaces the one finding that actually mattered.

This is precisely why Annex A.2's own illustrative model includes an override provision: any single factor scoring at the highest level for a confirmed sanctions or PEP hit, or a comprehensive-sanctions jurisdiction, is defined to trigger at minimum Tier 3 regardless of the aggregate score. The override exists because the standard's authors anticipated the dilution problem directly — a scoring model without an override provision will, sooner or later, average away the one finding a reviewer most needed to see.

Reading a Score Correctly

The table below illustrates, using the standard's own illustrative tier bands at Annex A.2, what a score tells a reviewer and what it deliberately does not. This is the discipline Section 6.4.6 is asking organisations to build into their own review practice: never accept the number without asking what it is standing in for.

Aggregate score

DD Tier

What the score cannot tell you

0–4

Tier 1 — Screening only

Whether a single factor scored 0 because it is genuinely low risk, or because the underlying research was thin

5–9

Tier 2 — Standard DD

Whether the score reflects five factors each moderately concerning, or four low-risk factors offsetting one serious one

10–14

Tier 3 — Enhanced DD

Which specific factor is driving the score, and whether that factor alone would justify escalation on its own

15+

Tier 4 — Field-verified DD

Whether the score was reached before or after an override provision should have applied regardless of the total

Building a Defensible Scoring Model

Keep the record behind the number

A scoring model satisfies Section 6.4.6 only if a reviewer, or a certification assessor, can take any given score and reconstruct exactly which factors produced it and why. A dashboard that displays scores without a click-through to the underlying factor ratings does not meet this bar, however clean it looks.

Build override provisions for anything non-negotiable

Section 6.4.4 requires the organisation to document override provisions for factors it treats as non-negotiable regardless of aggregate score, and the illustrative model at Annex A.2 shows one way to do this. Any organisation adapting the model should ask, module by module, which findings should never be averaged away — and encode that as an override, not as a heavily weighted factor that can still be outvoted by enough low-risk inputs.

Separate scoring from deciding

Section 10.22 requires that DD decisions — Go, Conditional Go, or No-Go — be made by a person with defined authority, applying judgement to the findings in front of them. A score is an input to that judgement, not a replacement for it. A process in which a score above a threshold automatically produces a decision, with no human review of the findings behind the number, does not satisfy Section 6.4.6's requirement that scoring not substitute for the analyst's and reviewer's own judgement.

Recalibrate against real outcomes

Section 12.1.7 requires the organisation to track and analyse outcomes of decisions made on the basis of due diligence, as feedback into whether the risk tiering methodology is well calibrated. A scoring model that has never been tested against what actually happened to the subjects it scored is a model running on assumption, not evidence.

Common Gaps Worth Checking

  • A dashboard displays aggregate scores with no way to see the individual factor ratings behind a given number.

  • The scoring model has no override provisions, so a confirmed serious finding can mathematically disappear into an averaged total.

  • Decisions are made automatically once a score crosses a threshold, with no documented human review of the underlying findings.

  • The scoring model was built once, at implementation, and has never been recalibrated against real outcomes tracked under Section 12.1.7.

  • Country or jurisdiction risk is scored using an unmodified public index, rather than the organisation's own adapted methodology required at Section 6.4.7.

How Speeki Sentinel Certification Assesses This

Certification against SPK DDMS2000:2026 tests whether a sample of scored cases can be reconstructed back to their underlying findings, whether override provisions exist and were correctly applied where triggered, and whether decisions in the sample show evidence of genuine review rather than automatic score-based approval. A scoring model that looks sophisticated on a dashboard but cannot be reconstructed case by case is treated as a gap against Section 6.4.6, not a matter of presentation style.

Speeki Sentinel is the certification product through which this assessment is delivered. Organisations may design and operate their own scoring methodology independently of Sentinel, consistent with the standard's own emphasis that risk methodology is the organisation's to define; certification is a separate, optional step available once an organisation believes its approach is ready to be independently tested.

Speeki is an accredited certification body. For current information on the specific accreditations Speeki holds and their scope, please refer to speeki.com rather than relying on this whitepaper, as accreditation status and scope are maintained centrally and can change.

Closing Note

A score is a summary, not a substitute for the facts it summarises. SPK DDMS2000:2026 permits scoring precisely because it is useful — for comparison, for prioritisation, for dashboards that let a DD Function Owner see across a large caseload at a glance. What the standard will not accept is a score that has become the whole story: a number nobody can unpack, protecting a serious finding nobody can see.