RACF-CC & Governance • Evidence-led assurance

Why headline security scores can hide real progress

Security scores are useful. But the number at the top of the dashboard is not the same thing as security improvement.

The central problem

A simple number can conceal a complicated result

A percentage rises. A risk score falls. A gauge moves from red towards green. For security teams, managers and boards, that simplicity is useful: complex environments need a way to communicate direction.

Risk begins when the headline number becomes the outcome rather than an indicator.

100/100 → 100/100PingCastle Domain Risk Level

Rule-level Active Directory improvements occurred while serious residual findings kept the maximum headline risk unchanged.

40.7% → 43.5%Microsoft Zero Trust Assessment

Substantial identity improvements sat beneath a modest 2.8 percentage-point movement and a changed assessment model.

Neither assessment tool had failed

A single aggregate number simply could not communicate everything that improved, everything that remained, and whether the two assessments were fully comparable.

Example 1 • Active Directory

The 100/100 score that did not move

Two PingCastle Active Directory health assessments were compared roughly three months apart. At headline level, the result looked disappointing:

Earlier assessment100/100
Follow-up100/100

PingCastle's Domain Risk Level uses the highest of its main indicator scores. Stale Objects, Privileged Accounts and Anomalies all remained at 100/100, so the overall result correctly continued to signal significant residual Active Directory risk.

But the rule-level evidence told a different—and equally important—part of the story.

Control-level findingEarlierFollow-upInterpretation
Golden Ticket / KRBTGT500Historic Kerberos persistence exposure reduced after controlled KRBTGT resets.
Certificate takeover250Risky certificate-template exposure reduced.
Object configuration3116Fewer object-configuration findings remained.
Account takeover6661Some improvement, with material residual exposure.
Reconnaissance72Fewer reconnaissance-related findings remained.

The KRBTGT account had previously retained a password dating from August 2016. The follow-up evidence showed two controlled resets, and the associated Golden Ticket rule no longer matched. Certificate-related exposure also reduced substantially.

Important attack paths have been reduced, but substantial residual risk remains.

That conclusion is more useful than either “nothing improved” or “the remaining 100 can be ignored”.

Example 2 • Zero Trust

A 2.8-point improvement that also hid the detail

The overall Microsoft Zero Trust Assessment increased from 40.7% to 43.5%. Viewed alone, that could suggest modest progress.

Single-factor users907 → 456

Approximately a 50% reduction.

Outstanding high-risk sign-ins8 → 0

No outstanding high-risk sign-ins at follow-up.

Inactive guest accounts471 disabled

Accounts had no recorded successful sign-in or more than 13 months of inactivity.

The denominator had effectively changed

The follow-up used a newer assessment model with additional controls and greater emphasis on cloud-native networking, data governance and capabilities dependent on particular licensing or architectural changes. The percentages were useful for direction-setting, but not a clean measure of every change delivered during the evaluation.

Control state mattered too. Several Conditional Access policies remained in report-only mode so impact could be evaluated before enforcement. That was evidence of progress towards a control—not evidence that a preventive control was already enforced.

Keep the tool—improve the interpretation

The dashboard is not the problem

Microsoft Secure Score, Zero Trust assessments, vulnerability scanners and Active Directory assessment platforms can expose weaknesses, create repeatable checks, identify priorities and make technical information easier to communicate. The research behind RACF-CC relied extensively on them.

The problem begins when an organisation allows the score to replace the analysis.

The model changes

New checks, weightings or denominators can move the score even when the underlying environment has not changed in the same way.

The scope changes

Discovering more assets or gaining authenticated visibility can increase findings while producing a more trustworthy assessment.

The maximum dominates

A high-weighted unresolved issue can keep an aggregate risk score unchanged while targeted weaknesses are remediated.

Easy points distract

Completing numerous low-value recommendations can improve a percentage without reducing the most credible attack path.

Movement in a score and movement in risk are related, but they are not necessarily the same thing.

A stronger reporting pattern

Use the score as one layer of evidence

  1. 01

    Start with the headline indicator

    Use it to understand broad posture and identify areas requiring investigation.

  2. 02

    Examine the controls underneath it

    Determine which individual findings improved, worsened or remained unresolved.

  3. 03

    Verify the implementation

    Use configuration, policy state, logs, vulnerability results or other direct evidence to establish what changed.

  4. 04

    Relate the change to risk

    Identify which credible attack path, exposure or operational consequence the control affects.

  5. 05

    Record what remains

    Keep residual risk, exceptions and controls still in pilot, audit or report-only states visible.

Instead of

“Our security score increased by five points.”

Report

“MFA exposure has reduced, this legacy authentication path has been removed, this network boundary is now enforced, these critical vulnerabilities have been remediated, and these material risks remain outstanding.”

Example 3 • Vulnerability management

Before comparing two numbers, confirm they measure the same thing

A scanner may report more critical findings because security deteriorated. It may also have authenticated successfully to more hosts, brought additional subnets into scope, reached previously unavailable assets or changed classification rules.

Raw Nessus totals in the RACF-CC research were treated cautiously because host availability and credentialed visibility changed between assessments. The more defensible comparison used 84 server-network hosts visible in both datasets.

Comparable population84 hosts

425 earlier critical instances

371 follow-up critical instances

54 fewer instances • 12.7% reduction

Comparability is part of the evidence

Record the population, scope, authentication success, assessment version and important exclusions before presenting two scan totals as a trend.

Measurement over assumption

Evidence-led does not mean metric-led

Being evidence-led means supporting a security claim with appropriate evidence—not collecting the largest possible number of metrics or allowing dashboards to dictate risk decisions.

If MFA improved

Authentication-method and policy evidence should demonstrate coverage and enforcement.

If segmentation improved

Reachability testing or firewall logs should demonstrate the boundary.

If an endpoint control is enforced

Configuration and telemetry should distinguish block mode from audit mode.

If backups provide resilience

Recovery evidence should show that critical data and services can actually be restored.

The measure should follow the security question—not the other way around.

Resource-aware decisions

Why this matters particularly for small teams

Resource-constrained organisations cannot remediate everything simultaneously. A team may spend significant time resolving one high-impact identity weakness while dozens of lower-priority dashboard recommendations remain visible and the overall score barely moves.

That does not necessarily make the work a poor use of time. The relevant questions are whether the action materially reduced a credible risk, whether it was implemented safely, and whether the improvement can be demonstrated.

This is why RACF-CC separates expected security benefit from implementation effort, operational impact and cost. The purpose is not to optimise a dashboard score. It is to decide what deserves attention.

The practical conclusion

Scores should start conversations, not finish them

Aggregate assessments remain valuable for discovering gaps and understanding direction. But the strongest evaluation evidence usually sits underneath the headline: before-and-after control metrics, configuration evidence, comparable vulnerability data, rule-level changes and operational validation.

A PingCastle Domain Risk Level of 100/100 still communicated that important Active Directory risks remained. A Zero Trust score in the low forties still showed that substantial work was unfinished. The mistake would be interpreting either number as a complete description of what had—or had not—changed.

What changed? Which attack path was reduced? Is the control actually enforced? What evidence proves it? What residual risk remains?

Headline security scores help us navigate. They should not become the destination.

Voluntary support

Found this useful? Support Trends4You

Trends4You's practical guides, RACF-CC resources and downloadable tools are provided free of charge. If they have helped you or your organisation, you can support the time and hosting that keeps them freely available.

Support is optional, handled securely by Stripe and does not provide additional access.