An OutThink open standard for measuring human risk. The library defines 90 metrics across 10domains - each tagged with what it measures, how it’s calculated, the HRM maturity level (L1 Reactive, L2 Adaptive, L3 Proactive, L4 Predictive) where it becomes meaningful, its direction of good, and why it matters. 45 of them are outcome-grade metrics (L3-L4).
01 - Engagement & Participation (12 metrics)
Whether people are actually paying attention - not just completing. Engagement metrics separate genuine attention from click-through compliance, and they are the earliest signal of whether anything downstream can work.
ENG-01 - Training completion rate
Percentage of assigned users who finished the training.
L1 Reactive · Calculated as: completed / assigned · Why it matters: The classic compliance number - and the classic trap. Completion measures activity, not outcome: report it for coverage and audit evidence, never as proof of behavior change.
ENG-02 - Assigned / enrolled users
The population targeted by a campaign.
L1 Reactive · Calculated as: count · Why it matters: The denominator behind every other metric. Check it first: impressive percentages on a partial population are how programs fool themselves.
ENG-03 - Not-started rate
Percentage of assigned users who never opened the training.
L1 Reactive · Calculated as: not_started / assigned · Why it matters: A distinct disengaged cohort - different people, different problem, different play than those who click through to complete. Find them early.
ENG-04 - In-progress rate
Percentage of users currently mid-training.
L1 Reactive · Calculated as: in_progress / assigned · Why it matters: Pacing context for campaign operations - useful for forecasting completion, not for judging risk.
ENG-05 - Overdue / past-due rate
Percentage of users past the assignment due date.
L1 Reactive · Calculated as: overdue / assigned · Why it matters: Compliance exposure and follow-up workload. Persistent overdue pockets usually trace to manager engagement, not individual laziness.
ENG-06 - Median time-to-complete
Days from assignment to completion.
L1 Reactive · Calculated as: median(days) · Why it matters: A procrastination and friction proxy. When it stretches, look at content relevance and workload before blaming the audience.
ENG-07 - Average time-in-content
Minutes spent per module per user.
L2 Adaptive · Calculated as: mean(minutes) · Why it matters: A depth proxy: suspiciously low time-in-content is the signature of click-through completion. Read it alongside the genuine engagement rate.
ENG-08 - Genuine engagement rate
Share of users who actively engage, versus click-through-to-complete, versus never start.
L2 Adaptive · Calculated as: engaged / assigned · Why it matters: The metric that answers 'who is actually paying attention?' - it separates three populations a completion rate merges into one. The flagship engagement KPI for any program that has outgrown completion reporting.
ENG-09 - Interaction depth score
Free-text given, questions asked, optional content explored.
L2 Adaptive · Calculated as: weighted index · Why it matters: True attention, quantified. Depth interactions also feed your workforce-listening intelligence - people who write back are telling you where risk lives.
ENG-10 - Nudge response rate
Nudges acted on as a share of nudges delivered.
L2 Adaptive · Calculated as: acted / delivered · Why it matters: Whether coaching lands at the moment of risk. A falling response rate is fatigue or poor timing - recalibrate cadence before adding volume.
ENG-11 - Discretionary / champion engagement
Opt-in content, peer help and self-initiated security activity.
L2 Adaptive · Calculated as: count / rate · Why it matters: The culture signal: people choosing security when nobody is making them. Your future champions surface here first.
ENG-12 - Content relevance rating
Self-reported 'this was made for my role / industry'.
L2 Adaptive · Calculated as: mean rating · Why it matters: Relevance drives attention. Low relevance scores predict disengagement long before completion rates move - and they tell you exactly which content to fix.
02 - Phishing & Attack-Sim Resilience (13 metrics)
How people perform against simulated attacks - from generic phishing to deepfake, vishing and AI-grade lures. The discipline here is reading resilience, not celebrating a low click rate on easy tests.
PHI-01 - Phishing click rate
Clicks as a share of simulations delivered.
L1 Reactive · Calculated as: clicked / delivered · Why it matters: The most quoted number in security awareness - and it says almost nothing on its own. A low click rate on generic tests is not resilience to targeted, AI-grade attacks. Benchmark it, caveat it, and pair it with difficulty-adjusted and reporting metrics.
PHI-02 - Emails sent / delivered
Simulation volume reaching inboxes.
L1 Reactive · Calculated as: count · Why it matters: The denominator for every simulation rate. Delivery problems masquerade as resilience wins.
PHI-03 - Open rate
Opens as a share of delivered.
L1 Reactive · Calculated as: opened / delivered · Why it matters: Exposure context - how many people the lure actually reached and tempted. Not a failure metric.
PHI-04 - Reporting rate
Reports as a share of delivered.
L1 Reactive · Calculated as: reported / delivered · Why it matters: An active resilience behavior, not just the absence of failure. Programs that grow reporting build a human detection layer - this is the number to grow, not just the click rate to shrink.
PHI-05 - Credential submission rate
Credentials entered as a share of delivered.
L1 Reactive · Calculated as: submitted / delivered · Why it matters: The closest simulation proxy to real compromise. Weight it far above clicks in any risk read-out.
PHI-06 - Attachment open rate
Attachments opened as a share of delivered.
L1 Reactive · Calculated as: opened_att / delivered · Why it matters: Payload-exposure behavior - the malware-relevant counterpart to credential submission.
PHI-07 - Repeat-clicker rate
Percentage clicking across two or more simulations.
L1 Reactive · Calculated as: repeat / clickers · Why it matters: Handle with care: repeat clicking is not the same as being high-risk. Always read it alongside targeting, access and the composite risk score before anyone lands on a 'list'.
PHI-08 - Time-to-click
Median seconds from delivery to click.
L2 Adaptive · Calculated as: median(sec) · Why it matters: An impulsivity and attention-fragmentation proxy - fast clicks point to workload and context, which are fixable, unlike 'carelessness'.
PHI-09 - Time-to-report
Median seconds from delivery to first report.
L2 Adaptive · Calculated as: median(sec) · Why it matters: Collective response speed - how fast your organization flags a live-style attack. One of the best operational resilience numbers a CISO can put in front of a board.
PHI-10 - Report:click ratio (net resilience)
Reporting strength relative to clicking.
L2 Adaptive · Calculated as: reported / clicked · Why it matters: Directional resilience in a single number: are defenders outpacing victims? Far more informative than either rate alone.
PHI-11 - Difficulty-adjusted click rate
Click rate weighted by simulation sophistication.
L3 Proactive · Calculated as: Σ(click × difficulty) · Why it matters: The honest click rate. A flat 8% on ever-harder simulations is improvement; a flat 8% on the same easy test is a plateau dressed as success. This is the comparable, defensible benchmark.
PHI-12 - AI / targeted-attack resilience
Click and report performance on deepfake, spear-phishing, vishing and smishing simulations.
L3 Proactive · Calculated as: per-type rate · Why it matters: Are your people ready for what attackers actually send now? Board-grade readiness for the AI-attack era - and the gap between generic and targeted performance is usually the most persuasive statistic in the program.
PHI-13 - Redemption rate (report-after-click)
Simulations reported after the user clicked, as a share of clicks.
L1 Reactive · Calculated as: reported after clicked / clicked · Why it matters: The recovery behavior: people who realize, own it and report anyway. A rising redemption rate signals psychological safety and growing awareness - celebrate it, never punish it.
03 - Knowledge & Competence (8 metrics)
What people actually know and can apply, measured continuously rather than at annual-training time. Competence scores make security skill visible, comparable and reportable - per person, per behavior domain.
CMP-01 - CyberQ - security competence score
Holistic per-person security-competence score.
L2 Adaptive · Calculated as: OutThink composite model · Why it matters: A defensible, ownable competence measure - think of it as a credit score for security behavior. Unlike a click rate, it can sit in a development conversation and an executive pack with equal credibility.
CMP-02 - Competence distribution
Share of workforce by competence band.
L2 Adaptive · Calculated as: banding · Why it matters: The shape of workforce competence - where mass sits, where the tail is, and how concentrated your exposure is.
CMP-03 - Competence trend / improvement
Change in competence score over time.
L2 Adaptive · Calculated as: period delta · Why it matters: Makes growth trackable: the difference between 'we trained people' and 'people measurably know more than last quarter'.
CMP-04 - Competence by behavior domain
Competence split by behavior area (phishing, data handling, credentials, AI use, OT, physical).
L3 Proactive · Calculated as: per-domain score · Why it matters: Pinpoints exactly where competence is weak, so the next quarter's program targets the gap instead of re-teaching the strength.
CMP-05 - Baseline knowledge score
Assessment score at baseline and ongoing.
L2 Adaptive · Calculated as: % correct · Why it matters: Your starting point and growth trajectory. Programs that never baseline can never prove improvement - run it before anything else.
CMP-06 - Knowledge decay rate
Score drop between touchpoints.
L2 Adaptive · Calculated as: slope over time · Why it matters: Quantifies the reinforcement need: annual training is forgotten in weeks, and this metric shows exactly how fast. The data case for always-on over annual.
CMP-07 - Competence calibration
Confidence versus actual competence (over- and under-confidence).
L3 Proactive · Calculated as: confidence − score gap · Why it matters: Surfaces the dangerous cohort: confidently wrong. Over-confident populations take risks that under-confident ones never would - and they don't self-select into help.
CMP-08 - % above competence threshold
Workforce share meeting a defined competence bar.
L2 Adaptive · Calculated as: n / total · Why it matters: A board-reportable assurance statement: 'X% of our workforce demonstrably meets the bar.' Clean, comparable, auditable.
04 - Attitudes, Motivation & Culture (6 metrics)
Why people behave the way they do. Attitudes predict behavior better than knowledge does - these metrics surface motivation, learned helplessness, champions and the culture trend a click rate will never reveal.
ATT-01 - Security attitude / intention-to-comply
Attitude and intent captured at baseline and ongoing.
L2 Adaptive · Calculated as: interaction composite · Why it matters: Attitudes drive behavior more reliably than knowledge does. This is the leading indicator most programs never measure - and the first place culture change shows up.
ATT-02 - Motivational driver mix
Workforce share by driver: personal relevance, professional duty, learned helplessness, champion energy.
L2 Adaptive · Calculated as: segmentation · Why it matters: You cannot motivate a workforce with one message. Knowing the driver mix lets you give every person a reason to care in their own language.
ATT-03 - Learned-helplessness prevalence
Percentage who are low-confidence or disengaged by belief ('I'll be fooled anyway').
L2 Adaptive · Calculated as: n / total · Why it matters: Hidden disengagement a click rate never reveals. This cohort needs confidence rebuilt, not more content - more training makes it worse.
ATT-04 - Champion identification rate
Percentage identified and activated as security champions.
L2 Adaptive · Calculated as: n / total · Why it matters: Culture multipliers, found early. A measured champion pipeline is the cheapest scaling mechanism a program has.
ATT-05 - Risk perception accuracy
Gap between self-perceived and actual risk.
L3 Proactive · Calculated as: perceived − actual · Why it matters: Blind-spot detection: people who think they're safe and aren't. Closing the perception gap changes behavior without a single new rule.
ATT-06 - Security culture index
Composite of attitudes across the organization.
L3 Proactive · Calculated as: index · Why it matters: The trackable culture trend - one executive-grade line that summarizes whether the human layer is strengthening or fraying.
05 - Secure Behavior Change (13 metrics)
The point of the whole program: are real-world behaviors actually changing? Measured through security-stack integrations - data handling, credentials, browsing, AI use, payments, clinical and OT workflows - not through training records.
BEH-01 - Real-world secure-behavior rate
Compliant actions per behavior, observed through security-stack integrations (data, credentials, browsing, media, AI, OT, physical).
L3 Proactive · Calculated as: compliant / total events · Why it matters: Outcome, not activity - the actual goal of human risk management, measured where work happens instead of where training happens. The metric the whole category is named after.
BEH-02 - Behavior-change count
Users who stopped a risky behavior in a period.
L3 Proactive · Calculated as: period delta (count) · Why it matters: Concrete, attributable wins: '42 people in Finance stopped using unapproved file sharing this quarter' is the sentence that wins the budget conversation.
BEH-03 - Shadow-IT / unapproved file-sharing incidence
Risky data-egress events.
L3 Proactive · Calculated as: events / user · Why it matters: Real data-loss exposure - and a window into broken processes: shadow IT is usually a symptom of friction, not malice.
BEH-04 - Credential hygiene
MFA and phishing-resistant MFA coverage; password reuse.
L3 Proactive · Calculated as: coverage % · Why it matters: Account-takeover exposure in one number. Phishing-resistant coverage is the strongest single control your population can adopt - measure it from the identity provider, not from self-report.
BEH-05 - DLP violations per segment
Data-handling policy violations by group.
L3 Proactive · Calculated as: rate / segment · Why it matters: Real data-handling risk, located. Segment-level views turn a compliance number into a targeting map for the next intervention.
BEH-06 - Behavioral recidivism
Repeat risky-behavior rate after an intervention.
L3 Proactive · Calculated as: repeats / flagged · Why it matters: Does correction actually stick? Falling recidivism is the cleanest evidence that coaching works; rising recidivism says the intervention is noise.
BEH-07 - Time-from-error-to-correction
Latency from risky event to corrective nudge.
L2 Adaptive · Calculated as: median(time) · Why it matters: Coaching works when it arrives at the moment of risk, not in next month's newsletter. This is the speed of your learning loop.
BEH-08 - Post-nudge behavior-change rate
Percentage who correct after a nudge.
L3 Proactive · Calculated as: corrected / nudged · Why it matters: Proof that correction closes the loop. With recidivism, this pair tells you whether your intervention engine actually changes anything.
BEH-09 - Cyber Range performance by behavior
Decision quality across realistic scenarios (deepfake, vishing, smishing, data, browsing, endpoint, social, physical).
L2 Adaptive · Calculated as: scenario score · Why it matters: Practiced resilience beyond phishing - measured where it can be safely tested. The scenario scores show which decisions hold up under pressure and which collapse.
BEH-10 - Behavior coverage breadth
Number of behaviors trained and practiced versus phishing-only.
L2 Adaptive · Calculated as: coverage % · Why it matters: Exposes the phishing-only gap: if you only practice one behavior, the other eighty percent of human risk is untrained. The maturity metric hiding in plain sight.
BEH-11 - Payment-verification adherence rate
Payment and bank-detail-change instructions where out-of-band callback verification was completed before execution.
L3 Proactive · Calculated as: verified / instructions · Why it matters: The anti-deepfake control, measured at the workflow gate where money actually moves. In a voice-cloning era this is the single most consequential behavior a finance function performs - and it is fully measurable from payment-workflow telemetry.
BEH-12 - Patient-record access integrity
Inappropriate or curiosity-driven record accesses per 1,000 accesses, from EHR audit logs.
L3 Proactive · Calculated as: flagged / 1,000 accesses · Why it matters: The healthcare-specific behavior regulators actually enforce: snooping. Measured from the EHR audit trail, not from policy attestations - and one of the few security metrics a clinical board instantly understands.
BEH-13 - OT procedure adherence & bypass attempts
Removable-media, boundary and vendor-access procedure adherence on OT networks, and attempts to bypass safety or security interlocks.
L3 Proactive · Calculated as: violations + bypass attempts / period · Why it matters: Where human risk meets physical safety. OT monitoring sees what awareness surveys never will: the workarounds. Any non-zero bypass-attempt trend is a leadership conversation, not a training module.
06 - Human Risk Quantification (10 metrics)
Turning behavior, attitudes, access, targeting and device posture into a defensible human risk score - per person, per segment, over time. The numbers a board and a SOC can both act on.
RSK-01 - Composite human risk score
Per-person risk from behavior, attitudes, access, real-world targeting, device posture and workplace factors.
L3 Proactive · Calculated as: OutThink risk model · Why it matters: The operational, defensible risk number a SOC and a board can both use - because it is built from real signals, not training records. Everything else in this domain explains or validates it.
RSK-02 - Risk segment distribution
Share of workforce at low, medium and high risk.
L3 Proactive · Calculated as: banding · Why it matters: The shape of organizational risk - and the chart that turns an abstract score into a population leadership can reason about.
RSK-03 - High-risk segment size & movement
Who is most at risk now, how many, and the trend.
L3 Proactive · Calculated as: top-N + delta · Why it matters: Answers the question legacy tools cannot: 'which fifty people are most at risk today, and why?' Watch movement more than size - a shrinking high-risk segment is the program working.
RSK-04 - Risk-score component contribution
Explainable breakdown of what drives each score.
L3 Proactive · Calculated as: attribution · Why it matters: Why someone is risky - the explainability that makes a score trusted and actionable instead of resented and ignored.
RSK-05 - Targeting intensity
How heavily a person is attacked in the real world.
L3 Proactive · Calculated as: attacks / user · Why it matters: The external-pressure dimension of risk: your most-attacked people are chosen by adversaries, not by you. Resilience only makes sense relative to targeting.
RSK-06 - Access / damage potential
Privilege and data access - the blast radius if this person is compromised.
L3 Proactive · Calculated as: weighted access · Why it matters: How much a compromise would cost. A medium-risk person with crown-jewel access outranks a high-risk person with none - this is the metric that encodes that.
RSK-07 - Device security posture
Endpoint-compliance contribution to human risk.
L3 Proactive · Calculated as: posture score · Why it matters: The technical exposure of the individual - the device they sit behind is part of their risk portrait, not a separate dashboard.
RSK-08 - Workplace risk factors
Email fatigue, frequent travel, collaboration-network exposure.
L4 Predictive · Calculated as: composite · Why it matters: The situational drivers legacy tools never capture: when and where people are set up to fail. Context that turns risk scores from judgments into diagnoses.
RSK-09 - Human risk reduction over time
Change in aggregate human risk score.
L3 Proactive · Calculated as: trend · Why it matters: The headline question, answered: is human risk actually decreasing? If you report one number upward, report this one.
RSK-10 - Risk-adjusted resilience
Resilience weighted by targeting and access.
L4 Predictive · Calculated as: composite · Why it matters: True exposure-adjusted readiness: how strong you are where it matters most, not on average.
07 - Operational Integration & Business Impact (10 metrics)
Where human risk management meets the rest of the security operation: incident correlation, automated response, SOC efficiency and financial value. The metrics that prove the program is infrastructure, not theater.
OPS-01 - Human-factor incident rate / reduction
Incidents driven by human error, over time.
L3 Proactive · Calculated as: rate / trend · Why it matters: The ultimate outcome the entire program exists to move. Every other metric in this library is upstream of this one.
OPS-02 - Incidents by risk segment (validation)
Incident rate across low / medium / high risk segments.
L3 Proactive · Calculated as: correlation · Why it matters: The credibility test: if high-risk segments have materially higher incident rates, your risk score predicts reality - and earns the right to drive decisions.
OPS-03 - Line-manager engagement × incidents
Manager engagement against business-unit incident rates.
L3 Proactive · Calculated as: correlation · Why it matters: Managers are a hidden risk variable: units whose managers disengage show more incidents. Engage the manager first - it moves the whole unit.
OPS-04 - Improvement actions generated / actioned
Prioritized recommendations across people, process and technology - and whether they get actioned.
L3 Proactive · Calculated as: count + action rate · Why it matters: The platform finding the work and the organization doing it. Generated-but-not-actioned is a governance gap, not a tooling gap.
OPS-05 - Self-remediation rate
High-risk users who self-remediate within the grace period.
L4 Predictive · Calculated as: remediated / flagged · Why it matters: Scaled correction without the SOC chasing anyone: people fixing their own risk when given visibility and a window. The most dignified automation in security.
OPS-06 - Mean time to remediate
Time for a risk score to recover after a flag.
L4 Predictive · Calculated as: median(time) · Why it matters: Organizational responsiveness to human risk, in the same shape as every other MTTR your operation already reports.
OPS-07 - Risk-driven access changes
Conditional-access updates driven automatically by human risk.
L4 Predictive · Calculated as: count · Why it matters: Security controls and human risk finally unified: risk signals adjusting access at machine speed. The definitive Level-4 metric.
OPS-08 - SOC hours saved / tickets reduced
Operational efficiency from automation.
L3 Proactive · Calculated as: hours / tickets · Why it matters: Hard ROI in the SOC's own currency: fewer tickets, faster triage, less manual chasing.
OPS-09 - Estimated risk / loss avoided ($)
Modeled financial value of risk reduction.
L4 Predictive · Calculated as: model · Why it matters: The CFO framing of human risk management. Use it as a modeled estimate with stated assumptions - it opens budget conversations that behavior charts cannot.
OPS-10 - Risk intelligence consumed by SOC / GRC
Share of access and triage decisions using human-risk context.
L4 Predictive · Calculated as: usage rate · Why it matters: Proof the score is operational, not a dashboard nobody acts on. Adoption by the SOC is the strongest trust signal a human-risk program can earn.
08 - Workforce Intelligence (Listening) (5 metrics)
Your workforce is a sensor network. Listening metrics quantify the intelligence people surface - risk gaps, policy friction, novel lures - that no scanner or phishing test would ever find.
INT-01 - Workforce insights surfaced
Structured insights from free-text, questions and nudge reactions.
L2 Adaptive · Calculated as: count / themes · Why it matters: Training turned into intelligence-gathering: thousands of human sensors reporting from where the risk actually is.
INT-02 - Risk-gap reports
Gaps users reveal - missing password manager, shadow tools, clunky processes.
L3 Proactive · Calculated as: count / themes · Why it matters: Finds what no phishing test or scanner would surface: the gaps people live with daily and nobody asked about.
INT-03 - Policy-friction signals
Where the secure path is clunky enough that people work around it.
L3 Proactive · Calculated as: themes · Why it matters: Pinpoints process fixes: when the workaround is easier than the rule, the rule is the vulnerability. Fix friction and the 'people problem' often evaporates.
INT-04 - Behavioral data-point volume
Total behavioral and interaction data points captured.
L2 Adaptive · Calculated as: count · Why it matters: The scale and credibility of the intelligence base under every score you report.
INT-05 - Emerging-threat early warning
Spikes in user-reported novel lures.
L3 Proactive · Calculated as: signal · Why it matters: A collective threat radar: your workforce sees new attack patterns before your tooling categorizes them. Reporting culture is what powers it.
09 - AI-Agent Risk Governance (4 metrics)
AI agents are a new workforce, and they need the same behavioral governance as people. These metrics extend human risk management to the agents acting on your organization's behalf.
AIA-01 - AI-agent risk score
Behavioral risk score for AI agents.
L4 Predictive · Calculated as: OutThink agent model · Why it matters: A new class of insider, governed like the workforce it joined. If agents act on your organization's behalf, they need a risk score on the same dashboard as their human principals.
AIA-02 - Agent behavioral drift detection
Rate of out-of-bounds agent behavior.
L4 Predictive · Calculated as: count / rate · Why it matters: Catches agents drifting from their job - in the context of their human principal, not just as anomalous traffic at 3am.
AIA-03 - Agents under governance (coverage)
Agents observed versus total in the environment.
L4 Predictive · Calculated as: coverage % · Why it matters: Governance reach as agents proliferate. The ungoverned remainder is your shadow-agent estate - and it grows by default.
AIA-04 - Human-principal-linked agent risk
Agent risk contextualized by the human principal's risk portrait.
L4 Predictive · Calculated as: composite · Why it matters: The contextual read only a human-risk platform can produce: a risky agent owned by a risky human is a different problem than either alone.
10 - Program Health & Adoption (9 metrics)
Whether the program itself is set up to succeed: coverage, cadence, integrations, baseline and feature adoption. Weak program-health numbers explain weak outcomes everywhere else.
PGM-01 - Continuous program coverage
Workforce share in an always-on program versus one-off campaigns.
L2 Adaptive · Calculated as: covered / total · Why it matters: Reactive versus continuous operating model in one number. October-only programs produce October-only behavior.
PGM-02 - Integration coverage
Security systems connected (EDR, DLP, email, IAM, web).
L3 Proactive · Calculated as: count connected · Why it matters: The gate between maturity levels: without integrations there is no real-world behavior data, and without behavior data a program is stuck reporting training activity.
PGM-03 - Baseline assessment run
Whether a baseline has been run, and its coverage.
L2 Adaptive · Calculated as: yes/no + % · Why it matters: Personalization is blind without a baseline. It is also the only way to ever prove improvement - run it first, not eventually.
PGM-04 - Competence-score coverage
Share of users with a competence (CyberQ) score.
L2 Adaptive · Calculated as: covered / total · Why it matters: A defensible score only helps for the population that has one. Coverage gaps are silent blind spots in every report downstream.
PGM-05 - Scenario deployment breadth
Practice scenarios live beyond phishing.
L2 Adaptive · Calculated as: count · Why it matters: Adoption health for the other eighty percent of behaviors: how much of the risk surface your program actually exercises.
PGM-06 - Autonomy level enabled
Advisory → selective → full automation.
L4 Predictive · Calculated as: tier · Why it matters: Trust in the platform, made visible: how much of the correct-and-respond loop runs without a human in the middle.
PGM-07 - Admin engagement / configuration recency
Admin logins, last configuration change, active admins.
L1 Reactive · Calculated as: activity · Why it matters: Program stewardship. A platform nobody tunes is a program nobody owns - and it shows up here months before it shows up in outcomes.
PGM-08 - Campaign cadence
Continuous versus episodic campaign frequency.
L2 Adaptive · Calculated as: frequency · Why it matters: The operating-model tell: episodic cadence is the awareness-month pattern; continuous cadence is risk management.
PGM-09 - Feature adoption vs entitlement
Entitled capabilities actually in use.
L2 Adaptive · Calculated as: used / entitled · Why it matters: Value realization: capability you own but don't run is risk reduction left on the table.
Frequently asked questions
What are human risk management (HRM) metrics?
Human risk management metrics measure how people's real security behaviors, attitudes, competence and risk change over time - using signals from security integrations, simulations, assessments and workforce listening, not just training completion records. This library catalogs 90 of them across 10 domains, from engagement through behavior change to AI-agent governance.
How are HRM metrics different from security awareness training metrics?
Security awareness metrics typically report activity: completion rates, generic phishing click rates, attendance. HRM metrics report outcomes: real-world secure-behavior rates, human risk scores, incident correlation and measurable behavior change. Activity metrics show a program ran; outcome metrics show whether risk went down.
What human risk metrics should a CISO report to the board?
Lead with outcome and trend: human risk reduction over time (RSK-09), human-factor incident rate (OPS-01), workforce share above the competence threshold (CMP-08), AI / targeted-attack resilience (PHI-12), and high-risk segment movement (RSK-03). Pair each with the validation metric that proves the score predicts reality (OPS-02).
Is phishing click rate a good security metric?
On its own, no - a low click rate on generic simulations says little about resilience to targeted, AI-grade attacks. Use difficulty-adjusted click rate (PHI-11), reporting rate (PHI-04) and AI-attack resilience (PHI-12) to read genuine resilience, and treat raw click rate as a caveated benchmark.
What is a composite human risk score?
A per-person score built from real signals - observed behavior, attitudes, access and blast radius, real-world targeting, device posture and workplace context (RSK-01). Done well it is explainable (RSK-04), validated against incidents (OPS-02), and consumed by the SOC for access and triage decisions (OPS-10).
What is CyberQ?
CyberQ is OutThink's per-person security-competence score (CMP-01) - a holistic, defensible measure of what someone actually knows and can apply, designed to be tracked over time, split by behavior domain, and reported to executives in place of raw quiz scores.
How do you measure security behavior change?
Through integrations with the systems where behavior happens: identity providers for credential hygiene (BEH-04), DLP for data handling (BEH-05), payment workflows for verification adherence (BEH-11), EHR audit logs in healthcare (BEH-12), OT monitoring in industrial environments (BEH-13). The real-world secure-behavior rate (BEH-01) aggregates them.
What are vanity metrics in security awareness?
Metrics that look reassuring but don't indicate risk reduction - most famously training completion rate (ENG-01) and raw phishing click rate (PHI-01). They have legitimate uses as coverage and benchmark context, but a program reporting only vanity metrics cannot say whether human risk changed.
How do you measure AI agent risk?
Treat agents like a workforce: score their behavior (AIA-01), detect drift from their intended job (AIA-02), measure governance coverage of the agent estate (AIA-03), and read agent risk in the context of the human principal who owns them (AIA-04).
What is the HRM maturity model (L1-L4)?
Four levels describing how a program operates: L1 Reactive (generic training, activity metrics), L2 Adaptive (per-person engagement and behavior-led practice), L3 Proactive (real-world behavior measured via integrations, composite risk scoring), L4 Predictive (risk signals drive automated response - including for AI agents). Every metric in this library is tagged with the level where it becomes meaningful.