โ† Back to Resources
SOLAR Resource Guide

Understanding Sex-Offense Risk Assessment

If a risk score, evaluation, treatment report, or supervision assessment has suddenly appeared in your life, this guide helps you figure out what you are looking at, what the result actually means, and what questions to ask before anyone treats it as certainty.

Jump to Sources

Start Here

Risk assessment is not one thing. A report may contain several different kinds of information at once: a historical score, a changeable-risk assessment, a treatment-need rating, a risk category, a percentage, supervision recommendations, and professional comments.

Those are not automatically the same thing. A treatment recommendation is not automatically a recidivism probability. A supervision decision is not automatically the instrument score. And a professional opinion may include information that was never part of the actuarial result.

You do not need to master the statistics before you can begin. Start by identifying what tool was used, what result it produced, what outcome it is talking about, what time period applies, and what group the result is being compared with.

If you have a report or score in front of you

Find these five things before deciding what the result actually means.

Do first

  • 1
    1. Find the instrument name. Static-99R? CPORT? STABLE-2007? PCRA? SOTIPS? Something else?
  • 2
    2. Find the actual result. Is it a raw score, risk category, relative-risk level, estimated percentage, or several of those?
  • 3
    3. Find the outcome. Sexual rearrest? Sexual reconviction? Any recidivism? Treatment need? Short-term supervision concern?

Then do next

  • 1
    4. Find the time period. Five years? Ten years? Ongoing supervision? A shorter-term monitoring period?
  • 2
    5. Find the comparison group. What population, norm, or reference group is being used to interpret the score?

Remember

If you can answer those five questions, you can usually begin interpreting the result. If you cannot, the report may be giving you a conclusion without enough information to understand what the conclusion means.

Go where you need to go

If you only need help understanding a report in front of you, Sections 1โ€“4 are the best place to start. The later sections explain particular tools, research findings, and questions to ask when the assessment is being used in a decision.

The core principle

Risk should be assessed as accurately, individually, transparently, and empirically as possible rather than inferred categorically from offense labels, intuition, or fear. Structured empirical assessment can add useful information without producing certainty about an individual future.

Why Am I Seeing a Risk Assessment?

What the process may look like before you ever get to the score.

Depending on the setting, you may be interviewed, records may be reviewed, one or more instruments may be scored, treatment providers or supervision officers may add information, and a final report may combine several different kinds of conclusions.

Different settings also use assessment for different purposes. Depending on the case, an assessment may inform sentencing, treatment planning, supervision intensity, release planning, institutional decisions, civil proceedings, or another specific decision. The purpose matters because a tool designed for one question should not automatically be treated as answering every other one.

If the assessment is part of a federal criminal case, the Federal Sex-Crime Process Guide explains where sentencing and related evaluations fit in the larger federal process.

You might see all of the following in one document:

  • Historical or static risk score: a score based on historical factors;
  • Dynamic or change-sensitive assessment: an assessment of factors intended to change over time;
  • Risk category or relative-risk level: a group classification or comparison;
  • Reference-group recidivism estimate: an observed or estimated rate tied to a comparison population;
  • Treatment targets or needs: issues identified for treatment or intervention;
  • Supervision recommendations: recommendations about case management or supervision;
  • Professional judgment or an override: a conclusion or adjustment beyond the instrument output;
  • Other case-specific comments: additional information the evaluator considers relevant.

The practical problem is that these pieces can look like one unified scientific conclusion even when they came from different methods and answer different questions.

The score and the final decision may not be the same thing

An instrument may produce one result while an evaluator, probation officer, agency, treatment provider, or decision-maker reaches a broader conclusion using additional information.

When that happens, separate the two: What did the instrument actually say? And what did the person or agency decide? Then ask what information caused the difference.

That distinction matters because a recommendation can be stricter, more lenient, or simply different from the instrument output. The difference may reflect another assessment, a treatment issue, an agency rule, case-specific information, or an explicit professional override.

How to Read the Result in Front of You

Start with the anatomy of the report before moving into the statistics behind it.

A fictional anatomy of a risk-assessment result

Imagine that a report contains the following kinds of entries. These are deliberately illustrative rather than real scoring instructions:

  • Instrument: Static-99R
  • Raw score: [score]
  • Risk level: [risk category]
  • Relative risk: [comparison with a reference group]
  • Estimated five-year rate: [percentage for the applicable norm group]
  • Professional conclusion: [treatment, supervision, or case recommendation]

Those lines are related, but they are not interchangeable.

A real report may include only some of these layers. For example, it may give a score and category without an absolute percentage, or it may discuss treatment needs without reporting a separate actuarial estimate. The point of this example is to show how different kinds of information fit together when they do appear.

What each line is doing

Raw score
The instrument's scored total. By itself, it is not automatically a probability.
Risk category
A group classification used to organize or interpret scores.
Relative risk
A comparison with another group. It answers higher or lower compared with whom.
Absolute estimate
A group-based observed or estimated event rate over a defined follow-up period.
Reference group
The population whose data are being used to interpret the score or percentage.
Professional recommendation
A conclusion that may incorporate information beyond the instrument itself.

If the recommendation and score seem inconsistent

Ask whether another assessment, dynamic information, agency policy, case-specific information, or an override changed the final conclusion. Do not assume the instrument itself produced every statement that appears in the report.

A five-step interpretation workflow

1

Identify

Identify the tool, version, and type of assessment.

2

Translate

Translate the raw score, category, percentage, or recommendation into plain language.

3

Compare

Find the population, norm, or reference group being used.

4

Check

Check the outcome, follow-up, coding rules, dynamic information, and any override.

5

Question

Question any conclusion that goes beyond what the tool was designed or validated to say.

Words You May See in a Report

A compact glossary for the terms that matter most.

Keep these definitions handy

Recidivism
A new measured criminal-justice or study outcome after a defined starting point. Always ask exactly how the study defined it.
Static
Based on historical facts that do not change because treatment occurs or time passes.
Dynamic
Based on factors intended to capture characteristics that can change over time.
Actuarial
Uses defined items and scoring rules derived from empirical data.
Norm / reference group
The comparison group used to interpret a score, category, relative-risk value, or estimated percentage.
Validation
Testing how a tool performs in a population, setting, or sample beyond the data used to create or develop it.
Risk category
A group classification. It is not a statement that one individual will or will not reoffend.
Override
A departure from, adjustment to, or broader conclusion beyond the instrument result.
AUC
A statistic describing how well a tool ranks higher- versus lower-risk cases. It is not an individual probability.
Calibration
How well estimated or expected event rates line up with what is actually observed in a population.
CSEM
Child sexual exploitation material. Some instruments are designed specifically for people with CSEM-related offenses rather than sexual offenses generally.

The Concepts Behind the Score

Once you know what kind of result you are looking at, these concepts explain how to interpret it responsibly.

Relative risk vs. absolute risk

Relative risk answers a comparison question: higher or lower compared with whom? Absolute riskis an observed or estimated event rate over a defined period, such as a five-year rate in a particular reference group.

A person can be higher than a low-risk comparison group while the absolute event rate remains modest. A dramatic-sounding relative difference does not automatically mean a high individual probability.

Group prediction vs. individual certainty

Risk instruments estimate patterns across groups and place an individual within those empirical patterns. They do not observe the future. A group rate is evidence about a reference group, not an individual destiny.

Outcome definition and follow-up period

Rearrest, charge, reconviction, reincarceration, self-report, and detected offending are not interchangeable. Official outcomes can miss undetected conduct; self-report has different limitations.

A recidivism rate is incomplete unless it tells you the outcome definition, population, follow-up period, and starting point.

A five-year rearrest rate beginning at supervision start is not the same quantity as a ten-year reconviction rate beginning at release. Comparing them as if they were the same can create false precision.

Validation population and population fit

Every validation study has a population: a jurisdiction, setting, offense mix, sex composition, age range, entry point, and follow-up design. A tool validated in one population is not automatically calibrated for another.

Before relying on a percentage, ask which reference group generated it and whether that group resembles the person and setting at issue.

Static vs. dynamic factors

Static factors are historical facts that do not change because time has passed or treatment has occurred: for example, parts of a person's prior offense or supervision history. Static tools are mainly about baseline group risk.

Dynamic factors are intended to capture risk-relevant characteristics that can change. Some change over months or years; others may shift much more quickly.

A dynamic score is not a promise that change has been measured perfectly. It is an attempt to add current, change-sensitive information to historical baseline information.

Actuarial, structured professional judgment, and unstructured judgment

Actuarial tools use specified empirical items and scoring rules to place people into relative risk groups or categories.

Structured professional judgment (SPJ) also uses a defined framework, but leaves more room for professional synthesis of case information.

Unstructured judgment is professional intuition without a comparable standardized empirical structure.

Meta-analytic evidence in the SOLAR evidence matrix supports empirically derived actuarial approaches over unstructured professional intuition on average, with SPJ performing differently from both and generally falling between them.

Base rates

A base rate is how often the outcome occurs in the population being studied before a particular score is considered. When the outcome is uncommon, precise individual prediction becomes harder.

Even a tool that sorts people better than chance will still make errors when applied to a low-frequency outcome.

AUC: ranking, not probability

AUC is a discrimination statistic. In plain English: if you randomly select one person who later had the measured recidivism outcome and one who did not, the AUC estimates how often the tool ranks the person with the later outcome as higher risk.

AUC .70 does not mean 70% chance of recidivism

AUC does not itself give an individual probability. It does not establish calibration, causation, or certainty. Moderate discrimination can still contain useful information, but better-than-chance ranking is not the same as knowing what one person will do.

Calibration

Calibration asks a different question: do the predicted or reference-group percentages line up with the observed rates in the population where the tool is being used?

A tool can rank people reasonably well and still overpredict or underpredict absolute rates in another setting.

The workbook flags this issue for Static-99R/Static-2002R and CPORT. Static meta-analytic research found more stability in relative predictive accuracy than in absolute rates across samples. A Spanish CPORT validation found observed sexual recidivism substantially below developer expectations, illustrating why population-specific norms and reference groups matter.

Readers who want to go deeper into SOLAR's research sources, definitions, and evidence base can continue to Research & Data Resources.

A useful risk estimate can be informative without being certain. The question is not whether the tool is perfect, but whether it is being used for the right question, population, outcome, and purpose.

Baseline / Static Sexual-Recidivism Risk

Tools that ask what historical factors suggest about relative long-term risk.

What question does it ask?

How does this person's static historical profile compare with other eligible adult men on long-term sexual-recidivism risk?

What is it?

A static actuarial sexual-recidivism instrument using ten historical factors.

Who was it built for?

Adult men with qualifying sexual-offense histories under the instrument's coding and eligibility rules.

Where might you encounter it?

In evaluations that need a baseline actuarial estimate of sexual-recidivism risk, including some sentencing, treatment, supervision, civil, or correctional settings.

What output does it produce?

A score that can be interpreted using risk levels, relative-risk information, and current normative recidivism estimates tied to reference groups.

What does the evidence say?

The workbook identifies Static-99R as an established static actuarial tool, while emphasizing that age weighting was revised because age contributes meaningful predictive information and that absolute rates vary across samples.

What should be checked before use?

Eligibility, the current coding rules, the version used, the norm or reference group, and whether the case type fits the instrument.

Why does eligibility matter?

CSEM-only cases require particular care. Not every person convicted of a sexual offense is automatically appropriate to score with Static-99R.

What a Static-99R score does not mean

It is not a diagnosis, a moral-severity ranking, proof that someone will reoffend, proof that someone will not reoffend, or an individualized certainty. Current coding and current norms matter.

What question does it ask?

Like Static-99R, it asks about relative long-term sexual-recidivism risk from static historical information, but it uses a different item and domain structure.

Where might you encounter it?

In settings seeking a baseline actuarial sexual-recidivism assessment using the Static-2002R framework.

What goes into it?

Static historical domains including age and offense history, scored under specific coding rules.

What does the evidence say?

The workbook reports that revised age weights improved fit for older people and that relative predictive accuracy was more stable across samples than absolute recidivism rates within score groups.

Main limitation

Eligibility and the selected reference group matter. It is not designed for every CSEM-only case.

What should not be assumed?

An old score-to-percentage table should not be treated as timeless. Norms and reference-group choices matter to interpretation.

SOLAR takeaway

Static tools can provide an empirical baseline, but baseline is not destiny. Their value depends on proper eligibility, coding, norms, and interpretation.

Changeable / Dynamic Risk and Needs

These tools ask different questions from static baseline tools.

STABLE-2007

STABLE-2007 is designed to assess relatively stable but changeable risk and need factors relevant to sexual recidivism. Ratings draw on structured interview, file, treatment, and supervision information.

The output is used for risk/need formulation and treatment or supervision planning rather than as a stand-alone long-term probability.

Where might you encounter it? Most often in treatment or community-supervision settings where professionals are trying to understand changeable risk and need factors rather than relying only on historical baseline information.

A prospective Canadian study in the workbook followed 768 community-supervised adult men and found that STABLE measures predicted sexual, violent, and any recidivism. The workbook treats this as evidence that dynamic information can add clinically and practically relevant information beyond static history.

ACUTE-2007

ACUTE-2007 is aimed at shorter-term, rapidly changing concerns during community supervision. It is meant to be reassessed repeatedly and interpreted in its supervision context.

Where might you encounter it? In ongoing community supervision or monitoring where short-term changes may matter to immediate case management.

It does not establish a person's long-term actuarial risk and does not diagnose dangerousness.

Stable dynamic is not the same as acute

STABLE-2007 focuses on changeable factors that generally move over a longer period. ACUTE-2007 focuses on shorter-term changes that may matter for ongoing supervision. Neither should be treated as if it were answering the same question as a static baseline score.

Dynamic improvement does not guarantee safety

A lower dynamic score can be meaningful evidence of change without proving that no future offending will occur. The same is true in the other direction: a concerning rating is not certainty that an offense will occur.

CSEM-Specific Assessment

Specialized tools should be judged within the populations and questions they were designed for.

What question does it ask?

CPORT estimates relative sexual-recidivism risk among adult men convicted of CSEM offenses.

What is it?

A primarily static actuarial tool using seven binary factors involving age, criminal or supervision history, contact sexual offending, sexual-interest evidence, and content indicators.

Where might you encounter it?

In assessments involving adult men with CSEM-related offenses, where a CSEM-specific empirical risk estimate is being considered.

What output does it produce?

A summed score used for relative-risk grouping. The score is not, by itself, an individualized probability.

What does the broader validation evidence say?

The workbook includes development and independent validation evidence showing meaningful predictive discrimination, including an independent validation AUC of .70 in a small 80-person cohort and combined sample AUCs of .72 for any sexual recidivism and .74 for a new CSEM offense.

What happened in the federal cohort?

In a large federal CSEM validation cohort, CPORT produced modest discrimination for five-year sexual rearrest: AUC .62, with a 95% confidence interval of .58โ€“.65.

Why does population fit matter?

A Spanish validation observed a much lower five-year sexual recidivism base rate than the development sample and found calibration concerns relative to developer expectations.

What coding issue matters in the federal study?

The federal implementation used MITRE-extracted data elements that differed from some standard CPORT scoring. That limits how broadly the federal result should be generalized.

CPORT is neither scientifically worthless nor a universal answer

The workbook supports a mixed, population-sensitive reading. CPORT has real empirical validation evidence; performance and calibration vary; federal implementation produced only modest discrimination; and later or stronger studies should not be erased because one cohort was weaker.

CASIC

CASIC is a structured proxy/index related to evidence of sexual interest in children. It was developed in part to operationalize a CPORT-related factor when direct admission or other evidence is unavailable.

Where might you encounter it? Within CPORT-related CSEM assessment work where an evaluator needs a structured way to code the relevant sexual-interest factor.

The workbook describes CASIC as relying on historical or behavioral correlates, not as a stand-alone recidivism instrument.

CASIC is not a diagnosis

A CASIC threshold does not diagnose pedophilia, does not establish that a person will sexually offend, and should not be presented as a stand-alone recidivism probability.

General Federal Risk / Needs Assessment

General recidivism tools can contain relevant information without becoming specialized sexual-risk instruments.

PCRA and PCRA-R

The federal Post Conviction Risk Assessment family is used for general recidivism risk, criminogenic needs, supervision planning, and allocation of intervention resources in federal post-conviction supervision.

It combines criminal-history information with dynamic needs and officer-assessment fields.

Where might you encounter it? In federal post-conviction supervision, where probation officers use a general risk-and-needs framework to support supervision planning. For the practical rules and decisions that may follow, see SOLAR's Supervision Conditions Survival Guide.

That scope distinction matters. PCRA can include factors that correlate with sexual recidivism, but its primary validated purpose is general federal risk and needs.

Correlation with a specialized outcome does not transform it into a specialized sexual-risk probability calculator.

Do not make PCRA answer a question it was not built to answer

In the federal CSEM cohort, PCRA's AUC for five-year sexual rearrest was approximately .61. That is a discrimination result for that cohort and outcome. It does not tell a court one individual person's probability of committing another sex offense.

Treatment Progress / Change-Sensitive Assessment

These instruments are designed to capture information that a purely historical score cannot.

SOTIPS

The Sex Offender Treatment Intervention and Progress Scale is a structured, change-sensitive assessment used in treatment and supervision contexts. It measures dynamic, treatment-relevant factors over time rather than asking only what happened in the past.

Where might you encounter it? In community treatment or supervision where practitioners are monitoring treatment-relevant needs and change over time.

The primary validation study in the workbook involved 759 adult men under correctional supervision in Vermont community treatment. SOTIPS ratings predicted sexual, violent, and any recidivism and return to prison.

Reductions in SOTIPS scores were associated with lower recidivism, and combining SOTIPS with Static-99R improved prediction in that validation study.

Read the SOTIPS AUC range carefully

The reported SOTIPS AUC range of .60โ€“.85 spans different outcomes and assessment times. The combined SOTIPS + Static-99R range of .67โ€“.89 also spans multiple outcomes. Neither range should be presented as one single sexual-recidivism AUC.

VRS-SO

The Violence Risk Scaleโ€“Sexual Offense version combines static and dynamic information. It is designed to assess baseline risk, treatment targets, and treatment-related change and is more intensive than a quick screening instrument.

Where might you encounter it? In more intensive treatment or evaluation settings where baseline risk and treatment-related change are both being assessed.

The workbook's foundational validation involved 321 adult men and found prediction of sexual and nonsexual violent recidivism over an average follow-up of about ten years.

Later multisite work developed updated risk categories and five- and ten-year recidivism estimates using pretreatment risk and change information.

Change scores are evidence, not verdicts

A VRS-SO change score does not prove someone is safe or unsafe. Treatment-related change can matter while the resulting inference remains probabilistic and group based.

Where Structured Professional Judgment Fits

Structured professional judgment is not the same thing as unsupported clinical intuition.

SPJ uses defined risk factors and a structured process, while leaving room for professional synthesis. That makes it meaningfully different from an unstructured impression such as โ€œthis person feels dangerousโ€ or โ€œmy experience tells me the score is wrong.โ€

The meta-analytic evidence in the workbook found empirically derived actuarial approaches more accurate than unstructured professional judgment across sexual, violent, and any recidivism outcomes.

SPJ performance was intermediate between actuarial and unstructured judgment in that synthesis.

The practical lesson is not โ€œscores only.โ€ Individualized professional information can matter.

The lesson is that a professional opinion does not become individualized science merely because a professional expresses it. A defensible assessment should connect judgment to a structured method, relevant case facts, and an empirical baseline.

Separate instrument output from professional judgment

If a report says the actuarial result is one thing but the final classification or recommendation is another, ask where the difference came from. Was another instrument used? Were dynamic factors considered? Was there a formal override? Was the change driven by policy rather than the assessment itself?

Worked Example: One Federal CSEM Cohort

A bounded example of outcome definition, base rates, population fit, and modest discrimination.

The federal CSEM risk-tool study examined a validation cohort of 5,768 male federal CSEM supervisees using a fixed 60-month follow-up beginning at supervision start or initial PCRA assessment.

  • Outcome: rearrest for any new sexual offense.
  • Observed five-year rate: 4.5% (262 of 5,768).
  • Contact sex-crime rearrest: fewer than 1%.
  • PCRA discrimination: AUC .61 (95% CI .58โ€“.64).
  • CPORT discrimination: AUC .62 (95% CI .58โ€“.65).

Those facts teach several different things at once. The outcome was rearrest, not all offending. The population was male, federal, and CSEM-specific. The follow-up was fixed at five years. The sexual-rearrest base rate was low. And both AUCs showed modest ranking ability rather than individualized certainty.

What this example does not show

It does not prove CPORT is useless, prove PCRA is a sexual-risk instrument, establish lifetime risk, or turn either AUC into a probability.

The federal CPORT data extraction also differed from some standard scoring elements, which limits how broadly the result should be generalized.

Cohort boundary

This worked example does not combine the separate nearly 6,900-person FY2017โ€“FY2021 federal supervision-override cohort with the 5,768-person validation cohort. They answer different questions and should remain separate.

Questions to Ask When Someone Gives You a Risk Score

You do not need to self-score a professional instrument to ask whether it is being interpreted responsibly.

Use the five-step workflow first

Identify โ†’ Translate โ†’ Compare โ†’ Check โ†’ Question.

That short workflow gets you oriented. The fuller checklist below is for the details that matter when a score affects treatment, supervision, sentencing, litigation, reporting, or policy.

Full risk-score review checklist

Verify before acting

Who to ask

The evaluator or agency using the score, plus counsel or another qualified professional when the score affects a legal decision.

What to ask

Ask for the instrument name, version, eligibility basis, coding rules, outcome, follow-up period, reference group, and any override rationale.

What to save

Save the written report, score sheet if disclosure is permitted, cited norms, evaluator explanation, corrections, and any written response to a disputed coding item.

If the disagreement involves a government action or legal process, SOLAR's Your Rights at Every Stage guide can help identify the separate legal questions. A disagreement about scoring or interpretation, by itself, does not establish a rights violation.

Do not turn this guide into a self-scoring exercise

Some instruments require professional training, controlled coding rules, or records that a reader may not have. The goal here is to understand and question interpretation, not to produce an unofficial score and assume it is valid.

What Risk Language Does Not Mean

These are the most common interpretation errors to stop before they spread.

  • โ€œLow riskโ€ does not mean zero risk.
  • โ€œHigh relative riskโ€ does not automatically mean a high absolute probability.
  • AUC .70 does not mean a 70% chance of recidivism.
  • A group recidivism rate is not an individual destiny.
  • A static score is not moral severity.
  • CASIC is not a pedophilia diagnosis.
  • CPORT is not proof someone will offend.
  • PCRA is not a specialized sexual-risk probability calculator.
  • Dynamic improvement does not guarantee no future offending.
  • A professional recommendation is not automatically the same thing as the instrument result.
  • Unstructured intuition is not individualized science simply because a professional expresses it.

Sources and Verification

The evidence below is the public-facing source backbone for the guide.

How to read the evidence

Each numerical claim in this guide is tied to the canonical SOLAR Evidence Matrix and its verified source record. Population, outcome, follow-up, and important caveats are preserved in the text rather than collapsed into a generic โ€œrecidivism rate.โ€

Sources & verification

Source inventory drawn from the canonical SOLAR Evidence Matrix. Workbook records list these sources as verified through August 22โ€“23, 2026.