Understanding Sex-Offense Risk Assessment
If a risk score, evaluation, treatment report, or supervision assessment has suddenly appeared in your life, this guide helps you figure out what you are looking at, what the result actually means, and what questions to ask before anyone treats it as certainty.
Start Here
Risk assessment is not one thing. A report may contain several different kinds of information at once: a historical score, a changeable-risk assessment, a treatment-need rating, a risk category, a percentage, supervision recommendations, and professional comments.
Those are not automatically the same thing. A treatment recommendation is not automatically a recidivism probability. A supervision decision is not automatically the instrument score. And a professional opinion may include information that was never part of the actuarial result.
You do not need to master the statistics before you can begin. Start by identifying what tool was used, what result it produced, what outcome it is talking about, what time period applies, and what group the result is being compared with.
If you have a report or score in front of you
Find these five things before deciding what the result actually means.
Do first
- 11. Find the instrument name. Static-99R? CPORT? STABLE-2007? PCRA? SOTIPS? Something else?
- 22. Find the actual result. Is it a raw score, risk category, relative-risk level, estimated percentage, or several of those?
- 33. Find the outcome. Sexual rearrest? Sexual reconviction? Any recidivism? Treatment need? Short-term supervision concern?
Then do next
- 14. Find the time period. Five years? Ten years? Ongoing supervision? A shorter-term monitoring period?
- 25. Find the comparison group. What population, norm, or reference group is being used to interpret the score?
Remember
Go where you need to go
If you only need help understanding a report in front of you, Sections 1โ4 are the best place to start. The later sections explain particular tools, research findings, and questions to ask when the assessment is being used in a decision.
The core principle
Risk should be assessed as accurately, individually, transparently, and empirically as possible rather than inferred categorically from offense labels, intuition, or fear. Structured empirical assessment can add useful information without producing certainty about an individual future.
Why Am I Seeing a Risk Assessment?
What the process may look like before you ever get to the score.
Depending on the setting, you may be interviewed, records may be reviewed, one or more instruments may be scored, treatment providers or supervision officers may add information, and a final report may combine several different kinds of conclusions.
Different settings also use assessment for different purposes. Depending on the case, an assessment may inform sentencing, treatment planning, supervision intensity, release planning, institutional decisions, civil proceedings, or another specific decision. The purpose matters because a tool designed for one question should not automatically be treated as answering every other one.
If the assessment is part of a federal criminal case, the Federal Sex-Crime Process Guide explains where sentencing and related evaluations fit in the larger federal process.
You might see all of the following in one document:
- Historical or static risk score: a score based on historical factors;
- Dynamic or change-sensitive assessment: an assessment of factors intended to change over time;
- Risk category or relative-risk level: a group classification or comparison;
- Reference-group recidivism estimate: an observed or estimated rate tied to a comparison population;
- Treatment targets or needs: issues identified for treatment or intervention;
- Supervision recommendations: recommendations about case management or supervision;
- Professional judgment or an override: a conclusion or adjustment beyond the instrument output;
- Other case-specific comments: additional information the evaluator considers relevant.
The practical problem is that these pieces can look like one unified scientific conclusion even when they came from different methods and answer different questions.
The score and the final decision may not be the same thing
An instrument may produce one result while an evaluator, probation officer, agency, treatment provider, or decision-maker reaches a broader conclusion using additional information.
When that happens, separate the two: What did the instrument actually say? And what did the person or agency decide? Then ask what information caused the difference.
That distinction matters because a recommendation can be stricter, more lenient, or simply different from the instrument output. The difference may reflect another assessment, a treatment issue, an agency rule, case-specific information, or an explicit professional override.
How to Read the Result in Front of You
Start with the anatomy of the report before moving into the statistics behind it.
A fictional anatomy of a risk-assessment result
Imagine that a report contains the following kinds of entries. These are deliberately illustrative rather than real scoring instructions:
- Instrument: Static-99R
- Raw score: [score]
- Risk level: [risk category]
- Relative risk: [comparison with a reference group]
- Estimated five-year rate: [percentage for the applicable norm group]
- Professional conclusion: [treatment, supervision, or case recommendation]
Those lines are related, but they are not interchangeable.
A real report may include only some of these layers. For example, it may give a score and category without an absolute percentage, or it may discuss treatment needs without reporting a separate actuarial estimate. The point of this example is to show how different kinds of information fit together when they do appear.
What each line is doing
If the recommendation and score seem inconsistent
Ask whether another assessment, dynamic information, agency policy, case-specific information, or an override changed the final conclusion. Do not assume the instrument itself produced every statement that appears in the report.
A five-step interpretation workflow
1
Identify
2
Translate
3
Compare
4
Check
5
Question
Words You May See in a Report
A compact glossary for the terms that matter most.
Keep these definitions handy
The Concepts Behind the Score
Once you know what kind of result you are looking at, these concepts explain how to interpret it responsibly.
Relative risk vs. absolute risk
Relative risk answers a comparison question: higher or lower compared with whom? Absolute riskis an observed or estimated event rate over a defined period, such as a five-year rate in a particular reference group.
A person can be higher than a low-risk comparison group while the absolute event rate remains modest. A dramatic-sounding relative difference does not automatically mean a high individual probability.
Group prediction vs. individual certainty
Risk instruments estimate patterns across groups and place an individual within those empirical patterns. They do not observe the future. A group rate is evidence about a reference group, not an individual destiny.
Outcome definition and follow-up period
Rearrest, charge, reconviction, reincarceration, self-report, and detected offending are not interchangeable. Official outcomes can miss undetected conduct; self-report has different limitations.
A recidivism rate is incomplete unless it tells you the outcome definition, population, follow-up period, and starting point.
A five-year rearrest rate beginning at supervision start is not the same quantity as a ten-year reconviction rate beginning at release. Comparing them as if they were the same can create false precision.
Validation population and population fit
Every validation study has a population: a jurisdiction, setting, offense mix, sex composition, age range, entry point, and follow-up design. A tool validated in one population is not automatically calibrated for another.
Before relying on a percentage, ask which reference group generated it and whether that group resembles the person and setting at issue.
Static vs. dynamic factors
Static factors are historical facts that do not change because time has passed or treatment has occurred: for example, parts of a person's prior offense or supervision history. Static tools are mainly about baseline group risk.
Dynamic factors are intended to capture risk-relevant characteristics that can change. Some change over months or years; others may shift much more quickly.
A dynamic score is not a promise that change has been measured perfectly. It is an attempt to add current, change-sensitive information to historical baseline information.
Actuarial, structured professional judgment, and unstructured judgment
Actuarial tools use specified empirical items and scoring rules to place people into relative risk groups or categories.
Structured professional judgment (SPJ) also uses a defined framework, but leaves more room for professional synthesis of case information.
Unstructured judgment is professional intuition without a comparable standardized empirical structure.
Meta-analytic evidence in the SOLAR evidence matrix supports empirically derived actuarial approaches over unstructured professional intuition on average, with SPJ performing differently from both and generally falling between them.
Base rates
A base rate is how often the outcome occurs in the population being studied before a particular score is considered. When the outcome is uncommon, precise individual prediction becomes harder.
Even a tool that sorts people better than chance will still make errors when applied to a low-frequency outcome.
AUC: ranking, not probability
AUC is a discrimination statistic. In plain English: if you randomly select one person who later had the measured recidivism outcome and one who did not, the AUC estimates how often the tool ranks the person with the later outcome as higher risk.
AUC .70 does not mean 70% chance of recidivism
AUC does not itself give an individual probability. It does not establish calibration, causation, or certainty. Moderate discrimination can still contain useful information, but better-than-chance ranking is not the same as knowing what one person will do.
Calibration
Calibration asks a different question: do the predicted or reference-group percentages line up with the observed rates in the population where the tool is being used?
A tool can rank people reasonably well and still overpredict or underpredict absolute rates in another setting.
The workbook flags this issue for Static-99R/Static-2002R and CPORT. Static meta-analytic research found more stability in relative predictive accuracy than in absolute rates across samples. A Spanish CPORT validation found observed sexual recidivism substantially below developer expectations, illustrating why population-specific norms and reference groups matter.
Readers who want to go deeper into SOLAR's research sources, definitions, and evidence base can continue to Research & Data Resources.
A useful risk estimate can be informative without being certain. The question is not whether the tool is perfect, but whether it is being used for the right question, population, outcome, and purpose.
Baseline / Static Sexual-Recidivism Risk
Tools that ask what historical factors suggest about relative long-term risk.
What question does it ask?
How does this person's static historical profile compare with other eligible adult men on long-term sexual-recidivism risk?
What is it?
A static actuarial sexual-recidivism instrument using ten historical factors.
Who was it built for?
Adult men with qualifying sexual-offense histories under the instrument's coding and eligibility rules.
Where might you encounter it?
In evaluations that need a baseline actuarial estimate of sexual-recidivism risk, including some sentencing, treatment, supervision, civil, or correctional settings.
What output does it produce?
A score that can be interpreted using risk levels, relative-risk information, and current normative recidivism estimates tied to reference groups.
What does the evidence say?
The workbook identifies Static-99R as an established static actuarial tool, while emphasizing that age weighting was revised because age contributes meaningful predictive information and that absolute rates vary across samples.
What should be checked before use?
Eligibility, the current coding rules, the version used, the norm or reference group, and whether the case type fits the instrument.
Why does eligibility matter?
CSEM-only cases require particular care. Not every person convicted of a sexual offense is automatically appropriate to score with Static-99R.
What a Static-99R score does not mean
It is not a diagnosis, a moral-severity ranking, proof that someone will reoffend, proof that someone will not reoffend, or an individualized certainty. Current coding and current norms matter.
What question does it ask?
Like Static-99R, it asks about relative long-term sexual-recidivism risk from static historical information, but it uses a different item and domain structure.
Where might you encounter it?
In settings seeking a baseline actuarial sexual-recidivism assessment using the Static-2002R framework.
What goes into it?
Static historical domains including age and offense history, scored under specific coding rules.
What does the evidence say?
The workbook reports that revised age weights improved fit for older people and that relative predictive accuracy was more stable across samples than absolute recidivism rates within score groups.
Main limitation
Eligibility and the selected reference group matter. It is not designed for every CSEM-only case.
What should not be assumed?
An old score-to-percentage table should not be treated as timeless. Norms and reference-group choices matter to interpretation.
SOLAR takeaway
Static tools can provide an empirical baseline, but baseline is not destiny. Their value depends on proper eligibility, coding, norms, and interpretation.
Changeable / Dynamic Risk and Needs
These tools ask different questions from static baseline tools.
STABLE-2007
STABLE-2007 is designed to assess relatively stable but changeable risk and need factors relevant to sexual recidivism. Ratings draw on structured interview, file, treatment, and supervision information.
The output is used for risk/need formulation and treatment or supervision planning rather than as a stand-alone long-term probability.
Where might you encounter it? Most often in treatment or community-supervision settings where professionals are trying to understand changeable risk and need factors rather than relying only on historical baseline information.
A prospective Canadian study in the workbook followed 768 community-supervised adult men and found that STABLE measures predicted sexual, violent, and any recidivism. The workbook treats this as evidence that dynamic information can add clinically and practically relevant information beyond static history.
ACUTE-2007
ACUTE-2007 is aimed at shorter-term, rapidly changing concerns during community supervision. It is meant to be reassessed repeatedly and interpreted in its supervision context.
Where might you encounter it? In ongoing community supervision or monitoring where short-term changes may matter to immediate case management.
It does not establish a person's long-term actuarial risk and does not diagnose dangerousness.
Stable dynamic is not the same as acute
STABLE-2007 focuses on changeable factors that generally move over a longer period. ACUTE-2007 focuses on shorter-term changes that may matter for ongoing supervision. Neither should be treated as if it were answering the same question as a static baseline score.
Dynamic improvement does not guarantee safety
A lower dynamic score can be meaningful evidence of change without proving that no future offending will occur. The same is true in the other direction: a concerning rating is not certainty that an offense will occur.
CSEM-Specific Assessment
Specialized tools should be judged within the populations and questions they were designed for.
What question does it ask?
CPORT estimates relative sexual-recidivism risk among adult men convicted of CSEM offenses.
What is it?
A primarily static actuarial tool using seven binary factors involving age, criminal or supervision history, contact sexual offending, sexual-interest evidence, and content indicators.
Where might you encounter it?
In assessments involving adult men with CSEM-related offenses, where a CSEM-specific empirical risk estimate is being considered.
What output does it produce?
A summed score used for relative-risk grouping. The score is not, by itself, an individualized probability.
What does the broader validation evidence say?
The workbook includes development and independent validation evidence showing meaningful predictive discrimination, including an independent validation AUC of .70 in a small 80-person cohort and combined sample AUCs of .72 for any sexual recidivism and .74 for a new CSEM offense.
What happened in the federal cohort?
In a large federal CSEM validation cohort, CPORT produced modest discrimination for five-year sexual rearrest: AUC .62, with a 95% confidence interval of .58โ.65.
Why does population fit matter?
A Spanish validation observed a much lower five-year sexual recidivism base rate than the development sample and found calibration concerns relative to developer expectations.
What coding issue matters in the federal study?
The federal implementation used MITRE-extracted data elements that differed from some standard CPORT scoring. That limits how broadly the federal result should be generalized.
CPORT is neither scientifically worthless nor a universal answer
The workbook supports a mixed, population-sensitive reading. CPORT has real empirical validation evidence; performance and calibration vary; federal implementation produced only modest discrimination; and later or stronger studies should not be erased because one cohort was weaker.
CASIC
CASIC is a structured proxy/index related to evidence of sexual interest in children. It was developed in part to operationalize a CPORT-related factor when direct admission or other evidence is unavailable.
Where might you encounter it? Within CPORT-related CSEM assessment work where an evaluator needs a structured way to code the relevant sexual-interest factor.
The workbook describes CASIC as relying on historical or behavioral correlates, not as a stand-alone recidivism instrument.
CASIC is not a diagnosis
A CASIC threshold does not diagnose pedophilia, does not establish that a person will sexually offend, and should not be presented as a stand-alone recidivism probability.
General Federal Risk / Needs Assessment
General recidivism tools can contain relevant information without becoming specialized sexual-risk instruments.
PCRA and PCRA-R
The federal Post Conviction Risk Assessment family is used for general recidivism risk, criminogenic needs, supervision planning, and allocation of intervention resources in federal post-conviction supervision.
It combines criminal-history information with dynamic needs and officer-assessment fields.
Where might you encounter it? In federal post-conviction supervision, where probation officers use a general risk-and-needs framework to support supervision planning. For the practical rules and decisions that may follow, see SOLAR's Supervision Conditions Survival Guide.
That scope distinction matters. PCRA can include factors that correlate with sexual recidivism, but its primary validated purpose is general federal risk and needs.
Correlation with a specialized outcome does not transform it into a specialized sexual-risk probability calculator.
Do not make PCRA answer a question it was not built to answer
In the federal CSEM cohort, PCRA's AUC for five-year sexual rearrest was approximately .61. That is a discrimination result for that cohort and outcome. It does not tell a court one individual person's probability of committing another sex offense.
Treatment Progress / Change-Sensitive Assessment
These instruments are designed to capture information that a purely historical score cannot.
SOTIPS
The Sex Offender Treatment Intervention and Progress Scale is a structured, change-sensitive assessment used in treatment and supervision contexts. It measures dynamic, treatment-relevant factors over time rather than asking only what happened in the past.
Where might you encounter it? In community treatment or supervision where practitioners are monitoring treatment-relevant needs and change over time.
The primary validation study in the workbook involved 759 adult men under correctional supervision in Vermont community treatment. SOTIPS ratings predicted sexual, violent, and any recidivism and return to prison.
Reductions in SOTIPS scores were associated with lower recidivism, and combining SOTIPS with Static-99R improved prediction in that validation study.
Read the SOTIPS AUC range carefully
The reported SOTIPS AUC range of .60โ.85 spans different outcomes and assessment times. The combined SOTIPS + Static-99R range of .67โ.89 also spans multiple outcomes. Neither range should be presented as one single sexual-recidivism AUC.
VRS-SO
The Violence Risk ScaleโSexual Offense version combines static and dynamic information. It is designed to assess baseline risk, treatment targets, and treatment-related change and is more intensive than a quick screening instrument.
Where might you encounter it? In more intensive treatment or evaluation settings where baseline risk and treatment-related change are both being assessed.
The workbook's foundational validation involved 321 adult men and found prediction of sexual and nonsexual violent recidivism over an average follow-up of about ten years.
Later multisite work developed updated risk categories and five- and ten-year recidivism estimates using pretreatment risk and change information.
Change scores are evidence, not verdicts
A VRS-SO change score does not prove someone is safe or unsafe. Treatment-related change can matter while the resulting inference remains probabilistic and group based.
Where Structured Professional Judgment Fits
Structured professional judgment is not the same thing as unsupported clinical intuition.
SPJ uses defined risk factors and a structured process, while leaving room for professional synthesis. That makes it meaningfully different from an unstructured impression such as โthis person feels dangerousโ or โmy experience tells me the score is wrong.โ
The meta-analytic evidence in the workbook found empirically derived actuarial approaches more accurate than unstructured professional judgment across sexual, violent, and any recidivism outcomes.
SPJ performance was intermediate between actuarial and unstructured judgment in that synthesis.
The practical lesson is not โscores only.โ Individualized professional information can matter.
The lesson is that a professional opinion does not become individualized science merely because a professional expresses it. A defensible assessment should connect judgment to a structured method, relevant case facts, and an empirical baseline.
Separate instrument output from professional judgment
If a report says the actuarial result is one thing but the final classification or recommendation is another, ask where the difference came from. Was another instrument used? Were dynamic factors considered? Was there a formal override? Was the change driven by policy rather than the assessment itself?
Worked Example: One Federal CSEM Cohort
A bounded example of outcome definition, base rates, population fit, and modest discrimination.
The federal CSEM risk-tool study examined a validation cohort of 5,768 male federal CSEM supervisees using a fixed 60-month follow-up beginning at supervision start or initial PCRA assessment.
- Outcome: rearrest for any new sexual offense.
- Observed five-year rate: 4.5% (262 of 5,768).
- Contact sex-crime rearrest: fewer than 1%.
- PCRA discrimination: AUC .61 (95% CI .58โ.64).
- CPORT discrimination: AUC .62 (95% CI .58โ.65).
Those facts teach several different things at once. The outcome was rearrest, not all offending. The population was male, federal, and CSEM-specific. The follow-up was fixed at five years. The sexual-rearrest base rate was low. And both AUCs showed modest ranking ability rather than individualized certainty.
What this example does not show
It does not prove CPORT is useless, prove PCRA is a sexual-risk instrument, establish lifetime risk, or turn either AUC into a probability.
The federal CPORT data extraction also differed from some standard scoring elements, which limits how broadly the result should be generalized.
Cohort boundary
This worked example does not combine the separate nearly 6,900-person FY2017โFY2021 federal supervision-override cohort with the 5,768-person validation cohort. They answer different questions and should remain separate.
Questions to Ask When Someone Gives You a Risk Score
You do not need to self-score a professional instrument to ask whether it is being interpreted responsibly.
Use the five-step workflow first
Identify โ Translate โ Compare โ Check โ Question.
That short workflow gets you oriented. The fuller checklist below is for the details that matter when a score affects treatment, supervision, sentencing, litigation, reporting, or policy.
Full risk-score review checklist
Verify before acting
Who to ask
What to ask
What to save
If the disagreement involves a government action or legal process, SOLAR's Your Rights at Every Stage guide can help identify the separate legal questions. A disagreement about scoring or interpretation, by itself, does not establish a rights violation.
Do not turn this guide into a self-scoring exercise
Some instruments require professional training, controlled coding rules, or records that a reader may not have. The goal here is to understand and question interpretation, not to produce an unofficial score and assume it is valid.
What Risk Language Does Not Mean
These are the most common interpretation errors to stop before they spread.
- โLow riskโ does not mean zero risk.
- โHigh relative riskโ does not automatically mean a high absolute probability.
- AUC .70 does not mean a 70% chance of recidivism.
- A group recidivism rate is not an individual destiny.
- A static score is not moral severity.
- CASIC is not a pedophilia diagnosis.
- CPORT is not proof someone will offend.
- PCRA is not a specialized sexual-risk probability calculator.
- Dynamic improvement does not guarantee no future offending.
- A professional recommendation is not automatically the same thing as the instrument result.
- Unstructured intuition is not individualized science simply because a professional expresses it.
Sources and Verification
The evidence below is the public-facing source backbone for the guide.
How to read the evidence
Each numerical claim in this guide is tied to the canonical SOLAR Evidence Matrix and its verified source record. Population, outcome, follow-up, and important caveats are preserved in the text rather than collapsed into a generic โrecidivism rate.โ
Sources & verification
- Federal CSEM risk-tool study โ Cohen (2023), Federal ProbationFederal 5,768-person CSEM validation cohort; five-year sexual rearrest; PCRA and CPORT discrimination; federal implementation limitations.
- Hanson & Morton-Bourgon (2009) โ risk-assessment accuracy meta-analysisComparative evidence on actuarial assessment, structured professional judgment, and unstructured professional judgment.
- Helmus et al. (2012) โ revised age weightsAge-related revision evidence for Static-99/Static-2002 and older-person fit.
- Helmus et al. (2012) โ Static absolute rates across samplesKey calibration evidence: relative accuracy can be more stable than absolute recidivism rates across samples.
- Static-99R Coding Rules Revised 2016Eligibility, coding, proper version use, and current-norm guidance.
- Static-99R Evaluators WorkbookReference-group, risk-level, relative-risk, and normative interpretation guidance.
- Static-2002 coding rules โ Public Safety CanadaAuthoritative coding, target-population, domain, and interpretation guidance for Static-2002/Static-2002R.
- Hanson, Helmus & Harris (2015) โ STABLE-2007 prospective studyProspective evidence on dynamic risk/need assessment alongside static measures.
- ACUTE-2007 โ BJA Public Safety Risk Assessment ClearinghouseOfficial tool profile describing ACUTE-2007 as a short-term dynamic monitoring instrument for sexual recidivism risk.
- Seto & Eke (2015) โ CPORT developmentDevelopment of the CSEM-specific CPORT and its initial five-year outcome evidence.
- Eke, Helmus & Seto (2019) โ CPORT validationIndependent validation and combined-sample discrimination evidence.
- Soldino et al. (2021) โ Spanish CPORT validationLow base-rate validation context and calibration concerns relative to developer expectations.
- Critical review of CPORT use in CSEM-exclusive cases (2023)Population-specific limitations, small recidivist counts, missing data, and U.S.-validation concerns at the time of review.
- Seto & Eke (2017) โ CASICPrimary CASIC source on behavioral correlates of admitted sexual interest in children and use within CPORT-related assessment.
- Johnson et al. (2011) โ PCRA construction and validationPrimary scope source for PCRA as a general federal post-conviction risk-and-needs instrument.
- McGrath, Lasher & Cumming (2012) โ SOTIPS validationPrimary 759-person validation source for SOTIPS predictive validity, change, and incremental value with Static-99R.
- Olver et al. (2007) โ VRS-SO validity and reliabilityFoundational validation for VRS-SO risk and treatment-change assessment.
- Olver et al. (2018) โ VRS-SO updated risk categoriesMultisite updated risk categories and five- and ten-year recidivism estimates incorporating pretreatment risk and change.
