RMF assessment software
Judge RMF assessment software on the second assessment, not the first. Prior determinations preserved, adjudicated false positives remembered, scope derived from what changed, and every result naming the system state it was performed against.
Most RMF assessment software demos beautifully against an empty system. You create a boundary, select a baseline, work through some controls, attach evidence, produce a report. It looks like the job.
The job is the second assessment. And the fifth. The one where four hundred of the determinations you are about to make are identical to last time, where the same eighty false positives are back, and where somebody has to work out which of it all is still true.
Evaluate for that.
Six requirements
1. Assessment happens below the control. Per-CCI results, because a control decomposes into assertions that can have different answers. Where the model holds one status per control, the detail moves into a comment field and stops being queryable.
2. Every result names what it was assessed against. Not just a date — the architecture version. A determination from fourteen months ago against a version the system is still on is more current than one from last month against a version since superseded. Dates cannot tell you that.
3. Prior adjudication survives. A false positive judged once stays judged across every subsequent upload of the same scan. This is the capability practitioners rate highest and vendors mention least, because it is unglamorous and it is the difference between a tool people use and a tool people work around.
4. Scope can be derived. When the system changes, the product should be able to tell you which objectives that change bears on, computed from a diff rather than declared in a meeting. Without it, the safe default is to reassess everything, and that default is why programs avoid changing systems.
5. Findings carry their lineage. A not-satisfied determination becomes a POA&M item that retains the link back through the observation to the evidence that produced it. When a reviewer asks where an item came from, the answer should be a link.
6. Reports are generated, not assembled. The SAR comes out of the assessment record. If your assessors work in the tool and somebody then writes the SAR separately, you have two records and they will disagree.
The distinction that has to be in the data model
Observation, finding and risk are three different things, and a product that conflates them cannot answer the most common assessor question there is.
- An observation is what was seen: a scan returned this, a setting has this value, this document says this. It carries a method and a date and no judgment.
- A finding is a determination about an objective, made by a named assessor on the basis of observations.
- A risk is the consequence of a negative finding, with a status and a remediation path, and it is what becomes a POA&M item.
Ask a vendor to show you all three as distinct objects. Where everything is a “finding” with a status field, on what basis did you decide that? tends to be answered in a comment rather than in the data.
The methods should be visible
SP 800-53A gives three assessment methods — EXAMINE, INTERVIEW, TEST — and a product should record which was used per observation.
This is not bookkeeping. Automation naturally pushes effort toward EXAMINE, because reading a configuration is easy and interviewing a person is not. A program that cannot see its method distribution will drift into assessing documentation, and a system can be beautifully documented and wrong.
If the product does not record method, you cannot notice the drift.
Questions for the demo
Ask for these in the live product. Each takes a couple of minutes and none can be rehearsed into something they are not.
- Reassess after a change. Change something in the architecture. Show me which objectives are now in scope and why.
- Upload the same scan twice. Show me that the findings I dismissed last time are still dismissed.
- Trace a POA&M item backwards. From the item to the finding to the observation to the evidence and its source.
- Show a stale determination. Not “assessed 400 days ago” — “assessed against version 24, the system is on 27, and here is what changed between them that touches this objective.”
- Generate the SAR. Then change a determination and regenerate it. Confirm the change appears and nothing else was lost.
Things that look like features and are not
- Compliance percentages. “87% compliant” is not an RMF outcome. Objectives are satisfied or they are not. A percentage averages a missing contingency plan against a missing comment.
- Automated determinations. If the product marks controls satisfied on its own, ask who the assessor of record is. The answer should never be the software.
- Framework breadth on its own. A long framework list describes breadth, which may be exactly what you need. It does not describe depth in the one you are assessing against — verify that separately.
- Evidence upload. Every product has file upload. The question is whether evidence is bound to the objectives and determinations it supports, or filed near them.
The organizational problem software will not fix
Worth saying before anyone spends money. Assessment software fixes searching, format-wrangling, re-adjudication and reconstruction. It does not fix an unclear boundary, an absent categorization rationale, engineering that will not tell you what shipped, or a program that has not decided who owns the ATO.
If those are the real problems, better tooling produces a well-instrumented version of the same mess, on a subscription. Fix the boundary first.