Public validation method · 2026 edition
A benchmark with the context left in.
Measure face coverage, false masks, track continuity and human review time on footage that resembles your real work. No unqualified “AI accuracy” percentage.
Six-scenario set
Test the conditions that break clean demos.
Use clips you are authorised to test. A useful internal set can be 60 total face appearances across six ten-minute scenarios, but teams should adjust the size to their risk and volume.
Frontal daylight
Baseline faces at near, medium and far scale.
Profiles & occlusion
Side views, partial faces, hats, masks and temporary obstruction.
Crowd & frame edges
Multiple simultaneous subjects entering, exiting and crossing.
Low light
Noise, uneven exposure, backlight and small faces.
Reflections & screens
Mirrors, vehicle glass and faces inside embedded displays.
Fast motion & cuts
Camera pans, subject motion, shot changes and re-entry.
What to record
Face coverage
Count expected visible face appearances and misses using a human-labelled reference.
False masks
Count stable censorship regions that cover non-face objects.
Track continuity
Count material breaks where protection leaves or loses an intended face.
Review effort
Measure active operator minutes required to produce an approved export.
Test context
Record software version, mode, hardware, codec, resolution, frame rate and scenario.
Interactive scorecard
Turn observations into a comparable run.
Enter the totals from one completed test. Save the context beside the result so a later version or workstation can be compared fairly.
Enter one completed test run
Run summary
Observed coverage
93.3%
Corrections needed
9
Review minutes per scenario
3.0 min
These are your entered observations, not CensorFlow-wide accuracy claims. Record camera, codec, resolution, hardware, version and selected privacy mode with every run.
Current evidence status
Protocol published; product-wide result not yet claimed.
CensorFlow is not presenting a universal precision or recall number on this page. A credible public result requires a frozen dataset, human-labelled reference, documented hardware and reproducible release. Until that run is completed, customer-specific validation is the honest decision tool.