Skip to content

Public validation method · 2026 edition

A benchmark with the context left in.

Measure face coverage, false masks, track continuity and human review time on footage that resembles your real work. No unqualified “AI accuracy” percentage.

Six-scenario set

Test the conditions that break clean demos.

Use clips you are authorised to test. A useful internal set can be 60 total face appearances across six ten-minute scenarios, but teams should adjust the size to their risk and volume.

01

Frontal daylight

Baseline faces at near, medium and far scale.

02

Profiles & occlusion

Side views, partial faces, hats, masks and temporary obstruction.

03

Crowd & frame edges

Multiple simultaneous subjects entering, exiting and crossing.

04

Low light

Noise, uneven exposure, backlight and small faces.

05

Reflections & screens

Mirrors, vehicle glass and faces inside embedded displays.

06

Fast motion & cuts

Camera pans, subject motion, shot changes and re-entry.

What to record

Face coverage

Count expected visible face appearances and misses using a human-labelled reference.

False masks

Count stable censorship regions that cover non-face objects.

Track continuity

Count material breaks where protection leaves or loses an intended face.

Review effort

Measure active operator minutes required to produce an approved export.

Test context

Record software version, mode, hardware, codec, resolution, frame rate and scenario.

Interactive scorecard

Turn observations into a comparable run.

Enter the totals from one completed test. Save the context beside the result so a later version or workstation can be compared fairly.

Enter one completed test run

Run summary

Observed coverage

93.3%

Corrections needed

9

Review minutes per scenario

3.0 min

These are your entered observations, not CensorFlow-wide accuracy claims. Record camera, codec, resolution, hardware, version and selected privacy mode with every run.

Current evidence status

Protocol published; product-wide result not yet claimed.

CensorFlow is not presenting a universal precision or recall number on this page. A credible public result requires a frozen dataset, human-labelled reference, documented hardware and reproducible release. Until that run is completed, customer-specific validation is the honest decision tool.