Cisco AI Defense skill scanner: evaluation results

What changes when the LLM judge and the System One model are switched on.

Reproducing the published baseline

Before any judged number is believable the harness has to reproduce the published one. This is that check, run on the locked source-disjoint test partition after materializing the corpus from its pinned revision and verifying every manifest digest against the lock.

MetricPublishedReproducedMatch
F113.74%13.74%matches
Precision60.75%60.75%matches
Recall7.75%7.75%matches
False-positive rate7.71%7.71%matches
Samples1,3841,384matches
Malicious839839matches
Benign545545matches

Confusion counts

True positiveFalse positiveFalse negativeTrue negativeScan errors
65427745030

Corpus identity

The population digest is pinned in the dataset lock, so a shrunken or substituted population cannot pass unnoticed.

population_sha256  e77564f010fe55ee368af237fabc89b94ca308bd0b3a8c49f052716609f89399

Scanner version n/a, CEL mode off.

Deterministic rules: every metric, corpus by corpus

Scored on the any-intervention lens, so a finding raised for review counts as a catch. Errors are requests the stage could not complete; they are excluded rather than counted as clean.

CorpusPopulationRecallPrecisionF1FPRErrors
HarmfulSkillBench200 bad, 0 harmless4.00%no harmless classno harmless classno harmless class0
MaliciousSkillBench balanced subset400 bad, 400 harmless33.00%94.29%48.89%2.00%0
MaliciousSkillBench source-disjoint (the published test split)839 bad, 545 harmless7.99%59.29%14.08%8.44%0
OpenSkillRisk: contextually risky214 bad, 0 harmless32.71%no harmless classno harmless classno harmless class0
OpenSkillRisk: obviously malicious160 bad, 0 harmless30.00%no harmless classno harmless classno harmless class0

Every member is positive-risk, so there is no harmless class and precision is undefined. Its licence also forbids any false-positive-rate claim. Every member is positive-risk, so there is no harmless class and precision is undefined.

Re-measuring every previously published figure

Each row is a published deterministic figure re-measured from the pinned corpus revision. Reproducing these was the precondition for believing any judged number, because a harness that cannot reproduce the baseline cannot be trusted with a new arm.

CorpusPopulationPreviously publishedRe-measuredAgrees
MaliciousSkillBench source-disjoint1,384 packagesF1 13.74%F1 13.742%yes
MaliciousSkillBench source-disjoint1,384 packagesprecision 60.75%60.748%yes
MaliciousSkillBench source-disjoint1,384 packagesrecall 7.75%7.747%yes
MaliciousSkillBench source-disjoint1,384 packagesFPR 7.71%7.706%yes
NotInject339 benign text cases0.00% flag rate at MEDIUM+0.00%yes
InjecAgent1,054 canonical signals100.00% signal recall100.00%yes
In-Page Prompt Injection1,101 canonical groups99.36% signal recall99.36%yes
HarmfulSkillBench200 positive-risk skills3.50% MEDIUM+3.5% block rateyes
OpenSkillRisk374 positive-risk skills28.90% HIGH+, 35.74% MEDIUM+30.0% and 32.7% detectionyes

What is good: All nine figures reproduce. That is the only reason the judged numbers elsewhere on this site are worth reading: the same code path produced both.

Three we could not re-measure, and why

CorpusPopulationReason
Bundled skills snapshot111 installed skillsthe lock is stale against the currently installed applications, and refreshing it requires human review; forcing a refresh would change the corpus identity and break comparability
DataDog malicious packages5 selected positivesquarantine-only acquisition policy, and too small to move any figure
MaliciousAgentSkillsBench98,380 rowsmetadata only, with no skill content to scan