Cisco AI Defense skill scanner: evaluation results
What changes when the LLM judge and the System One model are switched on.
NVIDIA SkillSpector vs Cisco skill-scanner
Both scanners run over one population of skill directories, byte-identical on each side, with every analyzer each tool ships enabled and neither tuned on these labels. Three arms: static analysis only, then both tools on Gemma 4 26B, then both on Claude Haiku 4.5.
The comparison is possible because SkillSpector emits the same LOW/MEDIUM/HIGH/CRITICAL levels this scanner does, so a threshold means the same thing on both sides. Package-level decisions do not translate, and are reported separately rather than normalised into a shared verdict neither tool emits.
Only records where both tools returned a usable result are scored. A crashed scan, an unparsable payload and a run with analyzers switched off all present as zero findings, so scoring them as clean would credit specificity that was never earned.
All 27 analyzers enabled; built-in YARA; live OSV; baselines off; resource ceilings raised; --transitive off (corpus safety terms)
Cisco skill-scanner
4f8bec3
Shipped core rule pack; judge + meta-judge; escalate-only verdict repair; policy suppressions empty; VirusTotal and AI Defense off (no credential)
Four decision lenses
This scanner emits findings plus a package verdict. SkillSpector emits findings plus a 0–100 risk score, banded into a recommendation, and gates an install above 50. Its score accumulates with diminishing returns per rule, so a single CRITICAL finding at 0.95 confidence scores 47 and returns CAUTION, which does not gate.
A gate-to-gate comparison therefore measures two different thresholds rather than two engines. Detection at a fixed severity is the primary lens for that reason; each tool's own gate is reported alongside it, and a sweep across every threshold is given below so the operating point is visible rather than assumed.
Each lens carries a paired difference: both tools are resampled together on the same records, so shared corpus difficulty cancels rather than widening both intervals, and an interval excluding zero means the two tools are separable on that lens. Each corpus also reports complementarity, the share of records exactly one tool flags, which is the interesting population because neither tool is a superset of the other.
Same model: Gemma 4 26B
Both tools on bedrock-mantle/google.gemma-4-26b-a4b.
MaliciousSkillBench, source-disjoint split
1367 records scored on both tools · 528 benign, 839 malicious
Cisco skill-scanner (Gemma 4 26B) matches the HIGH row exactly wherever that happens below. It is a property of the default policy, which gates at HIGH, not a repeated table: the decision was read from each tool's emitted verdict field on every record.
Wall clock was recorded per subprocess, which covers SkillSpector but not this scanner's in-process path, and neither tool reported token usage through the harness. The uninstrumented cells are marked rather than shown as zero, and no cost comparison is drawn from this table. Judge token cost was measured separately and is reported in the token-cost section.
Integrity
Tool
Rows
Usable
Errors
Capability degraded
Reported partial
Cisco skill-scanner (Gemma 4 26B)
1384
1384
0
0
0
NVIDIA SkillSpector (Gemma 4 26B)
1384
1367
0
17
1043
MaliciousSkillBench, balanced subset
793 records scored on both tools · 397 benign, 396 malicious
794 records scored on both tools · 397 benign, 397 malicious
Any finding at MEDIUM or above
Tool
Recall
Precision
F1
False-positive rate
Cisco skill-scanner (core rules)
33.00%
94.24%
48.88%
2.02%
Cisco skill-scanner (all rule packs)
94.46%
52.97%
67.87%
83.88%
NVIDIA SkillSpector
48.11%
62.62%
54.42%
28.72%
Any finding at HIGH or above
Tool
Recall
Precision
F1
False-positive rate
Cisco skill-scanner (core rules)
31.23%
95.38%
47.06%
1.51%
Cisco skill-scanner (all rule packs)
91.44%
56.45%
69.81%
70.53%
NVIDIA SkillSpector
30.48%
75.62%
43.45%
9.82%
Each tool's own shipped install gate
Tool
Recall
Precision
F1
False-positive rate
Cisco skill-scanner (core rules)
31.23%
95.38%
47.06%
1.51%
Cisco skill-scanner (all rule packs)
91.44%
56.45%
69.81%
70.53%
NVIDIA SkillSpector
8.56%
91.89%
15.67%
0.76%
Block on any finding at all
Tool
Recall
Precision
F1
False-positive rate
Cisco skill-scanner (core rules)
81.86%
47.86%
60.41%
89.17%
Cisco skill-scanner (all rule packs)
99.75%
50.06%
66.67%
99.50%
NVIDIA SkillSpector
50.13%
63.17%
55.90%
29.22%
Cost of a scan
Tool
Mean seconds per skill
Input tokens
Output tokens
Cisco skill-scanner (core rules)
0.017
not instrumented
not instrumented
Cisco skill-scanner (all rule packs)
0.280
not instrumented
not instrumented
NVIDIA SkillSpector
2.399
not instrumented
not instrumented
Integrity
Tool
Rows
Usable
Errors
Capability degraded
Reported partial
Cisco skill-scanner (core rules)
800
800
0
0
18
Cisco skill-scanner (all rule packs)
800
800
0
0
18
NVIDIA SkillSpector
800
794
0
6
674
OpenSkillRisk
335 records scored on both tools · 182 contextually_risky, 153 obviously_malicious
Any finding at MEDIUM or above
Tool
Recall
Precision
F1
False-positive rate
Cisco skill-scanner (core rules)
25.07%
no harmless class
no harmless class
no harmless class
Cisco skill-scanner (all rule packs)
96.42%
no harmless class
no harmless class
no harmless class
NVIDIA SkillSpector
77.91%
no harmless class
no harmless class
no harmless class
Any finding at HIGH or above
Tool
Recall
Precision
F1
False-positive rate
Cisco skill-scanner (core rules)
21.49%
no harmless class
no harmless class
no harmless class
Cisco skill-scanner (all rule packs)
91.64%
no harmless class
no harmless class
no harmless class
NVIDIA SkillSpector
62.69%
no harmless class
no harmless class
no harmless class
Each tool's own shipped install gate
Tool
Recall
Precision
F1
False-positive rate
Cisco skill-scanner (core rules)
21.49%
no harmless class
no harmless class
no harmless class
Cisco skill-scanner (all rule packs)
91.64%
no harmless class
no harmless class
no harmless class
NVIDIA SkillSpector
33.43%
no harmless class
no harmless class
no harmless class
Block on any finding at all
Tool
Recall
Precision
F1
False-positive rate
Cisco skill-scanner (core rules)
97.61%
no harmless class
no harmless class
no harmless class
Cisco skill-scanner (all rule packs)
100.00%
no harmless class
no harmless class
no harmless class
NVIDIA SkillSpector
78.21%
no harmless class
no harmless class
no harmless class
Cost of a scan
Tool
Mean seconds per skill
Input tokens
Output tokens
Cisco skill-scanner (core rules)
0.018
not instrumented
not instrumented
Cisco skill-scanner (all rule packs)
0.266
not instrumented
not instrumented
NVIDIA SkillSpector
2.457
not instrumented
not instrumented
Integrity
Tool
Rows
Usable
Errors
Capability degraded
Reported partial
Cisco skill-scanner (core rules)
374
374
0
0
5
Cisco skill-scanner (all rule packs)
374
365
0
9
14
NVIDIA SkillSpector
374
343
0
31
302
HarmfulSkillBench
197 records scored on both tools · 197 unlabelled
Any finding at MEDIUM or above
Tool
Recall
Precision
F1
False-positive rate
Cisco skill-scanner (core rules)
4.06%
not permitted
not permitted
not permitted
Cisco skill-scanner (all rule packs)
90.86%
not permitted
not permitted
not permitted
NVIDIA SkillSpector
56.85%
not permitted
not permitted
not permitted
Any finding at HIGH or above
Tool
Recall
Precision
F1
False-positive rate
Cisco skill-scanner (core rules)
3.55%
not permitted
not permitted
not permitted
Cisco skill-scanner (all rule packs)
83.76%
not permitted
not permitted
not permitted
NVIDIA SkillSpector
17.26%
not permitted
not permitted
not permitted
Each tool's own shipped install gate
Tool
Recall
Precision
F1
False-positive rate
Cisco skill-scanner (core rules)
3.55%
not permitted
not permitted
not permitted
Cisco skill-scanner (all rule packs)
83.76%
not permitted
not permitted
not permitted
NVIDIA SkillSpector
2.54%
not permitted
not permitted
not permitted
Block on any finding at all
Tool
Recall
Precision
F1
False-positive rate
Cisco skill-scanner (core rules)
96.45%
not permitted
not permitted
not permitted
Cisco skill-scanner (all rule packs)
100.00%
not permitted
not permitted
not permitted
NVIDIA SkillSpector
56.85%
not permitted
not permitted
not permitted
Cost of a scan
Tool
Mean seconds per skill
Input tokens
Output tokens
Cisco skill-scanner (core rules)
0.017
not instrumented
not instrumented
Cisco skill-scanner (all rule packs)
0.325
not instrumented
not instrumented
NVIDIA SkillSpector
2.400
not instrumented
not instrumented
Integrity
Tool
Rows
Usable
Errors
Capability degraded
Reported partial
Cisco skill-scanner (core rules)
200
200
0
0
19
Cisco skill-scanner (all rule packs)
200
200
0
0
19
NVIDIA SkillSpector
200
197
0
3
160
Limits on what these numbers support
Our rule packs were tuned against MaliciousSkillBench during development. SkillSpector has not seen it. The source-disjoint split and the corpora outside MSB are therefore the ones that carry weight, and figures on the balanced subset should be read as home ground. SkillSpector was tuned on a private 31,000-skill set that is not published, so contamination can be bounded in one direction only and is disclosed rather than measured in the other.
Enabling every rule pack this scanner ships is not its strongest configuration. On an 80/80 sample of the source-disjoint split, all packs raise recall from 8.8% to 73.8% and raise the benign flag rate from 7.5% to 92.5%. That is a triage setting rather than a gating one, so the shipped core profile is what appears above; the ATR pack is also excluded from source-disjoint claims by the corpus terms.
One SkillSpector capability is absent from these runs. Its --transitive mode follows references out of a skill, which on this corpus would mean fetching attacker-controlled URLs, and the corpus safety defaults forbid it. Its supply-chain recall would likely be higher with that enabled, so the figures here understate it on that axis.
MaliciousSkillBench stores each record as one SKILL.md. Analyzers that operate on bundled scripts or MCP manifests have nothing to read, on both sides, and report not_applicable rather than failing. Three of SkillSpector's 27 analyzers are inert on that corpus for this reason. OpenSkillRisk and HarmfulSkillBench carry the multi-file signal.
SkillSpector reports analysis as partial on most MSB records. Coverage is 100%, nothing is left uninspected, and the ledger exceptions are non-fatal: each SKILL.md references files the corpus does not ship, and the tool says so. That is a reporting capability this scanner does not have, and the partial rate is recorded in the integrity table rather than treated as a defect or used to exclude rows.
Finding counts are not compared directly. SkillSpector applies diminishing returns per rule and caps at three occurrences, so its issue count and this scanner's finding count measure different things. Distinct rules fired and records flagged are used instead.
Run notes
Both tools ran at shipped defaults with every capability enabled, and neither was tuned on these labels. The configuration is recorded in each run manifest so it can be checked rather than trusted.
Only records where both tools returned a usable result are scored. A crashed scan, an unparsable payload, or a run with analyzers switched off all present as “no findings”, and scoring those as clean would award free specificity.
Our Gemma 4 column reuses the per-record rows from the previously published judge run rather than re-scanning. The corpora are byte-identical, verified by digest over 400 records; only directory names changed and label sidecars moved out of scan scope.
Gemma 4 is reachable only through the Bedrock mantle route, which authenticates with SigV4. SkillSpector speaks OpenAI-compatible with a static bearer token, so a loopback signing proxy was used for that arm alone. Its Haiku arm used its native Bedrock provider.
HarmfulSkillBench figures on our side predate a fix that moved a label-bearing _meta.json out of the scanned directory, so that corpus is marked label-contaminated for us and should be read with that in mind.