Hello everyone,
I recently completed a sealed cross-model benchmark using 400 OWASP-style prompts: 337 attack prompts and 63 benign controls.
The same prompt set and strict-binary scoring method were applied across six frontier AI systems:
GPT-5.5, Gemini 3.1 Pro, Claude Opus 4.8, Llama 4 Maverick, Mistral Medium 3.5, and NVIDIA Nemotron-3 Ultra.
The most concerning result was not the ranking between models. It was the lack of a common security floor:
The main hypothesis behind the work is that conventional model safety largely judges semantic intent, while structural governance can independently evaluate authority, requested action, data boundaries, provenance, and decision boundaries.
The benchmark also produced structured conformity evidence for governed decisions, including the detected violation, UIA primitive and locus, enforcement rule, verdict, confidence, and rationale.
I am sharing this here to invite technical scrutiny—not to claim OWASP endorsement. I would especially welcome feedback on:
If you would like to read the full report, just tell me.
I would be glad to discuss the methodology, limitations, row-level evidence, and possibilities for independent replication.