Skip to main content

Model leaderboard

How AI models held up against the Secuvon corpus.

Every result here is a full-corpus scan Secuvon ran itself. A model is ranked only when an LLM judge other than the model under test graded the run and the run's provenance record is on file. Each model appears once.

No verified results yet

A model is listed here only after a full-corpus scan Secuvon ran itself, graded by a judge that is not the model under test, with its provenance record on file. No run meets that bar right now, so nothing is ranked.

How to read these results

  • The security score runs from 0 to 100 and higher is safer. It is 100 minus the severity-weighted share of gradeable tests that failed.
  • Tiers: S from 95, A from 85, B from 70, C from 55, D below 55.
  • A ranked run must have a gradeable result for every test. Execution errors and skipped tests cannot count as passes or failures.
  • Each model was called through its provider's official API with the provider's default safety settings and no Secuvon runtime protection in front.
  • A score is a snapshot of one run. Model sampling and the judge both vary, so a new run can score differently.

How Secuvon selects, runs and grades tests is described in the assessment methodology.

Runs we keep but do not rank

These runs stay in the benchmark data as dated history. They are not comparable with the ranked results, so they carry no rank.

  • OpenAI GPT-4o mini (OpenAI),

    No publishable score. Corpus secuvon-corpus-2.0, 3,491 tests. Grading method not recorded.

    Not ranked: the run was judged by the model under test and has no committed provenance record; its numeric result is withheld.