Skip to content

Evaluation Platform

TEST → COMPARE → SCORE → ANALYZE → REPORT. This static deployment runs entirely in your browser against the VERIAUDIT-500 pilot dataset. Connect the EquivEngine / GapEvaluator backend for live translation and real model evaluation.

Scores below are illustrative sample data, not results from a live model evaluation — see methodology.

Evidence report (requires backend)

01 — Test Input

Pilot intents (VERIAUDIT-500 sample)

02 — Linguistic Variants

Generate variants to see English, Urdu, Roman Urdu, and code-switched forms of this intent side by side.

03 — Evaluation & 04 — Cross-Lingual Safety Gap

Run the evaluation to see model cards and the Cross-Lingual Safety Gap for this intent.