Evaluation Platform
TEST → COMPARE → SCORE → ANALYZE → REPORT. This static deployment runs entirely in your browser against the VERIAUDIT-500 pilot dataset. Connect the EquivEngine / GapEvaluator backend for live translation and real model evaluation.
Scores below are illustrative sample data, not results from a live model evaluation — see methodology.
01 — Test Input
Pilot intents (VERIAUDIT-500 sample)
02 — Linguistic Variants
Generate variants to see English, Urdu, Roman Urdu, and code-switched forms of this intent side by side.
03 — Evaluation & 04 — Cross-Lingual Safety Gap
Run the evaluation to see model cards and the Cross-Lingual Safety Gap for this intent.