Know what the evidence means before you rely on it.
These answers define the launch scope. They do not replace the limitations shown inside each versioned report.
Know what the evidence means
What exactly does Jev Benchmark Lab test?
It measures how one Jev model version performs on the labeled examples and Choice or Noul question you provide. It does not predict every future input.
How many labeled rows do I need?
You can test up to 50 rows in the free preview and up to 2,000 in a paid report. Fewer than 100 rows receives an INSUFFICIENT_SAMPLE warning.
Do I need a Jev API key?
No. The hosted run uses the product's server-side route. Your browser never receives that credential.
Where does my data go, and how long is it kept?
Your original file is parsed in your browser and is not stored as an uploaded object. Paid report access lasts up to seven days and can be deleted sooner.
Is this an official TypeSafe product?
No. Jev Benchmark Lab is an independent third-party tool. It is not owned, sponsored, approved, or endorsed by TypeSafe AI or OpenRouter.
Does a threshold guarantee production safety?
No. A threshold summarizes observed tradeoffs on one dataset, question, and model version. Validate the full workflow and keep a fallback.
What do I receive for $9, and what if the run fails?
You receive one report for up to 2,000 rows. If the final result is PARTIAL or FAILED, choose a same-configuration rerun or full refund.