Skerpa test bench

Ingestion on skerpa's high tier: run the pipeline as it ships, with Jev doing the dedup, or fully Jev — and compare them on entities, quality, speed and cost.

Before you start

1 · Your file

pdf, docx, pptx, md, txt, html, xlsx. This is the one that gets timed.
The dedup judge only runs when the project already holds entities. This document is ingested and merged first, untimed, so the judge has something to compare against. Leave empty and the report will say the judge was not exercised.

2 · What to run

Tick one or more. Ticking option 1 as well makes it the yardstick for the others in this same run; leave it unticked and they use your last saved option 1 for this file.

3 · Options

TypeSafe's model is English-first. For other languages, read the agreement rate before the speed.
The yardstick for the "off topic" judgment. Left empty it is inferred from the transcript's title or its middle — never its opening, because calls start with greetings and everything substantive would then look off topic.
LLM latency varies run to run. 3 gives medians; each repeat costs time and money.
A benchmark is already running.

Running

starting…

Decisions, live

entity
chat LLM
Jev
agree

Result

Earlier runs