jev / iOS
PAIRED SIMULATOR RUNS
SAME APP. SAME GOAL. SAME LOCAL RUNNER.

The same task.
Two decision engines.

Jev · paired median full run
Baseline · paired median full run
Baseline time ÷ Jev time
Median of verified pair ratios
Verified attempts · Jev / baseline

Watch the difference

0.0 / 0.0 s

Each recording was captured separately. Playback starts at runner-zero, stays at 1×, and pauses to preserve alignment if decoding stalls. The shorter video holds its final frame. The clock shows recording time, including any captured tail. Simulator captures can omit a trailing static wait and end before the measured task; measured task duration appears above.

Every attempt is counted

Failures remain in the table and success rate.
PAIR / ORDERENGINE / OUTCOMEFULL RUNMEDIAN REQUESTACTIONS / CALLSREPORTED INPUT / OUTPUTEST. COST

What is held constant

The goal, expected labels, observed accessibility state, action menu, local freshness checks, input execution, and final verification are shared. Each attempt restarts the app; stored application data is preserved. Starting labels and a stable accessibility snapshot are checked. Any initial-state mismatch excludes the cohort from speed comparison.

Jev chooses from categorical answers. The baseline generates a structured decision using the same offered operations and targets. Their transports and output formats differ.

How to read these numbers

Full-run times include model calls and local device work. Request latency is measured at the client and includes network and provider processing. Its median is calculated over available call timings, including failed attempts; it is a separate measure from full-run speed.

Only pairs where both runs verified and shared the same initial state contribute speed ratios. This small local experiment measures these models, this scenario, and this machine. It does not establish a general winner. Costs are estimates from recorded usage and provider pricing, not billing readback.