Adding decision-tree instructions reduced agreement in 21 audited setups.
P2 versus P1: four improved and 13 tied among 38 audited hosted and subscription setups. These comparisons use the first recorded pass.
See paired changes ↓A public test of AI feedback classification
We asked AI models to classify 60 fictional comments from job candidates. For each comment, they judged sentiment, whether follow-up was needed, whether it raised a serious concern, and whether it could be a testimonial. Compare the saved results and read individual comments below.
Start here
These numbers describe agreement with provisional answers on fictional comments. They are a starting point for inspecting the runs, not a ranking for hiring use.
Reference review: DEV-006 has a proposed serious-concern correction; DEV-013 and DEV-030 still need human adjudication. Applying the single correction would move Jev from 54 to 55 all-four matches. These charts retain the original reference labels. Read the review and separate rescore.
P2 versus P1: four improved and 13 tied among 38 audited hosted and subscription setups. These comparisons use the first recorded pass.
See paired changes ↓Its six all-four misses span other issues, including a missed follow-up request and an off-topic review.
Read the six cases ↓Gemini 3.1 Pro low matched 56 / 60 for $0.063458; high matched 55 / 60 for $0.256826. The chart shows complete P0 runs with fully observed charges.
Inspect recorded charges ↓01 / Paired prompt comparisons
Matched setups are compared against the same 60 provisional answers.
Did the added instructions help the same setups?
Loading paired comparisons…
Prompt changes are associations within saved runs. The categories keep controlled pairs, hosted observations and Jev's native Choice changes separate. Read the analysis and comparison limits.
Repeatability / Compare repeated answers
Compare repeated scores before interpreting a small prompt improvement.
Loading repeat comparisons…
Choose a model configuration to compare its three planned passes. Completion is shown separately for each configuration. The accepted CLI patch change and hidden serving details prevent attributing every difference solely to randomness. Three passes are descriptive evidence, not a precise estimate of future variation. Read the analysis and source hashes.
02 / Decision errors
The four decisions can fail in different ways.
Where did Jev disagree, and did other runs miss the same comments?
Loading decision counts…
03 / Observed spending
Compare only charges attached to saved runs.
What did these scored runs cost?
Loading cost records…
04 / Comments worth a second look
A repeated miss may reveal an ambiguous comment or reference answer.
Which comments should a person review first?
Loading review counts…
Continue into the records
These rows show all-four agreement with the provisional reference. Select a run to inspect its setup, costs and comments.
Compare saved runs
Each row is one saved run. Models may appear more than once with different settings or attempts. Counts are out of 60 fictional comments.
Scores reflect saved outputs, not a controlled model ranking. Read each run’s setup and source report before comparing.
How to read the scores
All scores are counts out of 60 comments. A response can follow the required format yet disagree with the saved reference answer. Missing and invalid responses do not earn agreement points.
Responses in the required format that could be scored.
Comments where all four decisions matched the saved reference answers.
You can also see agreement for each decision separately.
Featured run · TypeSafe API
This is one saved run of TypeSafe Jev 1.13 using the base task. Those matches are against provisional reference answers, not a measure of hiring accuracy. Its API price is an estimate; the provider charge is unavailable.
Loading the saved Jev run…
Model comparison
Choose a run to see its agreement counts beside Jev. Prompt, settings and execution method may differ, so use the linked run notes before drawing conclusions.
Compare prompt versions
Choose a model and setup to compare the saved results for up to three prompt versions. A blank card means there is no public result for that version. Other differences between runs may also affect the scores.
The task, scoring rules and required response format.
The base prompt plus a classifier role and more explicit instructions.
The base prompt plus step-by-step rules for the four decisions.
Jev’s P1 and P2 change its native Choice questions. Other models’ P1 and P2 add chat instructions.
Saved runs
A model can appear more than once when it has different prompts, settings or attempts.
Loading saved outcomes…
Resource use
These figures describe the selected run. Missing provider usage and charges are shown as unavailable. Request duration is a client measurement, while inference time requires a separate server measurement.
Individual run
Select a run in the model comparison or full list to read its comments, answers and source report. Technical request timing and usage details follow the score summary.
Loading run evidence…
Method and limits
The comments are fictional and the reference answers are provisional. These scores show agreement within this development test.
Prompt versions: P0 is the base task; P1 and P2 add different instructions. Every score is out of 60, even when a run saved fewer responses. The same AI assistant drafted and reviewed the reference answers. No independent person checked them.
Time: Server inference time appears only when a provider reports it. The technical run detail shows recorded client request durations, their sum and the number of timed requests. Some requests classified one comment; others classified a batch. The sum does not show how long concurrent work took from start to finish. Local times reflect the test machine and workflow.
Cost: Reported API charges, price estimates and possible charges without a known bill are labeled separately. Subscription fees are not split across runs. An API price estimate is not an observed charge.
No. Valid means the response followed the format needed for scoring. Agreement is counted separately against the saved reference answers.
You can compare saved outcomes, but differences in where and when runs were made, missing records, and other settings make it hard to attribute a score change to the prompt alone. Five earlier local Qwen comparisons lack logs for the base runs. Their comparison reports show the saved changes and explain the missing evidence.
Complete means a response was saved for each of the 60 comments. Some saved responses did not follow the required format, so they could not be scored.
Some services do not report a charge for each run. An unavailable charge does not mean the run was free. Estimates and upper bounds are labeled separately.
No. This small fictional set and its provisional answers do not measure performance with real candidates or the effect of using a model in hiring.