A public test of AI feedback classification

Candidate Experience
Feedback Benchmark

We asked AI models to classify 60 fictional comments from job candidates. For each comment, they judged sentiment, whether follow-up was needed, whether it raised a serious concern, and whether it could be a testimonial. Compare the saved results and read individual comments below.

01 / SET60fictional candidate comments
02 / DECISIONS04decisions for each comment
03 / PROMPTS03up to three prompt versions

Start here

What the saved results say.

These numbers describe agreement with provisional answers on fictional comments. They are a starting point for inspecting the runs, not a ranking for hiring use.

Reference review: DEV-006 has a proposed serious-concern correction; DEV-013 and DEV-030 still need human adjudication. Applying the single correction would move Jev from 54 to 55 all-four matches. These charts retain the original reference labels. Read the review and separate rescore.

Prompt changes21 / 38

Adding decision-tree instructions reduced agreement in 21 audited setups.

P2 versus P1: four improved and 13 tied among 38 audited hosted and subscription setups. These comparisons use the first recorded pass.

See paired changes ↓
Jev decision profile25 / 25

Jev matched every reference marked serious concern: yes.

Its six all-four misses span other issues, including a missed follow-up request and an off-topic review.

Read the six cases ↓
Cost boundary

Higher effort cost more without adding matches in two saved pairs.

Gemini 3.1 Pro low matched 56 / 60 for $0.063458; high matched 55 / 60 for $0.256826. The chart shows complete P0 runs with fully observed charges.

Inspect recorded charges ↓

01 / Paired prompt comparisons

Loading paired prompt results…

Matched setups are compared against the same 60 provisional answers.

Did the added instructions help the same setups?

Loading paired comparisons…

Prompt changes are associations within saved runs. The categories keep controlled pairs, hosted observations and Jev's native Choice changes separate. Read the analysis and comparison limits.

Repeatability / Compare repeated answers

The same prompt can produce different answers.

Compare repeated scores before interpreting a small prompt improvement.

Loading repeat comparisons…

Choose a model configuration to compare its three planned passes. Completion is shown separately for each configuration. The accepted CLI patch change and hidden serving details prevent attributing every difference solely to randomness. Three passes are descriptive evidence, not a precise estimate of future variation. Read the analysis and source hashes.

02 / Decision errors

Loading Jev decision profile…

The four decisions can fail in different ways.

Where did Jev disagree, and did other runs miss the same comments?

Loading decision counts…

03 / Observed spending

Loading observed costs…

Compare only charges attached to saved runs.

What did these scored runs cost?

Loading cost records…

04 / Comments worth a second look

Loading difficult comments…

A repeated miss may reveal an ambiguous comment or reference answer.

Which comments should a person review first?

Loading review counts…

Continue into the records

Compare all saved runs.

These rows show all-four agreement with the provisional reference. Select a run to inspect its setup, costs and comments.

Compare saved runs

All four answers match the reference

Each row is one saved run. Models may appear more than once with different settings or attempts. Counts are out of 60 fictional comments.

Scores reflect saved outputs, not a controlled model ranking. Read each run’s setup and source report before comparing.

How to read the scores

A valid response may still disagree with the reference.

All scores are counts out of 60 comments. A response can follow the required format yet disagree with the saved reference answer. Missing and invalid responses do not earn agreement points.

01

Valid responses / 60

Responses in the required format that could be scored.

02

All four match / 60

Comments where all four decisions matched the saved reference answers.

03

Each decision / 60

You can also see agreement for each decision separately.

Featured run · TypeSafe API

Jev matched all four answers for 54 of 60 comments.

This is one saved run of TypeSafe Jev 1.13 using the base task. Those matches are against provisional reference answers, not a measure of hiring accuracy. Its API price is an estimate; the provider charge is unavailable.

ALL FOUR MATCH—/ 60

Loading the saved Jev run…

Model comparison

Compare Jev with another saved run.

Choose a run to see its agreement counts beside Jev. Prompt, settings and execution method may differ, so use the linked run notes before drawing conclusions.

Read the selected run's comments and source report ↓

Compare prompt versions

Compare results across prompt versions.

Choose a model and setup to compare the saved results for up to three prompt versions. A blank card means there is no public result for that version. Other differences between runs may also affect the scores.

P0

The task, scoring rules and required response format.

P1

The base prompt plus a classifier role and more explicit instructions.

P2

The base prompt plus step-by-step rules for the four decisions.

Jev’s P1 and P2 change its native Choice questions. Other models’ P1 and P2 add chat instructions.

Loading model setups…

Why some prompt versions are missing
Browse every saved run Search, filter and inspect all results

Saved runs

Compare all saved runs.

A model can appear more than once when it has different prompts, settings or attempts.

MODEL / RUNPROMPTVALID / 60ALL FOUR MATCH / 60USAGE / COST

Loading saved outcomes…

Resource use

Tokens, cost and measured time.

These figures describe the selected run. Missing provider usage and charges are shown as unavailable. Request duration is a client measurement, while inference time requires a separate server measurement.

Individual run

See the results for one run.

Select a run in the model comparison or full list to read its comments, answers and source report. Technical request timing and usage details follow the score summary.

Loading run evidence…

Method and limits

What this test can tell you.

The comments are fictional and the reference answers are provisional. These scores show agreement within this development test.

How the test was scored

Prompt versions: P0 is the base task; P1 and P2 add different instructions. Every score is out of 60, even when a run saved fewer responses. The same AI assistant drafted and reviewed the reference answers. No independent person checked them.

Time: Server inference time appears only when a provider reports it. The technical run detail shows recorded client request durations, their sum and the number of timed requests. Some requests classified one comment; others classified a batch. The sum does not show how long concurrent work took from start to finish. Local times reflect the test machine and workflow.

Cost: Reported API charges, price estimates and possible charges without a known bill are labeled separately. Subscription fees are not split across runs. An API price estimate is not an observed charge.

Common questions

Does a valid response mean the answer matched?

No. Valid means the response followed the format needed for scoring. Agreement is counted separately against the saved reference answers.

Do score changes show that one prompt is better?

You can compare saved outcomes, but differences in where and when runs were made, missing records, and other settings make it hard to attribute a score change to the prompt alone. Five earlier local Qwen comparisons lack logs for the base runs. Their comparison reports show the saved changes and explain the missing evidence.

Why can a complete run have invalid responses?

Complete means a response was saved for each of the 60 comments. Some saved responses did not follow the required format, so they could not be scored.

Why is a price missing?

Some services do not report a charge for each run. An unavailable charge does not mean the run was free. Estimates and upper bounds are labeled separately.

Can I use these scores to choose a hiring tool?

No. This small fictional set and its provisional answers do not measure performance with real candidates or the effect of using a model in hiring.