
Midjourney / Entroparc illustration.
Notes
The Number Nobody Gave
When an AI estimate and a tutor review disagree, the product still has to decide which judgment governs—and preserve enough evidence to test that rule later.
01 The disagreement
Threom is an IELTS practice product. It can estimate a learner’s band after a practice session, and a tutor can later confirm or adjust that estimate. As long as they agree, the product has a number. When they disagree, it must decide which number will shape the learner’s progress record.
One of Threom’s regression tests makes the problem deliberately stark. The AI estimate is 5. The tutor adjusts it to 7. The test expects the learner-facing result to be 7 and explicitly checks that it is not 6.
This is a synthetic test case, not a story about a learner. Its value is that it removes the usual fog around the decision. The AI gave 5. The tutor gave 7. Nobody gave 6.
02 The number nobody gave
An average can look neutral because it sits politely between two disagreeing judgments. But the arithmetic does not choose the average; the product does. Showing 6 would mean adopting a rule that combines the two judgments and lets the compromise govern the learner’s record.
That rule might be defensible. It is not neutral merely because a calculator can perform it. A product that displays 6 still needs a reason why its aggregation rule deserves authority.
Threom uses a different rule. Before tutor review, the AI practice estimate is the selected value. If a tutor confirms it, the number stays the same. If the tutor adjusts it, the tutor’s value becomes the selected value. The implementation does not create a third judgment between them.
03 One value for the learner, two for the record
This choice affects more than a label or tooltip. The selected value supplies the point shown in the learner’s progress trend and contributes to the overall practice average. A single trajectory requires the product to decide which value governs it.
Choosing one value does not require the system to forget the other. Threom keeps the AI estimate and the tutor’s judgment separately, together with whether the estimate was confirmed or adjusted and who performed the review. The progress chart can show one value while still disclosing how it acquired that value.
The same pair of numbers therefore serves three different jobs. The learner-facing view needs a current value. The evidence record needs the disagreement and its provenance. Calibration needs many such pairs, qualified human ratings, and actual results. A tidy interface cannot substitute for that evidence.
04 Authority is not accuracy
This is why “the human overrides the AI” is an incomplete description of the rule. The tutor-adjusted value has operational authority in the interface. The AI estimate remains part of the evidence. Neither fact proves that the tutor was correct.
A tutor can be wrong. Two tutors can disagree. Threom’s calibration work remains unfinished, and this rule does not establish that an adjusted value is more accurate, let alone an official score. It answers a narrower question: after an identified tutor changes the assessment, which recorded judgment should govern the learner-facing view?
An average is still a fair alternative if a product can defend it. It could display 6 while preserving 5 and 7 for later evaluation. The harder question is what evidence makes 6 the right value to act on. Without a defensible weighting rule, the average hides a product judgment inside a mathematical operation.
The distinction also appears when the tutor agrees with the AI. If both values are 6.5, the number does not change after review, but its standing does. It moves from an unreviewed practice estimate to an estimate an identified tutor confirmed. The value and the standing of the value are not the same thing.
05 What makes a judgment legitimate?
So what makes an AI judgment legitimate: the number, the evidence behind it, or the person attached to the review?
For immediate use, Threom’s answer is an explicit authority rule with visible provenance. For empirical confidence, that is not enough. The preserved evidence must still be tested through calibration. The product can decide whose judgment governs the interface without pretending that this decision has settled which judgment was true.
The point is not that a human should always defeat a machine, or that an average is always wrong. It is that the product should name the authority rule it chose and preserve the evidence needed to test it. In Threom’s synthetic case, one value governs the learner-facing record, while both judgments remain available for later comparison.
Threom provides practice feedback, not official IELTS scores. IELTS is jointly owned by the British Council, IDP IELTS, and Cambridge University Press & Assessment. Threom is not affiliated with or endorsed by IELTS.
