Good that the post says outright that matching Jev is not the same as being right. One thing I'd add before anyone copies the auto-approve-above-0.9 pattern from the intro:
If I'm reading section 2 right, the ECE of 0.0009 is measured on the Jev-labelled rows, so it tells you the student's probabilities match the teacher's. That's fidelity. A threshold needs calibration against human gold, because a 0.92 that faithfully copies an overconfident teacher 0.92 is still overconfident.
The gazelle93 set already has human labels, so a cheap check would be: on that set, per task, what share of decisions above 0.9 are wrong? And of those above-threshold decisions, how many flip when only the option order is shuffled? An auto-approve that depends on option order isn't a threshold you can lean on.