A high jump competition does not ask everybody the same question.
The bar starts low enough that most of the field will clear it, and after each round it goes up. Jumpers who clear it stay in. Jumpers who fail three times at one height are finished, and the height they last cleared is their result. Nobody jumps a hundred times; the competition finds each athlete's level in a handful of attempts, because every attempt is chosen in the light of the last one.
The efficiency of that is easy to miss. To measure everyone accurately with a fixed set of heights you would need dozens of them: low ones that tell you nothing about the leaders, high ones that tell you nothing about anybody else. An attempt only carries information when it might go either way.
Modern computer-based tests are built on the same principle. A question far below a candidate's level, or far above it, uses up time and reveals almost nothing. So the test watches the early answers and chooses what comes next — harder if the early work was solid, easier if it was not.
Which means a hard second half is not bad news. It is the sign that the bar has been raised, and it is raised for exactly one reason.