The factory tests a random sample of a hundred items from each week's production and reports the share that are faulty. Last month the figure was three percent. This month it is two.
The manager wrote to the team congratulating them on a third fewer faults. The quality engineer asked him to look at the second column of the report first.
Each figure came with a margin of error of one and a half percent. So last month's true rate was somewhere between one and a half and four and a half percent, and this month's is somewhere between one half and three and a half. Those two ranges overlap across most of their length. It is entirely possible that nothing at all changed and the sample simply came out differently.
The engineer was not saying the improvement was unreal. She was saying the numbers do not yet show it, which is a different statement and the only honest one available.
The team had in fact worked hard that month, and the manager was right to notice. Being right and having evidence are separate things.
Her proposal was to test four hundred items a week instead of a hundred. A larger sample narrows both ranges, and if the improvement is real it will show as a gap between them within two months.
She added one warning. Narrower ranges only help if the items are still chosen at random from the whole week. Testing four hundred items all from Monday morning would give a very confident answer to the wrong question.