More Data Helps - with Conditions
Quantity reduces noise, not systematic error
The law of large numbers says a sample average tends toward the population average as independent observations accumulate under stable conditions. This is why estimates from 100,000 representative events are usually steadier than estimates from 20. It does not say every large dataset is truthful.
Dependence reduces effective information: one million nearly identical frames from the same five-minute video are not equivalent to one million independent scenes. Distribution change also breaks the stable-population assumption. And systematic collection bias does not wash out with volume.
Analogy: Repeating a miscalibrated scale measurement a million times tells you its wrong answer very precisely. More data narrows random wobble but cannot repair the scale.
small + representative -> uncertain but relevant
large + representative -> more precise and relevant
large + biased -> precisely misleading
Note: Ask three separate questions: how many observations exist, how independent they are, and how well their collection matches the deployment population.