It is common practice when training deep learning algorithms to augment the training data. For example, in computer vision an image might be flipped, rotated, cropped etc., adding new training samples and effectively increasing the training set size. In general, models trained with data augmentation perform better than models that didn't include data augmentation during training. The better performance is commonly attributed to the fact that the model was trained on a more diverse set of examples, effectively making it more robust. However, I fail to see how this can be explained on a theoretical basis. In machine learning, we typically assume that both the training and test set come from the same distribution $p(x, y)$ i.e $p(x, y) = p_\text{train}(x, y) = p_\text{test}(x, y)$ . By augmenting the training set, it is not guaranteed that the new distribution $p_\text{train}'(x, y)$ will match the original $p_\text{train}(x, y)$ and as such, $p(x, y)$ . The only way that this can happen is if the applied transformations are consistent with $p(x, y)$ . Is the following correct? Data augmentation works because we artificially increase the dataset with examples that "respect" the data generating distribution. An example Suppose we train a neural network for classifying digits. If all images in the dataset are in the same scale, crystal-clear and with a fixed orientation, is there any reason to augment by cropping, blurring and randomly rotating the training images? If the above r…

Full article content could not be extracted automatically. Read the original below.