In regards specifically to mode collapse in the DINO head, I am not asking exactly why the teacher and student eventually gravitate towards producing the same output as thats one way to make the loss minimal even if its not constructively building proper feature representations in the output probability distribution vectors. What I am asking about is how the student would even get to such a state in the first place. If the teacher and student are initialised together randomly before training, what would cause the student to eventually start producing the same output vectors for every input before the teacher gravitates towards the same behaviour creating a positive reinforcement loop? The best explanation I can come up with is if multiple inputs happen to produce bias towards certain dimensions, then the aggressive optimisation caused by the high peaking from a lower temperature in the teacher softmax when trying to match student and teacher predictions would push the student model too hard in a direction which would teacht it to stay consistently high in certain dimensions, which overtime would cause the teacher to follow the same pattern through the EMA, and that the centering solves this issue by tempering the distribution of prototype scores. Please let me know how wrong I am and where I went wrong.

Full article content could not be extracted automatically. Read the original below.