In the early 2010s, deep learning ran on a simple faith: deeper networks, with more layers, should be more powerful. Reality disagreed. Past a certain depth, adding layers made networks harder to train and, strangely, worse even on the data they were trained on, a failure researchers called degradation. In 2015 a team at Microsoft Research published "Deep Residual Learning for Image Recognition" and fixed it with an idea so simple it fits in a sentence: let each layer learn the difference it makes, not the whole answer.

Your free preview ends here
Keep reading with a free account
Sign up and redeem two free articles per month. No credit card required.
Already a member? Log in
