Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

https://www.cs.toronto.edu/~hinton/absps/JMLRdropout.pdf

Looks pretty close, yeah.

So this is basically dividing the network into two groups, and give each group a second weight: one for those that are on, zero for those that are off.

Has anyone tried more variants of this? Using random bundling + various levels of amplification of such bundles, so that backpropagation affects various networks differently?



There is also this: Mixture of experts. A massive neural net from Google, with multiple submodules that are activated depending on input.

https://medium.com/@thoszymkowiak/google-brains-new-super-fa...




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: