Publications
2020
Issue 59

MiS Preprint Repository

We have decided to discontinue the publication of preprints on our preprint server end of 2024. The publication culture within mathematics has changed so much due to the rise of repositories such as ArXiV (www.arxiv.org) that we are encouraging all institute members to make their preprints available there. An institute's repository in its previous form is, therefore, unnecessary. The preprints published to date will remain available here, but we will not add any new preprints here.

MiS Preprint

59/2020

On the Locality of the Natural Gradient for Deep Learning

Nihat Ay

Abstract

We study the natural gradient method for learning in deep Bayesian networks, including neural networks. There are two natural geometries associated with such learning systems consisting of visible and hidden units. One geometry is related to the full system, the other one to the visible sub-system. These two geometries imply different natural gradients. In a first step, we demonstrate a great simplification of the natural gradient with respect to the first geometry, due to locality properties of the Fisher information matrix. This simplification does not directly translate to a corresponding simplification with respect to the second geometry. We develop the theory for studying the relation between the two versions of the natural gradient and outline a method for the simplification of the natural gradient with respect to the second geometry based on the first one. This method suggests to incorporate a recognition model as an auxiliary model for the efficient application of the natural gradient method in deep networks.

Contact the author per mail Download full preprint 2 MB

Received:: 21.05.20

Published:: 22.05.20

Keywords:: natural gradient, Fisher-Rao metric, Deep learning, Helmholtz machines, wake-sleep algorithm

Related publications

inJournal

2023 Journal Open Access

Nihat Ay

On the locality of the natural gradient for learning in deep Bayesian networks

In: Information geometry, 6 (2023) 1, pp. 1-49

BibTex DOI: 10.1007/s41884-020-00038-y ArXiv: 2005.10791