Topology reduction in deep convolutional feature extraction networks

Authors

Thomas Wiatowski, Philipp Grohs, and Helmut Bölcskei

Reference

Proc. of SPIE (Wavelets and Sparsity XVII), San Diego, CA, USA, Vol. 10394, pp. 1039418:1-1039418:12, Aug. 2017, (invited paper).

[BibTeX, LaTeX, and HTML Reference]

Abstract

Deep convolutional neural networks (CNNs) used in practice employ potentially hundreds of layers and 10,000s of nodes. Such network sizes entail significant computational complexity due to the large number of convolutions that need to be carried out; in addition, a large number of parameters needs to be learned and stored. Very deep and wide CNNs may therefore not be well suited to applications operating under severe resource constraints as is the case, e.g., in low-power embedded and mobile platforms. This paper aims at understanding the impact of CNN topology, specifically depth and width, on the network’s feature extraction capabilities. We address this question for the class of scattering networks that employ either Weyl-Heisenberg filters or wavelets, the modulus non-linearity, and no pooling. The exponential feature map energy decay results in Wiatowski et al., 2017, are generalized to O(a^{-N}), where an arbitrary decay factor a > 1 can be realized through suitable choice of the Weyl-Heisenberg prototype function or the mother wavelet. We then show how networks of fixed (possibly small) depth N can be designed to guarantee that ((1 − epsilon) · 100)% of the input signal’s energy are contained in the feature vector. Based on the notion of operationally significant nodes, we characterize, partly rigorously and partly heuristically, the topology-reducing effects of (effectively) band-limited input signals, band-limited filters, and feature map symmetries. Finally, for networks based on Weyl-Heisenberg filters, we determine the prototype function bandwidth that minimizes—for fixed network depth N—the average number of operationally significant nodes per layer.

Keywords

Machine learning, deep convolutional neural networks, scattering networks, feature extraction, wavelets, Weyl-Heisenberg frames

Comments

This is a slightly updated version of the paper published in the SPIE proceedings. Specifically, we corrected errors in arguments on spectral decay of Sobolev functions. Moreover, we replaced part of the decay results (Sections 5-7) by corresponding statements for effectively band-limited functions.

Download this document:

This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each author's copyright. In most cases, these works may not be reposted without the explicit permission of the copyright holder.