BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//RLASKEY//CALENDEROUS//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
BEGIN:VEVENT
DTSTAMP:20260922T023026Z
LAST-MODIFIED:20170403T205802Z
DTSTART:20170404T150000Z
DTEND:20170404T160000Z
UID:event1756@bu.edu
URL:http://physics.bu.edu/internal/events/show/1756
SUMMARY:Training and generalization dynamics in deep linear neural networks
DESCRIPTION:Featuring Andrew Saxe \, Harvard\n\nPart of the Biophysics Semi
	nars.\n\nAnatomically\, the brain is deep; and computationally\, deep learn
	ing is known to be hard. How might depth impact learning in the brain? To u
	nderstand the specific ramifications of depth\, I develop the theory of lea
	rning in deep linear neural networks. I will describe exact solutions to th
	e dynamics of learning which specify how every weight in the network evolve
	s over the course of training. The theory answers fundamental questions suc
	h as how learning speed scales with depth\, how structured data sets are em
	bedded into hidden neural representations\, and why unsupervised pretrainin
	g accelerates learning. Turning to generalization error\, we use random mat
	rix theory to analyze the cognitively-relevant "high-dimensional" regime\, 
	where the number of training examples is on the order of or even less than 
	the number of adjustable synapses. We find that generalization error can di
	verge in certain instances if training is run forever\, but that implicit r
	egularization in the form of early stopping and small initial weights subst
	antially improves performance. Next\, we turn to the question of how comple
	x a model should be for optimal generalization. We describe a counter-intui
	tive regime where increasing the complexity of a model can lower both the a
	pproximation error and the estimation error of the system\, resulting in su
	bstantial generalization benefits. This result may help explain the strikin
	g performance of even very large deep network models in practice\, which of
	ten have more parameters than training samples. Finally\, if time permits\,
	 I will describe an example of how these results may begin to inform the dy
	namics of nonlinear networks. In particular\, I will describe a setting in 
	which a nonlinear network can be understood as a collection of linear netwo
	rks that learn in parallel\, yielding learning dynamics which are dominated
	 by the fastest linear network in the collection.
LOCATION:SCI 328\, 590 Commonwealth Avenue\, 02215
STATUS:CONFIRMED
CLASS:PUBLIC
END:VEVENT
END:VCALENDAR
