16. Morrell, F., “Electrophysiological Contributions to the Neural Basis of Learning,” Physiological Reviews =41(No. 3)= (1961)
17. Pask, G., “The Growth Process Inside the Cybernetic Machine,” Proc. 2nd Congress International Association Cybernetics, Gauthier-Villars, Paris:Namur, 1958
18. Retzlaff, E., “Neurohistological Basis for the Functioning of Paired Half-Centers,” J. Comp. Neurology =101=:407-443 (1954)
19. Sperry, R. W., “Neurology and the Mind-Brain Problem,” Amer. Scientist =40(No. 2)=: 291-312 (1952)
20. Tasaki, I., and Bak, A. F., J. Gen. Physiol. =42=:899 (1959)
21. Thorpe, W. H., “The Concepts of Learning and Their Relation to Those of Instinct,” S. E. B. Symposia, No. IV, “Physiological Mechanisms in Animal Behavior,” Cambridge:University Press, USA:Academic Press, Inc., 1950
22. Yamagiwa, K., “The Interaction in Various Manifestations (Observations on Lillie’s Nerve Model),” Jap. J. Physiol. =1=:40-54 (1950)
23. Young, J. Z., “The Evolution of the Nervous System and of the Relationship of Organism and Environment,” G. R. de Beer, ed., “Evolution,” Oxford:Clarendon Press, pp. 179-204, 1938
24. Young, J. Z., “Doubt and Certainty in Science, A Biologist’s Reflections on the Brain,” New York:Oxford Press, 1951
Multi-Layer Learning Networks
R. A. STAFFORD
Philco Corp., Aeronutronic Division Newport Beach, California
INTRODUCTION
This paper is concerned with the problem of designing a network of linear threshold elements capable of efficiently adapting its various sets of weights so as to produce a prescribed input-output relation. It is to accomplish this adaptation by being repetitively presented with the various inputs along with the corresponding desired outputs. We will not be concerned here with the further requirement of various kinds of ability to “generalize”—i.e., to tend to give correct outputs for inputs that have not previously occurred when they are similar in some transformed sense to other inputs that have occurred.
In putting forth a model for such an adapting or “learning” network, a requirement is laid down that the complexity of the adaption process in terms of interconnections among elements needed for producing appropriate weight changes, should not greatly exceed that already required to produce outputs from inputs with a static set of weights. In fact, it has been found possible to use the output-from-input computing capacity of the network to help choose proper weight changes by observing the effect on the output of a variety of possible weight changes.
No attempt is made here to defend the proposed network model on theoretical grounds since no effective theory is known at present. Instead, the plausibility of the various aspects of the network model, combined with empirical results must suffice.
SINGLE ELEMENTS
To simplify the problem it is assumed that the network receives a set of two-valued inputs, x₁, x₂, ..., xₙ, and is required to produce only a single two-valued output, y. It is convenient to assign the numerical quantities +1 and -1 to the two values of each variable.
The simplest network would consist of a single linear threshold element with a set of weights, c₀, c₁, c₂, ..., cₙ. These determine the output-input relation or function so that y is +1 or -1 according as the quantity, c₀ + c₁x₁ + c₂x₂ + ... + cₙxₙ, is positive or not, respectively. It is possible for such a single element to exhibit an adaptive behavior as follows. If, for a given set, x₁, x₂, ..., xₙ, the output, y, is correct, then make no changes to the weights. Otherwise change the weights according to the equations
Δc₀ = y* Δcᵢ = y*xᵢ, i = 1,2, ...,n
where y* is the desired output.
It has been shown by a number of people that the weights of such an element are assured of arriving at a set of values which produce the correct output-input relation after a sufficient number of errors, provided that such a set exists. An upper bound on the number of possible errors can be given which depends only on the initial weight values and the logical function to be learned. This does not, however, solve our network problem for two reasons.
First, as the number, n, of inputs gets large, the number of errors to be expected for most functions which can be learned increases to unreasonable values. For example, for n = 6, most such functions result in 500 to 1000 errors compared to an average of 32 errors to be expected in a perfect learning device.
Second, and more important, the fraction of those logical functions which can be generated in a single element becomes vanishingly small as n increases. For example, at n = 6 less than one in each three trillion logical functions is so obtainable.
NETWORKS OF ELEMENTS
It can be demonstrated that if a sufficiently large number of linear threshold elements is used, with the outputs of some being the inputs of others, then a final output can be produced which is any desired logical function of the inputs. The difficulty in such a network lies in the fact that we are no longer provided with a knowledge of the correct output for each element, but only for the final output. If the final output is incorrect there is no obvious way to determine which sets of weights should be altered.
As a result of considerable study and experimentation at Aeronutronic, a network model has been evolved which, it is felt, will get around these difficulties. It consists of four basic features which will now be described.
Positive Interconnecting Weights
Self-Organizing Systems, 1963 · The Wunder Library — complete classics, free to read, with narration.