Language
If man is to use the computer to teach himself, he must be able to converse with it. In the early days of computers it was said with a good deal of justification that the machine was not only stupid but decidedly insular as well. In other words, man spoke to it in its own language or not at all. A host of different languages, or “compilers” as they are often called, were constructed and their originators beat the drums for them. With tongues like ALGY, ALGOL, COBOL, FACT, FLOWMATIC, FORTRAN, INTERCOM, IT, JOVIAL, LOGLAN, MAD, PICE, and PROLAN, to name a few, the computer has become a tower of Babel, and a programmer’s talents must include linguistics.
One language called ALGOL, for Algorithmic Oriented Language, had pretty smooth sailing, since it consists of algebraic and arithmetic notation. Out of the welter of business languages a compromise Common Business Oriented Language, or COBOL, evolved. What COBOL does for programming computer problems is best shown by comparing it with instructions once given the machine. The sample below is typical of early machine language:
SUBTRACT QUANTITY-SOLD FROM BALANCE-ON-HAND. IF BALANCE-ON-HAND IS NOT LESS THAN REORDER-LEVEL THEN GO TO BALANCE-OK ELSE COMPUTE QUANTITY-TO-BUY = TOTAL-SALES-3-MOS/3. ]
Recommended by a task force for the Department of Defense, industry, and other branches of the government, COBOL nevertheless has had a tough fight for acceptance, and there is still argument and confusion on the language scene. New tongues continue to proliferate, some given birth by ALGOL and COBOL themselves. Examples of this generation are GECOM, BALGOL, and TABSOL. One worthy attempt at a sort of machine Esperanto is called a pun-inviting UNCOL, for Universal Computer-Oriented Language and seems to be a try for the computer’s vote. One harried machine-language user has suggested formation of an “ALGOLICS Anonymous” group for others of his ilk, while another partisan accuses his colleagues in Arizona of creating a new language while “maddened by the scent of saguaro blossoms.”
It was recently stated that perhaps by the time a decision is ultimately reached as to which will be the general language, there will be no need of it because by then the computer will have learned to read and write, and perhaps to listen and to speak as well. Recent developments bear out the contention.
Although it has used intermediate techniques, the computer has proved it can do a lot with our language in some of the tasks it has been given. Among these is the preparation of a Bible concordance, listing principal words, frequency of appearance, and where they are found. The computer tackled the same job on the poems of Matthew Arnold. For this chore, Professor Stephen Maxfield Parrish of Cornell worked with three colleagues and two technicians to program an IBM 704 data-processing system. In addition to compiling the list of more than 10,000 words used most often by Arnold, the computer arranged them alphabetically and also compiled an appendix listing the number of times each word appeared. To complete the job, the computer itself printed the 965-page volume. The Dead Sea Scrolls and the works of St. Thomas Aquinas have also been turned over to the computer for preparation of analytical indexes and concordances.
At Columbia University, graduate student James McDonough gave an IBM 650 the job of sleuthing the author of The Iliad and The Odyssey. Since the computer can detect metric-pattern differences otherwise practically undiscoverable, McDonough felt that the machine could prove if Homer had written both poems, or if he had help on either. Thus far he is sure the entire Iliad is the work of one man, after computer analysis of its 112,000 words. The project is part of his doctoral thesis. A recent article in a technical journal used a title suggested by an RCA 501, and suspicion is strong that the machines themselves are guilty of burning midnight kilowatts to produce the acronyms that abound in the industry. The computer is even beginning to prove its worth as an abstracter.
Other literary jobs the computer has done include the production of a book of fares for the International Air Transport Association. The computer compiled and then printed out this 420-page book which gives shortest operating distances between 1,600 cities of the world. Now newspapers are beginning to use computers to do the work of typesetting. These excursions into the written language of human beings, plus its experience as a poet and in translation from language to language, have undoubtedly brought the computer a long way from its former provincialism.
As pointed out, computer work with human language generally is not accomplished without intermediate steps. For example, in one of the concordances mentioned, although the computer required only an hour to breeze through the work, a programmer had spent weeks putting it in the proper shape. What is needed is a converter which will do the work directly, and this is exactly what firms like Digitronics supply to the industry. This computer-age Berlitz school has produced converters for Merrill Lynch, Pierce, Fenner & Smith for use in billing its stock-market customers, Wear-Ever as an order-taking machine, Reader’s Digest for mailing-list work, and Schering Corporation for rat-reaction studies in drug research, to mention a few.
The importance of such converters is obvious. Prior to their use it was necessary to type English manually into the correct code, a costly and time-consuming business. Converters are not cheap, of course, but they operate so rapidly that they pay for themselves in short order. Merrill Lynch’s machine cost $120,000, but paid back two-thirds of that amount in savings the first year. There is another important implication in converter operation. It can get computer language out of English—or Japanese, or even Swahili if the need arises. A more recent Digitronics’ converter handles information in English or Japanese.
If the computer has its language problems, man has them also, to the nth degree. There are about 3,000 tongues in use today; mercifully, scientific reports are published in only about 35 of these. Even so, at least half the treatises published in the world cannot be read by half the world’s scientists. Unfortunately, UNESCO estimates that while 50 per cent of Russian scientists read English, less than 1 per cent of United States scientists return the compliment! The ramifications of these facts we will take up a little later on; for now it will be sufficient to consider the language barrier not only to science but also to culture and the international exchange of good will that can lead to and preserve peace. Esperanto, Io, and other tongues have been tried as common languages. One recent comer to the scientific scene is called Interlingua and seems to have considerable merit. It is used in international medical congresses, with text totaling 300,000 words in the proceedings of one of these. But a truly universal language is, like prosperity, always just around the corner. Even the scientific community, recognizing the many benefits that would accrue, can no more adopt Interlingua or another than it can settle on the metric system of measurement. Our integration problems are not those of race, color, and creed only.
Before Sputnik our interest in foreign technical literature was not as keen as it has been since. One immediate result of the satellite launching by the Russians was amendment of U.S. Public Law 480 to permit money from the sale of American farm equipment abroad to be used for translation of foreign technical literature. We are vitally concerned with Russia, but have also arranged for thousands of pages of scientific literature from Poland, Yugoslavia, and Israel. Communist China is beginning to produce scientific reports too, and Japanese capability in such fields as electronics is evident in the fact that the revolutionary “tunnel diode” was invented by Esaki in Japan.
It is understandable that we should be concerned with the output of Russian literature, and much attention has been given to the Russian-English translator developed by IBM for the Air Force. It is estimated that the Russians publish a billion words a year, and that about one-third of this output is technical in nature. Conventional translating techniques, in addition to being tedious for the translators, are hopelessly slow, retrieving only about 80 million words a year. Thus we are falling behind twelve years each year! Outside of a moratorium on writing, the only solution is faster translation.
The Air Force translator was a phenomenal achievement. Based on a photoscopic memory—a glass disc 10 inches in diameter capable of storing 55,000 words of Russian-English dictionary in binary code—the system used a “one-to-one” method of translation. The result initially was a translation at the rate of about 40 words per minute of Russian into an often terribly scrambled and confusing English. The speed was limited not by the memory or the computer itself but by the input, which had to be prepared on tape by a typist. Subsequently a scanning system capable of 2,400 words a minute upped the speed considerably.
Impressive as the translator was, its impact was dulled after a short time when it was found that a second “translation” was required of the resulting pidgin English, particularly when the content was highly technical. As a result, work is being done on more sophisticated translation techniques. Making use of predictive analysis, and “lexical buffers” which store all the words in a sentence for syntactical analysis before final printout, scientists have improved the translation a great deal. In effect, the computer studies the structure of the sentence, determining whether modifiers belong with subject or object, and checking for the most probable grammatical form of each word as indicated by other words in the sentence.
The advanced nature of this method of translation requires the help of linguistics experts. Among these is Dr. Sydney Lamb of the University of California at Berkeley who is developing a computer program for analysis of the structure of any language. One early result of this study was the realization that not enough is actually known of language structure and that we must backtrack and build a foundation before proceeding with computer translation techniques. Dr. Lamb’s procedure is to feed English text into the computer and let it search for situations in which a certain word tends to be preceded or followed by other words or groups of words. The machine then tries to produce the grammatical structure, not necessarily correctly. The researcher must help the machine by giving it millions of words to analyze contextually.
What the computer is doing in hours is reproducing the evolution of language and grammar that not only took place over thousands of years, but is subject to emotion, faulty logic, and other inaccuracies as well. Also working on the translation problem are the National Bureau of Standards, the Army’s Office of Research and Development, and others. The Army expects to have a computer analysis in 1962 that will handle 95 per cent of the sentences likely to be encountered in translating Russian into English, and to examine foreign technical literature at least as far as the abstract stage.
Difficult as the task seems, workers in the field are optimistic and feel that it will be feasible to translate all languages, even the Oriental, which seem to present the greatest syntactical barriers. An indication of success is the announcement by Machine Translations Inc. of a new technique making possible contextual translation at the rate of 60,000 words an hour, a rate challenging the ability of even someone coached in speed-reading! The remaining problem, that of doing the actual reading and evaluation after translation, has been brought up. This considerable task too may be solved by the computer. The machines have already displayed a limited ability to perform the task of abstracting, thus eliminating at the outset much material not relevant to the task at hand. Another bonus the computer may give us is the ideal international and technical language for composing reports and papers in the first place. A logical question that comes up in the discussion of printed language translation is that of another kind of translation, from verbal input to print, or vice versa. And finally from verbal Russian to verbal English. The speed limitation here, of course, is human ability to accept a verbal input or to deliver an output. Within this framework, however, the computer is ready to demonstrate its great capability.
A recent article in Scientific American asks in its first sentence if a computer can think. The answer to this old chestnut, the authors say, is certainly yes. They then proceed to show that having passed this test the computer must now learn to perceive, if it is to be considered a truly intelligent machine. A computer that can read for itself, rather than requiring human help, would seem to be perceptive and thus qualify as intelligent.
Even early computers such as adding machines printed out their answers. All the designers have to do is reverse this process so that printed human language is also the machine’s input. One of the first successful implementations of a printed input was the use of magnetic ink characters in the Magnetic Ink Character Recognition (MICR) system developed by General Electric. This technique called for the printing of information on checks with special magnetic inks. Processed through high-speed “readers,” the ink characters cause electrical currents the computer can interpret and translate into binary digits.
Close on the heels of the magnetic ink readers came those that use the principle of optical scanning, analogous to the method man uses in reading. This breakthrough came in 1961, and was effected by several different firms, such as Farrington Electronics, National Cash Register, Philco, and others, including firms in Canada and England. We read a page of printed or written material with such ease that we do not realize the complex way our brains perform this miracle, and the optical scanner that “reads” for the computer requires a fantastically advanced technology.
As the material to be read comes into the field of the scanner, it is illuminated so that its image is distinct enough for the optical system to pick up and project onto a disc spinning at 10,000 revolutions per minute. In the disc are tiny slits which pass a certain amount of the reflected light onto a fixed plate containing more slits. Light which succeeds in getting through this second series of slits activates a photoelectric cell which converts the light into proportionate electrical impulses. Because the scanned material is moving linearly and the rotating disc is moving transversely to this motion, the character is scanned in two directions for recognition. Operating with great precision and speed, the scanner reads at the rate of 240 characters a second.
National Cash Register claims a potential reading rate for its scanner of 11,000 characters per second, a value not reached in practice only because of the difficulty of mechanically handling documents at this speed. Used in post-office mail sorting, billing, and other similar reading operations, optical scanners generally show a perfect score for accuracy. Badly printed characters are rejected, to be deciphered by a human supervisor.
It is the optical scanner that increased the speed of the Russian-English translating computer from 40 to 2,400 words per minute. In post-office work, the Farrington scanner sorts mail at better than 9,000 pieces an hour, rejecting all handwritten addresses. Since most mail—85 per cent, the Post Office Department estimates—is typed or printed, the electronic sorter relieves human sorters of most of their task. Mail is automatically routed to proper bins or chutes as fast as it is read.
The electronic readers have not been without their problems. A drug firm in England had so much difficulty with one that it returned it to the manufacturer. We have mentioned the one that was confused by Christmas seals it took for foreign postage stamps. And as yet it is difficult for most machines to read anything but printed material.
An attempt to develop a machine with a more general reading ability, one which recognizes not only material in which exact criteria are met, but even rough approximations, uses the gestalt or all-at-once pattern principle. Using a dilating circular scanning method, the “line drawing pattern recognizer” may make it possible to read characters of varying sizes, handwritten material, and material not necessarily oriented in a certain direction. A developmental model recognizes geometric figures regardless of size or rotation and can count the number of objects in its scope. Such experimental work incidentally yields much information on just how the eye and brain perform the deceptively simply tasks of recognition. Once 1970 had been thought a target date for machine recognition of handwritten material, but researchers at Bell Telephone Laboratories have already announced such a device that reads cursive human writing with an accuracy of 90 per cent.
The computer, a backward child, learned to write long before it could read and does so at rates incomprehensible to those of us who type at the blinding speed of 50 to 60 words a minute. A character-generator called VIDIAC comes close to keeping up with the brain of a high-speed digital computer and has a potential speed of 250,000 characters, or about 50,000 words, per second. It does this, incidentally, by means of good old binary, 1-0 technique. To add to its virtuosity, it has a repertoire of some 300 characters. Researchers elsewhere are working on the problems to be met in a machine for reading and printing out 1,000,000 characters per second!
Computers—the Machines We Think With · The Wunder Library — complete classics, free to read, with narration.