None of us can talk or listen at much over 250 words a minute, even though we may convince ourselves we read several thousand words in that period of time. A simple test of ability to hear is to play a record or tape at double speed or faster. Our brains just won’t take it. For high-speed applications, then, verbalized input or output for computers is interesting in theory only. However, there are occasions when it would be nice to talk to the computer and have it talk back.
In the early, difficult days of computer development, say when Babbage was working on his analytical engine, the designer probably often spoke to his machine. He would have been stunned to hear a response, of course, but today such a thing is becoming commonplace. IBM has a computer called “Shoebox,” a term both descriptive of size and refreshing in that is not formed of initial capitals from an ad writer’s blurb. You can speak figures to Shoebox, tell it what you want done with them, and it gets busy. This is admittedly a baby computer, and it has a vocabulary of just 16 words. But it takes only 31 transistors to achieve that vocabulary, and jumping the number of transistors to a mere 2,000 would increase its word count to 1,000, which is the number required for Basic English.
The Russians are working in the field of speech recognition too, as are the Japanese. The latter are developing an ambitious machine which will not only accept voice instructions, but also answer in kind. To make a true speech synthetizer, the Japanese think they will need a computer about 5,000 times as fast as any present-day type, so for a while it would seem that we will struggle along with “canned” words appropriately selected from tape memory.
We have mentioned the use of such a tape voice in the computerized ground-controlled-approach landing system for aircraft, and the airline reservation system called Unicall in which a central computer answers a dialed request for space in less than three seconds—not with flashing lights or a printed message but in a loud clear voice. It must pain the computer to answer at the snail-like human speed of 150 words a minute, so it salves its conscience by handling 2,100 inputs without getting flustered.
The writer’s dream, a typewriter that has a microphone instead of keys and clacks away merrily while you talk into it, is a dream no longer. Scientists at Japan’s Kyoto University have developed a computer that does just this. An early experimental model could handle a hundred Japanese monosyllables, but once the breakthrough was made, the Japanese quickly pushed the design to the point where the “Sonotype” can handle any language. At the same time, Bell Telephone Laboratories works on the problem from the other end and has come up with a system for a typewriter that talks. Not far behind these exotic uses of digital computer techniques are such things as automatic translation of telephone or other conversations.
Information Retrieval
It has been estimated that some 445 trillion words are spoken in each 16-hour day by the world’s inhabitants, making ours a noisy planet indeed. To bear out the “noisy” connotation, someone else has reckoned that only about 1 per cent of the sounds we make are real information. The rest are extraneous, incidentally telling us the sex of the speaker, whether or not he has a cold, the state of his upper plate, and so on. It is perhaps a blessing that most of these trillions of words vanish almost as soon as they are spoken. The printed word, however, isn’t so transient; it not only hangs around, but also piles up as well. The pile is ever deeper, technical writings alone being enough to fill seven 24-volume encyclopedias each day, according to one source. As with our speech, perhaps only 1 per cent of this outpouring of print is of real importance, but this does not necessarily make what some have called the Information Explosion any less difficult to cope with.
The letters IR once stood for infra-red; but in the last year or so they have been appropriated by the words “information retrieval,” one of the biggest bugaboos on the scientific horizon. It amounts to saving ourselves from drowning in the fallout from typewriters all over the earth. There are those cool heads who decry the pushing of the panic button, professing to see no exponential increase in literature, but a steady 8 per cent or so each year. The button-pushers see it differently, and they can document a pretty strong case. The technical community is suffering an embarrassment of riches in the publications field.
While a doubling in the output of technical literature has taken the last twelve years or so, the next such increase is expected in half that time. Perhaps the strongest indication that IR is a big problem is the obvious fact that nobody really knows just how much has been, is being, or will be written. For instance, one authority claims technical material is being amassed at the rate of 2,000 pages a minute, which would result in far more than the seven sets of encyclopedias mentioned earlier. No one seems to know for sure how many technical journals there are in the world; it can be “pinpointed” somewhere between 50,000 and 100,000. Selecting one set of figures at random, we learn that in 1960 alone 1,300,000 different technical articles were published in 60,000 journals. Of course there were also 60,000 books on technical subjects, plus many thousands of technical reports that did not make the formal journals, but still might contain the vital bit of information without which a breakthrough will be put off, or a war lost. Our research expenses in the United States ran about $13 billion in 1960, and the guess is they will more than double by 1970. An important part of research should be done in the library, of course, lest our scientist spend his life re-inventing the wheel, as the saying goes.
To back up this saying are specific examples. For instance, a scientific project costing $250,000 was completed a few days before an engineer came across practically the identical work in a report in the library. This was a Russian report incidentally, titled “The Application of Boolean Matrix Algebra to the Analysis and Synthesis of Relay Contact Networks.” In another, happier case, information retrieval saved Esso Research & Engineering Co. a month of work and many thousands of dollars when an alert—or lucky—literature searcher came across a Swedish scientist’s monograph detailing Esso’s proposed exploration. Another literature search obviated tests of more than a hundred chemical compounds. Unfortunately not all researchers do or can search the literature in all cases. There is even a tongue-in-cheek law which governs this phenomenon—“Mooer’s” Law states, “An information system will tend not to be used whenever it is more painful for a customer to have information than for him not to have it.”
As a result, it has been said that if a research project costs less than $100,000 it is cheaper to go ahead with it than to conduct a rigorous search of the literature. Tongue in cheek or not, this state of affairs points up the need for a usable information retrieval system. Fortune magazine reports that 10 per cent of research and development expense could be saved by such a system, and 10 per cent in 1960, remember, would have amounted to $1.3 billion. Thus the prediction that IR will be a $100 million business in 1965 does not seem out of line.
The Center for Documentation at Western Reserve University spends about $6-1/2 simply in acquiring and storing a single article in its files. In 1958 it could search only thirty abstracts of these articles in an hour and realized that more speed was vital if the Center was to be of value. As a result, a GE 225 computer IR system was substituted. Now researchers go through the entire store of literature—about 50,000 documents in 1960—in thirty-five minutes, answering up to fifty questions for “customers.”
International Business Machines Corp.
The document file of this WALNUT information retrieval system contains the equivalent of 3,000 books. A punched-card inquiry system locates the desired filmstrip for viewing or photographic reproduction. ]
International Business Machines Corp.
This image converter of the WALNUT system optically reduces and transfers microfilm to filmstrips for storage. Each strip contains 99 document images. As a document image is transferred from microfilm to filmstrip, the image converter simultaneously assigns image file addresses and punches these addresses into punched cards controlling the conversion process. ]
The key to information retrieval lies in efficient abstracting. It has been customary to let people do this task in the past because there was no other way of getting it done. Unfortunately, man does not do a completely objective job of either preparing or using the abstract, and the result is a two-ended guessing game that wastes time and loses facts in the process. A machine abstracting system, devised by H. Peter Luhn of IBM, picks the words that appear most often and uses them as keys to reduce articles to usable, concise abstracts. A satisfactory solution seems near and will be a big step toward a completely computerized IR system.
For several years there has been a running battle between the computer IR enthusiast and the die-hard “librarian” type who claims that information retrieval is not amenable to anything but the human touch. It is true that adapting the computer to the task of information retrieval did not prove as simple as was hoped. But detractors are in much the same fix as the man with a shovel trying to build a dike against an angry rising sea, who scoffs at the scoop-shovel operator having trouble starting his engine. The wise thing to do is drop the shovel and help the machine. There will be a marriage of both types of retrieval, but Verner Clapp, president of the Washington, D.C., Council on Library Resources, stated at an IR symposium that computers offer the best chance of keeping up with the flood of information.
One sophisticated approach to IR uses symbolic logic, the forte of the digital computer. In a typical reductio ad logic, the following request for information:
An article in English concerning aircraft or spacecraft, written neither before 1937 or after 1957; should deal with laboratory tests leading to conclusions on an adhesive used to bond metal to rubber or plastic; the adhesive must not become brittle with age, must not absorb plasticizer from the rubber adherent, and must have a peel-strength of 20 lbs/in; it must have at least one of these properties—no appreciable solution in fuel and no absorption of solvent.
becomes the logical statement:
KKaVbcPdeCfg, and KAhiKKKNjNklSmn.
Armed with this symbolic abbreviation, the computer can dig quickly into its memory file and come up with the sought-for article or articles.
It has been suggested that the abstracting technique be applied at the opposite end of the cycle with a vengeance amounting to birth control of new articles. A Lockheed Electronics engineer proposes a technical library that not only accepts new material, but also rejects any that is not new. Here, of course, we may be skirting danger of the type risked by human birth control exponents—that of unwittingly depriving the world of a president, or a powerful scientific finding. Perhaps the screening, the function of “garbage disposal,” as one blunt worker puts it, should be left as an after-the-fact measure.
Despite early setbacks, the computer is making progress in the job of information retrieval. Figures of a 300 per cent improvement in efficiency in this new application are cited over the last several years. Operation HAYSTAQ, a Patent Office project in the chemical patent section accounting for one-fifth of all patents, showed a 50 per cent improvement in search speed and 100 per cent in accuracy as a result of using automated methods. Desk-size computer systems with solid-state circuits are being offered for information retrieval.
The number of scientific information centers in this country, starting with one in 1830, reached 59 in 1940 and now stands at 144. Significantly, of 2,000 scientists and engineers working at these centers, 381 are computer people.
Some representative information retrieval applications making good use of computer techniques are the selection of the seven astronauts for the Mercury Project from thousands of jet pilots, Procter & Gamble’s Technical Information Service, demonstration of an electronic law library to the American Bar Association, and Food Machinery and Chemical Corporation’s Central Research Laboratory. The National Science Foundation, the National Bureau of Standards, and the U.S. Patent Office are among the government agencies in addition to the military services that are interested in electronic information retrieval.
Summary
Computers—the Machines We Think With · The Wunder Library — complete classics, free to read, with narration.